3. 模型与基准 (Models & Benchmarks)
- 2026 LLM大模型权威排行榜评测网站指南(持续更新)-网站名
2 minutes ago · 开发团队简介SWE-bench 最初由来自 Princeton University、Princeton Language and IntelligencePLI和 University of Chicago 的研究人员提出。 原始论文作者包括 Carlos E. Jimenez、John Yang、Alexander Wettig、Shunyu Yao、Kexin Pei、Ofir Press 和 Karthik Narasimhan。
- 2026年AI大模型综合排行榜 - AnyRank - AnyRank
2 minutes ago · 截至 2026 年 3 月,目前的 LLM 综合排名,主要参考 Code Arena 和各项基准测试得分。 | 本月榜单显示,Anthropic 和 Google 在第一梯队竞争极其激烈,二者在编程能力(Code Arena / S | 前三名:Claude Opus 4.6、Gemini 3.1 Pro、GPT-5.
- SWE-Bench Verified Leaderboard - llm-stats.com
2 minutes ago · SWE-Bench Verified is a text benchmark evaluating models on reasoning, frontend development, and code tasks. LLM Stats tracks 116 models on this benchmark, scored on a 0– 1 scale.
- Arena AI(LMArena)大模型排行榜·官网镜像 9月15号更新
2 minutes ago · Arena AI 中文镜像 前沿AI大模型双盲评测榜单 基于真实人类偏好(Bradley-Terry 模型)计算大模型 Elo 评分的中文镜像站。数据每 20 分钟自动与官方同步,为您提供最新、最公正的大模型能力排名。
- AI Benchmarks 2026 - MMLU, GPQA, SWE-bench | LM Market Cap
1 day ago · Compare AI models across 17 benchmarks including MMLU, GPQA Diamond, MATH-500, HumanEval, SWE-bench, and Arena Elo. See current leaders, score history, and interactive charts for 350+ models.
- Open-Source LLM Leaderboard 2026 | LM Market Cap
1 day ago · Open LLM leaderboard ranking the best open-source language models by benchmarks, pricing, and capabilities. Compare Llama, DeepSeek, Qwen, Mistral, and Gemma with live scores.