3. 模型与基准 (Models & Benchmarks)
- AI & LLM Benchmarks 2026: Rankings, Scores & Results
13 minutes ago · AI & LLM Benchmarks 2026 Explore live AI and LLM benchmark rankings across reasoning, coding, math, vision, agents, and tool use. Compare composite indexes first, then open individual evaluations for score provenance, coverage, and methodology.
- 2026 年 8 月大模型性能全景:Claude 登顶文本、Kimi 称霸编码,多模...
2 minutes ago · SWE-bench Verified 各家评测口径差异较大未列入本表以各厂商官方发布为准。 再叠加一个人类偏好信号。 LMArena 前端编码榜Frontend Code Arena上Kimi K3 以约 1679 Elo 登顶第一超过 Claude Fable 5~1631与 GPT-5.6 Sol~1618七个前端子领域里拿下六个唯一输掉的是 Gaming第一名是 Fable 5。
- DeepSeek-V4-Pro - AI模型价格对比 (2026/9/29)
3 minutes ago · DeepSeek V4 Pro 是 DeepSeek 推出的大规模混合专家模型,总参数量为 1.6T(万亿),激活参数量为 49B(十亿),支持 100 万 Token 的上下文窗口。 该模型专为高级推理、编程以及长周期智能体工作流而设计,在知识、数学和软件工程基准测试中均表现出色。
- AI Benchmarks 2026 - MMLU, GPQA, SWE-bench | LM Market Cap
1 day ago · Compare AI models across 17 benchmarks including MMLU, GPQA Diamond, MATH-500, HumanEval, SWE-bench, and Arena Elo. See current leaders, score history, and interactive charts for 350+ models.
- 火山引擎 DeepSeek 落地实践分享:企业如何用好推理模型?|算法|大模...
1 day ago · AI新浪潮观察 16min read 火山引擎 DeepSeek 落地实践分享:企业如何用好推理模型? Founder Park 2026/09/28 摘要 教育、设计、代码生成这三类最适合使用推理模型。 本文首发于 Founder Park 公众号 · 2025 年 2 月 26 日 DeepSeek R1 上线之后,火山引擎是部署 R1 最快的云平台之一。 如今 R1 发布已经过去一个多月的 ...
- SWE-bench Verified Benchmark - AI Code Generation Leaderboard ...
1 day ago · Software Engineering Benchmark (Verified): Can a model resolve real GitHub issues from popular Python repositories? Human-validated subset ensures accurate evaluation. Tests end-to-end software engineering ability. See which AI models score highest on SWE-bench Verifi