3. 模型与基准 (Models & Benchmarks)
- AI & LLM Benchmarks 2026: Rankings, Scores & Results
35 minutes ago · AI & LLM Benchmarks 2026 Explore live AI and LLM benchmark rankings across reasoning, coding, math, vision, agents, and tool use. Compare composite indexes first, then open individual evaluations for score provenance, coverage, and methodology.
- DeepSeek-V4-Flash - AI模型价格对比 (2026/9/26)
2 minutes ago · DeepSeek V4 Flash 是 DeepSeek 开发的效率优化的专家混合模型,拥有 284B 总参数量和 13B 激活参数量,支持 1M Token的上下文窗口。它专为快速推理和高吞吐量工作负载设计,同时保持强大的推理和编码性能。 该模型包含混合注意力机制,用于高效处理长上下文,并支持可配置的推理模式。它非
- AI Model Rankings: Live LLM Leaderboard | ModelCap
5 hours ago · The list ranks models measured on at least two admitted coding boards (Arena Coding, SWE-bench bash-only, Terminal-Bench 2.1 with Terminus 2) by the Index's coding component, then lists models measured on a single board by that board's score, with the figures beside
- AI Benchmarks 2026 - MMLU, GPQA, SWE-bench | LM Market Cap
1 day ago · Compare AI models across 17 benchmarks including MMLU, GPQA Diamond, MATH-500, HumanEval, SWE-bench, and Arena Elo. See current leaders, score history, and interactive charts for 350+ models.
- DeepSeek API Pricing — R1, V3 & Chat Model Costs (2026)
1 day ago · Complete DeepSeek API pricing for R1, V3, and Chat models. Compare DeepSeek costs per million tokens with OpenAI and Anthropic. Updated hourly.
- 2026 年 OpenRouter 排行榜深度解读:Top 10...
为什么 2026 年要看 OpenRouter 排行榜而不是只看 Benchmark?Hy3 Preview 则以腾讯混元 3 的开源 MoE(295B 总量、约 21B 激活)承接私有化与 STEM Agent 需求,SWE-bench Verified 约 74.4% 档,与 Kimi K2.5 同级竞争。