3. 模型与基准 (Models & Benchmarks)
- DeepSeek与开源大模型:本地部署实战指南-长河编程
2 minutes ago · DeepSeek与开源大模型:本地部署实战指南开源大模型让AI普惠成为可能。本文详解如何利用DeepSeek等开源模型实现本地部署,兼顾性能与隐私。一、开源大模型时代:2026年的格局 1.1 从闭源垄断到开源崛起 2023年以前,大语言模型领域几…
- 【模型架构篇09】国产大模型生态:DeepSeek、Qwen与智谱
2 minutes ago · 总结与展望国产大模型的三大阵营阵营代表策略优势技术驱动型DeepSeek、智谱、阿里Qwen开源 技术领先全球影响力、社区生态场景驱动型百度、字节、腾讯闭源 生态绑定产品化、商业闭环垂直专精型Kimi、MiniMax聚焦特定场景差异化、用户体验2026年国产大模型趋势 ...
- The Eval Index — every LLM evaluation & benchmark tool ...
11 hours ago · A living leaderboard of LLM & agent evaluation, benchmark and red-teaming tools, ranked daily from live GitHub signals.
- AI Benchmarks 2026 - MMLU, GPQA, SWE-bench | LM Market Cap
1 day ago · Compare AI models across 17 benchmarks including MMLU, GPQA Diamond, MATH-500, HumanEval, SWE-bench, and Arena Elo. See current leaders, score history, and interactive charts for 350+ models.
- Arena Elo Benchmark - AI Model Leaderboard (2026) | LM Market Cap
1 day ago · LMSYS Chatbot Arena Elo Rating: Human preference rating from 6M+ crowdsourced blind head-to-head comparisons. Users chat with two anonymous models and pick the better response. See which AI models score highest on Arena Elo. Updated rankings with scores from 350+ mode
- Agent效果测评:评测用例怎么从0攒起来 - CSDN博客
1 day ago · Agent 的能力由模型和 harness 共同决定,同一模型换一套 harness,在 SWE-bench、Terminal-bench 等 评测 上的得分可相差十几甚至二十多个百分点,差距堪比模型换代。 SWE-bench 分数由三个变量决定:底层模型、harness 设计,以及具体的 评测 任务集。