3. 模型与基准 (Models & Benchmarks)
- AI 编程工具—Cursor进阶使用deepseek V3 模型(deepseek + cursor...
今日模型配置页面,这里我们勾掉之前已经激活的模型.创建deepseek-chat 的模型后,选中这个模型,然后在下面的配置框中输入我们刚才生成的API keys.
- LLM Leaderboard 2026 — Compare 256 AI Models... | BenchLM.ai
The BenchLM LLM leaderboard 2026 provisionally ranks 122+ models and tracks 256+ large language models side by side across 236 benchmarks — from SWE-bench and LiveCodeBench for coding to GPQA Diamond and MMLU-Pro for knowledge and reasoning.
- Multi-SWE-bench...
Multi-SWE-bench 旨在补全现有同类基准语言覆盖方面的不足,系统性评估大模型在复杂开发环境下的“多语言泛化能力”,推动多语言软件开发 Agent 的评估与研究,其.
- SWE-bench Leaderboards
SWE-bench Multilingual features 300 tasks across 9 programming languages [Post]. SWE-bench Lite is a subset curated for less costly evaluation [Post].
- SWE-Bench Pro Leaderboard | LLM Stats
SWE-Bench Pro is a text benchmark that evaluates large language models on reasoning, agents, and code tasks. LLM Stats tracks 23 models on this benchmark, with a maximum possible score of 1. Current average across reported models is 0.6, with the leader reaching 0.8.
- GitHub - cynic12138/Webui-llm-eval · GitHub
Contribute to cynic12138/Webui-llm-eval development by creating an account on GitHub.自研 Python 评测框架,支持 OpenAI/Anthropic/自定义 API.