Saotri Bench — coding benchmark for evaluating LLM agents on multi-phase programming tasks with hidden requirements.
-
Updated
Feb 21, 2026 - Python
Saotri Bench — coding benchmark for evaluating LLM agents on multi-phase programming tasks with hidden requirements.
What extra reasoning buys in code and tests: 216 Sol coding attempts, MVCC defect detection, DAG execution, effort premiums and API-equivalent USD, with a reproducible stand and English report.
面向 AI IDE / 编码 Agent 的编程能力基准测试框架,内置防背题三防线:公私用例分离、泛化差距检测、参数化生成
Raw logs of Claude Code running on local Qwen3.5-27B (llama.cpp). Builds a Python todo app with 50 tests. Real-world performance data: 30 min, cache thrashing, 38 t/s generation.
Benchmark and tune every locally installed Ollama model on your own GPUs. Finds the best model per GPU setup plus its optimal temperature and context, with a no-spill VRAM guard, coding + agentic test suites, and a single-file HTML report. 100% local: no cloud, no API keys, no model judges.
To associate your repository with the coding-benchmark topic, visit your repo's landing page and select "manage topics."