Bugs fixed /105 · Bug Hunt Bench Results Across LLMs
Independent evaluator Bugs fixed /105 · Bug Hunt Bench results across showcased LLMs. Recorded Bugs fixed /105 measurements from Bug Hunt Bench, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.
What this table contains
Recorded Bugs fixed /105 measurements from Bug Hunt Bench, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.
20 models with results · 10 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.
Interactive results & effort preference → · Results over time → · Methodology
| Model | Headline score | Source/record date | Evidence |
|---|---|---|---|
| Claude Sonnet 5.5Anthropic | 51.3 fixesmax effort | 2026-10-03 | All 5 recorded results & sources51.3 fixes · raw 51.3 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort max; 3 runs; evaluation 2026-09-29. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench36 fixes · raw 36 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort xhigh; 3 runs; evaluation 2026-09-29. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench34.3 fixes · raw 34.3 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort high; 3 runs; evaluation 2026-09-29. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench23.3 fixes · raw 23.3 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort medium; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench19 fixes · raw 19 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort low; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| GPT-6 AstraOpenAI | 45 fixesmax effort | 2026-10-03 | All 5 recorded results & sources45 fixes · raw 45 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort max; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench43 fixes · raw 43 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort xhigh; 1 runs; evaluation 2026-09-04. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench35 fixes · raw 35 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort high; 1 runs; evaluation 2026-09-04. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench34 fixes · raw 34 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort medium; 1 runs; evaluation 2026-09-05. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench27 fixes · raw 27 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort low; 1 runs; evaluation 2026-09-05. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| GPT-6.1 SolOpenAI | 44.3 fixesmax effort | 2026-10-03 | All 5 recorded results & sources44.3 fixes · raw 44.3 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort max; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench42.7 fixes · raw 42.7 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort xhigh; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench36.5 fixes · raw 36.5 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort high; 2 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench29 fixes · raw 29 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort medium; 2 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench22.5 fixes · raw 22.5 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort low; 2 runs; evaluation 2026-10-01. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Claude Fable 5.1Anthropic | 43 fixesmax effort | 2026-10-03 | All 5 recorded results & sources43 fixes · raw 43 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort max; 1 runs; evaluation 2026-09-01. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench33 fixes · raw 33 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort high; 1 runs; evaluation 2026-09-01. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench29 fixes · raw 29 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort low; 1 runs; evaluation 2026-09-02. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench29 fixes · raw 29 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort xhigh; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench21 fixes · raw 21 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort medium; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Claude Opus 5.5Anthropic | 41.7 fixesmax effort | 2026-10-03 | All 5 recorded results & sources41.7 fixes · raw 41.7 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort max; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench36 fixes · raw 36 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort xhigh; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench31.7 fixes · raw 31.7 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort high; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench30.3 fixes · raw 30.3 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort medium; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench22.3 fixes · raw 22.3 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code; effort low; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Muse Spark 1.3Meta | 32.2 fixesmax effort | 2026-10-03 | All 5 recorded results & sources32.2 fixes · raw 32.2 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Muse Code / Meta API; effort max; 5 runs; evaluation 2026-09-17. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench20.3 fixes · raw 20.3 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Muse Code / Meta API; effort xhigh; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench18.7 fixes · raw 18.7 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Muse Code / Meta API; effort high; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench13 fixes · raw 13 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Muse Code / Meta API; effort medium; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench9.7 fixes · raw 9.7 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Muse Code / Meta API; effort low; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| GPT-5.6 TerraOpenAI | 32 fixesmax effort | 2026-10-03 | All 4 recorded results & sources32 fixes · raw 32 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort max; 1 runs; evaluation 2026-08-27. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench20 fixes · raw 20 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort xhigh; 1 runs; evaluation 2026-08-28. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench18 fixes · raw 18 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort high; 1 runs; evaluation 2026-08-27. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench15 fixes · raw 15 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort medium; 1 runs; evaluation 2026-08-27. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Grok 4.7xAI | 28.8 fixesxhigh effort | 2026-10-03 | All 4 recorded results & sources28.8 fixes · raw 28.8 fixes Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Grok Build CLI (ACP); effort xhigh; 4 runs; evaluation 2026-09-21. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench26.7 fixes · raw 26.7 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Grok Build CLI (ACP); effort medium; 3 runs; evaluation 2026-09-21. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench19.3 fixes · raw 19.3 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Grok Build CLI (ACP); effort high; 3 runs; evaluation 2026-09-21. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench15.7 fixes · raw 15.7 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Grok Build CLI (ACP); effort low; 3 runs; evaluation 2026-09-21. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Qwen3.8 FlashAlibaba (Qwen) | 26 fixesmax effort | 2026-10-03 | All 2 recorded results & sources26 fixes · raw 26 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / Alibaba API; effort max; 1 runs; evaluation 2026-09-11. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench23 fixes · raw 23 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / Alibaba API; effort low; 1 runs; evaluation 2026-09-11. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Qwen3.8 Max (0902)Alibaba (Qwen) | 25.7 fixesmax effort | 2026-09-23 | All 1 recorded result & sources25.7 fixes · raw 25.7 fixes Headline · max effort · Independent evaluator Source/record date: 2026-09-23 Claude Code / Alibaba API max; 3-run mean — listed as Qwen3.8-Max; board data/benchmark.json updated 2026-09-23 https://github.com/phuryn/bug-hunt-bench |
| MiMo-V2.6-FlashXiaomi | 23.3 fixesunknown effort | 2026-10-03 | All 1 recorded result & sources23.3 fixes · raw 23.3 fixes Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / OpenRouter; effort default; 3 runs; evaluation 2026-09-22. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench |
| MiMo-V2.6-ProXiaomi | 22.7 fixesunknown effort | 2026-10-03 | All 1 recorded result & sources22.7 fixes · raw 22.7 fixes Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / OpenRouter; effort default; 3 runs; evaluation 2026-09-22. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench |
| DeepSeek V4.1 FlashDeepSeek | 21.7 fixesmax effort | 2026-10-03 | All 2 recorded results & sources21.7 fixes · raw 21.7 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / DeepSeek API; effort max; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench19 fixes · raw 19 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / DeepSeek API; effort high; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Kimi K3Moonshot AI | 21 fixesunknown effort | 2026-10-03 | All 1 recorded result & sources21 fixes · raw 21 fixes Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / OpenRouter; effort default; 1 runs; evaluation 2026-07-26. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench |
| GLM-5.3Z.ai | 19 fixesmax effort | 2026-10-03 | All 1 recorded result & sources19 fixes · raw 19 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / Z.ai API; effort max; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench |
| GPT-6 LunaOpenAI | 18.3 fixesmax effort | 2026-10-03 | All 5 recorded results & sources18.3 fixes · raw 18.3 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort max; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench14 fixes · raw 14 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort xhigh; 1 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench9 fixes · raw 9 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort high; 1 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench4 fixes · raw 4 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort medium; 1 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench4 fixes · raw 4 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Codex CLI; effort low; 1 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Gemini 3.8 FlashGoogle | 18 fixeshigh effort | 2026-10-03 | All 1 recorded result & sources18 fixes · raw 18 fixes Headline · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Antigravity CLI; effort high; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench |
| Hy4 previewTencent | 18 fixesunknown effort | 2026-10-03 | All 1 recorded result & sources18 fixes · raw 18 fixes Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / OpenRouter; effort default; 1 runs; evaluation 2026-08-28. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench |
| GLM-5.3-FlashZ.ai | 17.7 fixesmax effort | 2026-10-03 | All 4 recorded results & sources17.7 fixes · raw 17.7 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / Z.ai API; effort max; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench16 fixes · raw 16 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / Z.ai API; effort high; 1 runs; evaluation 2026-09-13. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench13 fixes · raw 13 fixes Alternative · unknown effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / OpenRouter; effort default; 1 runs; evaluation 2026-08-27. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench9 fixes · raw 9 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / Z.ai API; effort low; 1 runs; evaluation 2026-09-13. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Qwen3.8 27BAlibaba (Qwen) | 15 fixesxhigh effort | 2026-10-03 | All 1 recorded result & sources15 fixes · raw 15 fixes Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / Alibaba API; effort xhigh; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench |
| Gemini 4 ArgonGoogle | — | — | No recorded result |
| Kolibri-1Aleph Alpha | — | — | No recorded result |
| Ling 3.0 Flash VLinclusionAI | — | — | No recorded result |
| Ling 3.1 FlashinclusionAI | — | — | No recorded result |
| Mercury 2.5Inception | — | — | No recorded result |
| MiMo-V2.6-Distill-Qwen-9BXiaomi | — | — | No recorded result |
| Naive-N0.5-FlashNaiveAI | — | — | No recorded result |
| Nex-N2.5-ProNex AGI | — | — | No recorded result |
| Pareto 26.10 PreviewUnbiased | — | — | No recorded result |
| Qwen3.8 Omni FlashAlibaba (Qwen) | — | — | No recorded result |
Use the fair-comparison guide before interpreting results from different configurations.