Examenos

Bugs fixed /105 · Bug Hunt Bench Results Across LLMs

Independent evaluator Bugs fixed /105 · Bug Hunt Bench results across showcased LLMs. Recorded Bugs fixed /105 measurements from Bug Hunt Bench, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.

Independent evaluator Higher is better · unit: fixes

What this table contains

Recorded Bugs fixed /105 measurements from Bug Hunt Bench, with source links, evaluation dates and available effort or harness records. Vendor results remain separate.

20 models with results · 10 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.

Interactive results & effort preference → · Results over time → · Methodology

Bugs fixed /105 · Bug Hunt Bench: independent evaluator results only
ModelHeadline scoreSource/record dateEvidence
Claude Sonnet 5.5Anthropic51.3 fixesmax effort2026-10-03
All 5 recorded results & sources

51.3 fixes · raw 51.3 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort max; 3 runs; evaluation 2026-09-29. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

36 fixes · raw 36 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort xhigh; 3 runs; evaluation 2026-09-29. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

34.3 fixes · raw 34.3 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort high; 3 runs; evaluation 2026-09-29. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

23.3 fixes · raw 23.3 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort medium; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

19 fixes · raw 19 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort low; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
GPT-6 AstraOpenAI45 fixesmax effort2026-10-03
All 5 recorded results & sources

45 fixes · raw 45 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort max; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

43 fixes · raw 43 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort xhigh; 1 runs; evaluation 2026-09-04. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

35 fixes · raw 35 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort high; 1 runs; evaluation 2026-09-04. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

34 fixes · raw 34 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort medium; 1 runs; evaluation 2026-09-05. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

27 fixes · raw 27 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort low; 1 runs; evaluation 2026-09-05. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
GPT-6.1 SolOpenAI44.3 fixesmax effort2026-10-03
All 5 recorded results & sources

44.3 fixes · raw 44.3 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort max; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

42.7 fixes · raw 42.7 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort xhigh; 3 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

36.5 fixes · raw 36.5 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort high; 2 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

29 fixes · raw 29 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort medium; 2 runs; evaluation 2026-09-30. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

22.5 fixes · raw 22.5 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort low; 2 runs; evaluation 2026-10-01. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Claude Fable 5.1Anthropic43 fixesmax effort2026-10-03
All 5 recorded results & sources

43 fixes · raw 43 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort max; 1 runs; evaluation 2026-09-01. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

33 fixes · raw 33 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort high; 1 runs; evaluation 2026-09-01. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

29 fixes · raw 29 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort low; 1 runs; evaluation 2026-09-02. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

29 fixes · raw 29 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort xhigh; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

21 fixes · raw 21 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort medium; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Claude Opus 5.5Anthropic41.7 fixesmax effort2026-10-03
All 5 recorded results & sources

41.7 fixes · raw 41.7 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort max; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

36 fixes · raw 36 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort xhigh; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

31.7 fixes · raw 31.7 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort high; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

30.3 fixes · raw 30.3 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort medium; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

22.3 fixes · raw 22.3 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code; effort low; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Muse Spark 1.3Meta32.2 fixesmax effort2026-10-03
All 5 recorded results & sources

32.2 fixes · raw 32.2 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Muse Code / Meta API; effort max; 5 runs; evaluation 2026-09-17. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

20.3 fixes · raw 20.3 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Muse Code / Meta API; effort xhigh; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

18.7 fixes · raw 18.7 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Muse Code / Meta API; effort high; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

13 fixes · raw 13 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Muse Code / Meta API; effort medium; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

9.7 fixes · raw 9.7 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Muse Code / Meta API; effort low; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
GPT-5.6 TerraOpenAI32 fixesmax effort2026-10-03
All 4 recorded results & sources

32 fixes · raw 32 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort max; 1 runs; evaluation 2026-08-27. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

20 fixes · raw 20 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort xhigh; 1 runs; evaluation 2026-08-28. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

18 fixes · raw 18 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort high; 1 runs; evaluation 2026-08-27. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

15 fixes · raw 15 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort medium; 1 runs; evaluation 2026-08-27. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Grok 4.7xAI28.8 fixesxhigh effort2026-10-03
All 4 recorded results & sources

28.8 fixes · raw 28.8 fixes

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Grok Build CLI (ACP); effort xhigh; 4 runs; evaluation 2026-09-21. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

26.7 fixes · raw 26.7 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Grok Build CLI (ACP); effort medium; 3 runs; evaluation 2026-09-21. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

19.3 fixes · raw 19.3 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Grok Build CLI (ACP); effort high; 3 runs; evaluation 2026-09-21. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

15.7 fixes · raw 15.7 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Grok Build CLI (ACP); effort low; 3 runs; evaluation 2026-09-21. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Qwen3.8 FlashAlibaba (Qwen)26 fixesmax effort2026-10-03
All 2 recorded results & sources

26 fixes · raw 26 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Alibaba API; effort max; 1 runs; evaluation 2026-09-11. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

23 fixes · raw 23 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Alibaba API; effort low; 1 runs; evaluation 2026-09-11. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Qwen3.8 Max (0902)Alibaba (Qwen)25.7 fixesmax effort2026-09-23
All 1 recorded result & sources

25.7 fixes · raw 25.7 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

Claude Code / Alibaba API max; 3-run mean — listed as Qwen3.8-Max; board data/benchmark.json updated 2026-09-23

https://github.com/phuryn/bug-hunt-bench
MiMo-V2.6-FlashXiaomi23.3 fixesunknown effort2026-10-03
All 1 recorded result & sources

23.3 fixes · raw 23.3 fixes

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / OpenRouter; effort default; 3 runs; evaluation 2026-09-22. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench
MiMo-V2.6-ProXiaomi22.7 fixesunknown effort2026-10-03
All 1 recorded result & sources

22.7 fixes · raw 22.7 fixes

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / OpenRouter; effort default; 3 runs; evaluation 2026-09-22. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench
DeepSeek V4.1 FlashDeepSeek21.7 fixesmax effort2026-10-03
All 2 recorded results & sources

21.7 fixes · raw 21.7 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / DeepSeek API; effort max; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

19 fixes · raw 19 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / DeepSeek API; effort high; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Kimi K3Moonshot AI21 fixesunknown effort2026-10-03
All 1 recorded result & sources

21 fixes · raw 21 fixes

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / OpenRouter; effort default; 1 runs; evaluation 2026-07-26. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench
GLM-5.3Z.ai19 fixesmax effort2026-10-03
All 1 recorded result & sources

19 fixes · raw 19 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Z.ai API; effort max; 1 runs; evaluation 2026-09-10. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench
GPT-6 LunaOpenAI18.3 fixesmax effort2026-10-03
All 5 recorded results & sources

18.3 fixes · raw 18.3 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort max; 3 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

14 fixes · raw 14 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort xhigh; 1 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

9 fixes · raw 9 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort high; 1 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

4 fixes · raw 4 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort medium; 1 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

4 fixes · raw 4 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Codex CLI; effort low; 1 runs; evaluation 2026-09-23. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Gemini 3.8 FlashGoogle18 fixeshigh effort2026-10-03
All 1 recorded result & sources

18 fixes · raw 18 fixes

Headline · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Antigravity CLI; effort high; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench
Hy4 previewTencent18 fixesunknown effort2026-10-03
All 1 recorded result & sources

18 fixes · raw 18 fixes

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / OpenRouter; effort default; 1 runs; evaluation 2026-08-28. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench
GLM-5.3-FlashZ.ai17.7 fixesmax effort2026-10-03
All 4 recorded results & sources

17.7 fixes · raw 17.7 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Z.ai API; effort max; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

16 fixes · raw 16 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Z.ai API; effort high; 1 runs; evaluation 2026-09-13. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

13 fixes · raw 13 fixes

Alternative · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / OpenRouter; effort default; 1 runs; evaluation 2026-08-27. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

9 fixes · raw 9 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Z.ai API; effort low; 1 runs; evaluation 2026-09-13. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Qwen3.8 27BAlibaba (Qwen)15 fixesxhigh effort2026-10-03
All 1 recorded result & sources

15 fixes · raw 15 fixes

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Alibaba API; effort xhigh; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench
Gemini 4 ArgonGoogle——No recorded result
Kolibri-1Aleph Alpha——No recorded result
Ling 3.0 Flash VLinclusionAI——No recorded result
Ling 3.1 FlashinclusionAI——No recorded result
Mercury 2.5Inception——No recorded result
MiMo-V2.6-Distill-Qwen-9BXiaomi——No recorded result
Naive-N0.5-FlashNaiveAI——No recorded result
Nex-N2.5-ProNex AGI——No recorded result
Pareto 26.10 PreviewUnbiased——No recorded result
Qwen3.8 Omni FlashAlibaba (Qwen)——No recorded result

Use the fair-comparison guide before interpreting results from different configurations.