Beam
Explore Beam from Reflection AI: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare Beam with other models →Calculate costExplore data coverage
Pricing · OpenRouter
No OpenRouter pricing is recorded for this model. Missing rates are not free.
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| DeepSWE v1.1 | 44.4%unknown effort | All 1 recorded result & sources44.4% · raw 44.4 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| SWE-bench Pro v2-Hard | 77.2%unknown effort | All 1 recorded result & sources77.2% · raw 77.2 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| SWE-bench Pro v1 | 65.5%unknown effort | All 1 recorded result & sources65.5% · raw 65.5 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| Terminal-Bench 2.1 | 80.1%unknown effort | All 1 recorded result & sources80.1% · raw 80.1 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| SWE-Atlas Codebase QnA | 34.6%unknown effort | All 1 recorded result & sources34.6% · raw 34.6 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| SWE-bench Multilingual | 78%unknown effort | All 1 recorded result & sources78% · raw 78 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| SWE-bench Verified | 80.9%unknown effort | All 1 recorded result & sources80.9% · raw 80.9 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| AIME 2026 | 97.8%unknown effort | All 1 recorded result & sources97.8% · raw 97.8 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. Language and evaluation protocol not specified. https://reflection.ai/blog/introducing-beam |
| HLE (text, no tools) | 36.2%unknown effort | All 1 recorded result & sources36.2% · raw 36.2 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. No tools; corresponding efficiency figure labels HLE text-only. https://reflection.ai/blog/introducing-beam |
| SciCode | 49.7%unknown effort | All 1 recorded result & sources49.7% · raw 49.7 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| CriPT (AA) | 16.3%unknown effort | All 1 recorded result & sources16.3% · raw 16.3 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| GPQA Diamond | 90.5%unknown effort | All 1 recorded result & sources90.5% · raw 90.5 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| AutomationBench (public split) | 37%unknown effort | All 1 recorded result & sources37% · raw 37 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| MCP Atlas | 78.7%unknown effort | All 1 recorded result & sources78.7% · raw 78.7 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| τ³-Bench Banking | 38%unknown effort | All 1 recorded result & sources38% · raw 38 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| BrowseComp | 77.4%unknown effort | All 1 recorded result & sources77.4% · raw 77.4 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. Context management enabled; other agent settings unspecified. https://reflection.ai/blog/introducing-beam |
| DeepSearchQA | 80.1%unknown effort | All 1 recorded result & sources80.1% · raw 80.1 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. Context management enabled; other agent settings unspecified. https://reflection.ai/blog/introducing-beam |
| AA-LCR | 79.3%unknown effort | All 1 recorded result & sources79.3% · raw 79.3 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| LongBench v2 | 65.5%unknown effort | All 1 recorded result & sources65.5% · raw 65.5 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| IFBench | 79.7%unknown effort | All 1 recorded result & sources79.7% · raw 79.7 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. https://reflection.ai/blog/introducing-beam |
| AA-Omniscience Index (public set) | 13unknown effort | All 1 recorded result & sources13 · raw 13 index Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection launch evaluation; full agent, revision and token-budget settings not published. Beam early-access preview, reported by Reflection. Public split only, raw index (not percent), evaluated by Reflection; no full/private AA subset implied. https://reflection.ai/blog/introducing-beam |
| Land or Water grid (180×90 demo) | 95.5%unknown effort | All 1 recorded result & sources95.5% · raw 95.5 % Headline · unknown effort · Own vendor Source/record date: 2026-10-05 Harness: Reflection single-prompt land/water demo, 180×90 grid / 16,200 points; vendor reports 95.5% correct coverage. Grader, repeat count and effort not published. Comparison names Opus 5 and Fable 5 are not the site’s Opus 5.5/Fable 5.1 snapshots; peer results kept only in research notes. https://reflection.ai/blog/introducing-beam |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
No independent evaluator results are recorded for this model.
Read how we select and source scores or the comparison guide.