Kolibri-1 Benchmarks, Specifications & Availability
Explore Kolibri-1 from Aleph Alpha: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare Kolibri-1 with other models →Explore data coverage
Published specifications
- Provider
- Aleph Alpha
- Access
- Open weights
- License
- Apache 2.0
- Context window
- 1.05M
- Total parameters
- 78B
- Active parameters
- 3.46B
- Released
- 2026-10-03
- Modalities
- text
- Family
- Kolibri
Model card · Announcement · Website
Model notes
Official ungated FP8 weights on Hugging Face; Apache 2.0. Card states 78,103,074,560 total and 3,457,573,120 active parameters. English/German reasoning and tool calling. Native context 262,144 tokens; validated extrapolation to 1,048,576 tokens; vendor recommends <=262,144 for serving efficiency and complex tasks. Benchmarks at high reasoning effort in Aleph Alpha eval-framework/Harbor. No OpenRouter listing on 2026-10-03. BF16 is a precision variant, not a separate model.
Pricing · OpenRouter
No OpenRouter pricing is recorded for this model. Missing rates are not free.
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Aleph Alpha post-training Overall (EN) | 75.5%high effort | All 1 recorded result & sources75.5% · raw 75.5 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha post-training Overall (DE) | 70.8%high effort | All 1 recorded result & sources70.8% · raw 70.8 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Knowledge average (EN) | 50.1%high effort | All 1 recorded result & sources50.1% · raw 50.1 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Knowledge average (DE) | 57.6%high effort | All 1 recorded result & sources57.6% · raw 57.6 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| GPQA Diamond | 84.3%high effort | All 1 recorded result & sources84.3% · raw 84.3 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Chain-of-thought, avg@8. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| GPQA Diamond (DE) | 81.3%high effort | All 1 recorded result & sources81.3% · raw 81.3 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| HLE (text, no tools) | 21.5%high effort | All 1 recorded result & sources21.5% · raw 21.5 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Text-only questions, no output cap. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Humanity's Last Exam (DE) | 15.9%high effort | All 1 recorded result & sources15.9% · raw 15.9 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AA-Omniscience Accuracy (public set) | 14.8%high effort | All 1 recorded result & sources14.8% · raw 14.8 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AA-Omniscience Index (public set) | -32.8high effort | All 1 recorded result & sources-32.8 · raw -32.8 index Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| MMLU-Pro CoT (EN) | 80%high effort | All 1 recorded result & sources80% · raw 80 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| MMLU-ProX CoT (DE) | 75.5%high effort | All 1 recorded result & sources75.5% · raw 75.5 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Math average (EN) | 96.5%high effort | All 1 recorded result & sources96.5% · raw 96.5 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Math average (DE) | 88.8%high effort | All 1 recorded result & sources88.8% · raw 88.8 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AIME 2025 (EN) | 96.9%high effort | All 1 recorded result & sources96.9% · raw 96.9 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). avg@16. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AIME 2025 (DE) | 87.5%high effort | All 1 recorded result & sources87.5% · raw 87.5 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. avg@16. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AIME 2026 (EN) | 96%high effort | All 1 recorded result & sources96% · raw 96 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). avg@16. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AIME 2026 (DE) | 90%high effort | All 1 recorded result & sources90% · raw 90 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. avg@16. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Agentic average (EN) | 63.4%high effort | All 1 recorded result & sources63.4% · raw 63.4 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Terminal-Bench 2.1 | 27.7%high effort | All 1 recorded result & sources27.7% · raw 27.7 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Terminus 2 agent; max 250 steps; reply cap 80k tokens or half the context; avg@3. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Tau2-Bench (Telecom) | 94.7%high effort | All 1 recorded result & sources94.7% · raw 94.7 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Domain score, avg@3; not the three-domain aggregate. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Tau2-Bench (Retail) | 69.9%high effort | All 1 recorded result & sources69.9% · raw 69.9 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Domain score, avg@3; not the three-domain aggregate. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Tau2-Bench (Airline) | 76.7%high effort | All 1 recorded result & sources76.7% · raw 76.7 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Domain score, avg@3; not the three-domain aggregate. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| τ³-Bench Banking | 38.1%high effort | All 1 recorded result & sources38.1% · raw 38.1 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Banking alltools retrieval configuration; avg@4; context overflow scores zero. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v3 (multi-turn) | 39.8%high effort | All 1 recorded result & sources39.8% · raw 39.8 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v4 (overall) | 61.4%high effort | All 1 recorded result & sources61.4% · raw 61.4 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v4 (non-live AST) | 79.1%high effort | All 1 recorded result & sources79.1% · raw 79.1 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v4 (live) | 78.9%high effort | All 1 recorded result & sources78.9% · raw 78.9 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v4 (multi-turn) | 47.5%high effort | All 1 recorded result & sources47.5% · raw 47.5 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v4 (memory) | 62.8%high effort | All 1 recorded result & sources62.8% · raw 62.8 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v4 (web search) | 62.5%high effort | All 1 recorded result & sources62.5% · raw 62.5 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BrowseComp | 29.4%high effort | All 1 recorded result & sources29.4% · raw 29.4 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Code average (EN) | 89.3%high effort | All 1 recorded result & sources89.3% · raw 89.3 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| LiveCodeBench v6 | 85.9%high effort | All 1 recorded result & sources85.9% · raw 85.9 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). August 2024-April 2025 problems; avg@3. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| HumanEval+ | 92.7%high effort | All 1 recorded result & sources92.7% · raw 92.7 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SWE-bench Verified | 66.4%high effort | All 1 recorded result & sources66.4% · raw 66.4 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). OpenCode agent; at most 250 steps; dedicated deployment. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Instruction Following average (EN) | 78.1%high effort | All 1 recorded result & sources78.1% · raw 78.1 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| IFBench (loose-prompt) | 78.1%high effort | All 1 recorded result & sources78.1% · raw 78.1 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Loose-prompt metric, avg@5; distinct from strict IFBench scoring. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Grounding / Hallucinations average (EN) | 59.4%high effort | All 1 recorded result & sources59.4% · raw 59.4 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SQuAD (M/A Grounding Score) | 23.4%high effort | All 1 recorded result & sources23.4% · raw 23.4 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SQuAD (Utility Accuracy) | 83.4%high effort | All 1 recorded result & sources83.4% · raw 83.4 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| RGB Closed-Book | 51%high effort | All 1 recorded result & sources51% · raw 51 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| RGB Negative (Abstention) | 85.6%high effort | All 1 recorded result & sources85.6% · raw 85.6 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AA-Omniscience Non-Hallucination Rate (1 − Hallucination Rate, public set) | 44%high effort | All 1 recorded result & sources44% · raw 44 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| RGB Fact-Check (Error Correction) | 34%high effort | All 1 recorded result & sources34% · raw 34 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| FRAMES (<24k) | 71.2%high effort | All 1 recorded result & sources71.2% · raw 71.2 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| FRAMES (>24k) | 73%high effort | All 1 recorded result & sources73% · raw 73 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SealQA (no distractors, <24k) | 80.7%high effort | All 1 recorded result & sources80.7% · raw 80.7 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SealQA (12 distractors, <24k) | 61%high effort | All 1 recorded result & sources61% · raw 61 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SealQA (no distractors, >24k) | 100%high effort | All 1 recorded result & sources100% · raw 100 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SealQA (12 distractors, >24k) | 65.1%high effort | All 1 recorded result & sources65.1% · raw 65.1 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Agentic Retrieval average (EN) | 77.3%high effort | All 1 recorded result & sources77.3% · raw 77.3 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Agentic Retrieval average (DE) | 69.4%high effort | All 1 recorded result & sources69.4% · raw 69.4 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| MuSiQue (EN) | 77.3%high effort | All 1 recorded result & sources77.3% · raw 77.3 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Honeypot | 80.8%high effort | All 1 recorded result & sources80.8% · raw 80.8 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Agentic Wiki QA (DE) | 69.4%high effort | All 1 recorded result & sources69.4% · raw 69.4 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Industry RAG average (EN) | 89.7%high effort | All 1 recorded result & sources89.7% · raw 89.7 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Industry RAG average (DE) | 67.5%high effort | All 1 recorded result & sources67.5% · raw 67.5 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Semiconductors | 80.4%high effort | All 1 recorded result & sources80.4% · raw 80.4 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| German Public Sector | 75%high effort | All 1 recorded result & sources75% · raw 75 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aerospace | 58.9%high effort | All 1 recorded result & sources58.9% · raw 58.9 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Automotive Supplier | 99%high effort | All 1 recorded result & sources99% · raw 99 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Industrial Drive Technology | 60%high effort | All 1 recorded result & sources60% · raw 60 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| LongBench Pro | 64.5%high effort | All 1 recorded result & sources64.5% · raw 64.5 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AA-LCR | 68.3%high effort | All 1 recorded result & sources68.3% · raw 68.3 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor-run AA-LCR, avg@3; not an independent AA measurement. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| RGB: invents nothing | 87.3%high effort | All 1 recorded result & sources87.3% · raw 87.3 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha grounding evaluation. Release radar visually inspected; exact value from its Show the numbers companion table. Not the RGB Negative (Abstention) metric. https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/ |
| AA-Omniscience Index (scaled public set) | 33.6%high effort | All 1 recorded result & sources33.6% · raw 33.6 % Headline · high effort · Own vendor Source/record date: 2026-10-03 Harness: Aleph Alpha public-set evaluation. Published product-page radar companion table gives 33.6% on the scaled display; the raw index is -32.8, preserved separately. Not an independent AA score. https://aleph-alpha.com/kolibri/ |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
No independent evaluator results are recorded for this model.
Read how we select and source scores or the comparison guide.