Examenos

Kolibri-1 Benchmarks, Specifications & Availability

Explore Kolibri-1 from Aleph Alpha: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare Kolibri-1 with other models →Explore data coverage

Published specifications

Provider
Aleph Alpha
Access
Open weights
License
Apache 2.0
Context window
1.05M
Total parameters
78B
Active parameters
3.46B
Released
2026-10-03
Modalities
text
Family
Kolibri

Model card · Announcement · Website

Model notes

Official ungated FP8 weights on Hugging Face; Apache 2.0. Card states 78,103,074,560 total and 3,457,573,120 active parameters. English/German reasoning and tool calling. Native context 262,144 tokens; validated extrapolation to 1,048,576 tokens; vendor recommends <=262,144 for serving efficiency and complex tasks. Benchmarks at high reasoning effort in Aleph Alpha eval-framework/Harbor. No OpenRouter listing on 2026-10-03. BF16 is a precision variant, not a separate model.

Pricing · OpenRouter

No OpenRouter pricing is recorded for this model. Missing rates are not free.

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Aleph Alpha post-training Overall (EN)75.5%high effort
All 1 recorded result & sources

75.5% · raw 75.5 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha post-training Overall (DE)70.8%high effort
All 1 recorded result & sources

70.8% · raw 70.8 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Knowledge average (EN)50.1%high effort
All 1 recorded result & sources

50.1% · raw 50.1 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Knowledge average (DE)57.6%high effort
All 1 recorded result & sources

57.6% · raw 57.6 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
GPQA Diamond84.3%high effort
All 1 recorded result & sources

84.3% · raw 84.3 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Chain-of-thought, avg@8.

https://huggingface.co/Aleph-Alpha/Kolibri-1
GPQA Diamond (DE)81.3%high effort
All 1 recorded result & sources

81.3% · raw 81.3 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks.

https://huggingface.co/Aleph-Alpha/Kolibri-1
HLE (text, no tools)21.5%high effort
All 1 recorded result & sources

21.5% · raw 21.5 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Text-only questions, no output cap.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Humanity's Last Exam (DE)15.9%high effort
All 1 recorded result & sources

15.9% · raw 15.9 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks.

https://huggingface.co/Aleph-Alpha/Kolibri-1
AA-Omniscience Accuracy (public set)14.8%high effort
All 1 recorded result & sources

14.8% · raw 14.8 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
AA-Omniscience Index (public set)-32.8high effort
All 1 recorded result & sources

-32.8 · raw -32.8 index

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
MMLU-Pro CoT (EN)80%high effort
All 1 recorded result & sources

80% · raw 80 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
MMLU-ProX CoT (DE)75.5%high effort
All 1 recorded result & sources

75.5% · raw 75.5 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Math average (EN)96.5%high effort
All 1 recorded result & sources

96.5% · raw 96.5 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Math average (DE)88.8%high effort
All 1 recorded result & sources

88.8% · raw 88.8 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
AIME 2025 (EN)96.9%high effort
All 1 recorded result & sources

96.9% · raw 96.9 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). avg@16.

https://huggingface.co/Aleph-Alpha/Kolibri-1
AIME 2025 (DE)87.5%high effort
All 1 recorded result & sources

87.5% · raw 87.5 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. avg@16.

https://huggingface.co/Aleph-Alpha/Kolibri-1
AIME 2026 (EN)96%high effort
All 1 recorded result & sources

96% · raw 96 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). avg@16.

https://huggingface.co/Aleph-Alpha/Kolibri-1
AIME 2026 (DE)90%high effort
All 1 recorded result & sources

90% · raw 90 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. avg@16.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Agentic average (EN)63.4%high effort
All 1 recorded result & sources

63.4% · raw 63.4 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Terminal-Bench 2.127.7%high effort
All 1 recorded result & sources

27.7% · raw 27.7 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Terminus 2 agent; max 250 steps; reply cap 80k tokens or half the context; avg@3.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Tau2-Bench (Telecom)94.7%high effort
All 1 recorded result & sources

94.7% · raw 94.7 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Domain score, avg@3; not the three-domain aggregate.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Tau2-Bench (Retail)69.9%high effort
All 1 recorded result & sources

69.9% · raw 69.9 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Domain score, avg@3; not the three-domain aggregate.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Tau2-Bench (Airline)76.7%high effort
All 1 recorded result & sources

76.7% · raw 76.7 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Domain score, avg@3; not the three-domain aggregate.

https://huggingface.co/Aleph-Alpha/Kolibri-1
τ³-Bench Banking38.1%high effort
All 1 recorded result & sources

38.1% · raw 38.1 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Banking alltools retrieval configuration; avg@4; context overflow scores zero.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v3 (multi-turn)39.8%high effort
All 1 recorded result & sources

39.8% · raw 39.8 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v4 (overall)61.4%high effort
All 1 recorded result & sources

61.4% · raw 61.4 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v4 (non-live AST)79.1%high effort
All 1 recorded result & sources

79.1% · raw 79.1 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v4 (live)78.9%high effort
All 1 recorded result & sources

78.9% · raw 78.9 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v4 (multi-turn)47.5%high effort
All 1 recorded result & sources

47.5% · raw 47.5 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v4 (memory)62.8%high effort
All 1 recorded result & sources

62.8% · raw 62.8 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v4 (web search)62.5%high effort
All 1 recorded result & sources

62.5% · raw 62.5 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BrowseComp29.4%high effort
All 1 recorded result & sources

29.4% · raw 29.4 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Code average (EN)89.3%high effort
All 1 recorded result & sources

89.3% · raw 89.3 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
LiveCodeBench v685.9%high effort
All 1 recorded result & sources

85.9% · raw 85.9 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). August 2024-April 2025 problems; avg@3.

https://huggingface.co/Aleph-Alpha/Kolibri-1
HumanEval+92.7%high effort
All 1 recorded result & sources

92.7% · raw 92.7 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
SWE-bench Verified66.4%high effort
All 1 recorded result & sources

66.4% · raw 66.4 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). OpenCode agent; at most 250 steps; dedicated deployment.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Instruction Following average (EN)78.1%high effort
All 1 recorded result & sources

78.1% · raw 78.1 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
IFBench (loose-prompt)78.1%high effort
All 1 recorded result & sources

78.1% · raw 78.1 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Loose-prompt metric, avg@5; distinct from strict IFBench scoring.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Grounding / Hallucinations average (EN)59.4%high effort
All 1 recorded result & sources

59.4% · raw 59.4 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
SQuAD (M/A Grounding Score)23.4%high effort
All 1 recorded result & sources

23.4% · raw 23.4 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
SQuAD (Utility Accuracy)83.4%high effort
All 1 recorded result & sources

83.4% · raw 83.4 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
RGB Closed-Book51%high effort
All 1 recorded result & sources

51% · raw 51 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
RGB Negative (Abstention)85.6%high effort
All 1 recorded result & sources

85.6% · raw 85.6 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
AA-Omniscience Non-Hallucination Rate (1 − Hallucination Rate, public set)44%high effort
All 1 recorded result & sources

44% · raw 44 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
RGB Fact-Check (Error Correction)34%high effort
All 1 recorded result & sources

34% · raw 34 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
FRAMES (<24k)71.2%high effort
All 1 recorded result & sources

71.2% · raw 71.2 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
FRAMES (>24k)73%high effort
All 1 recorded result & sources

73% · raw 73 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
SealQA (no distractors, <24k)80.7%high effort
All 1 recorded result & sources

80.7% · raw 80.7 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
SealQA (12 distractors, <24k)61%high effort
All 1 recorded result & sources

61% · raw 61 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
SealQA (no distractors, >24k)100%high effort
All 1 recorded result & sources

100% · raw 100 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
SealQA (12 distractors, >24k)65.1%high effort
All 1 recorded result & sources

65.1% · raw 65.1 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Agentic Retrieval average (EN)77.3%high effort
All 1 recorded result & sources

77.3% · raw 77.3 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Agentic Retrieval average (DE)69.4%high effort
All 1 recorded result & sources

69.4% · raw 69.4 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
MuSiQue (EN)77.3%high effort
All 1 recorded result & sources

77.3% · raw 77.3 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
Honeypot80.8%high effort
All 1 recorded result & sources

80.8% · raw 80.8 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
Agentic Wiki QA (DE)69.4%high effort
All 1 recorded result & sources

69.4% · raw 69.4 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Industry RAG average (EN)89.7%high effort
All 1 recorded result & sources

89.7% · raw 89.7 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Industry RAG average (DE)67.5%high effort
All 1 recorded result & sources

67.5% · raw 67.5 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Semiconductors80.4%high effort
All 1 recorded result & sources

80.4% · raw 80.4 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
German Public Sector75%high effort
All 1 recorded result & sources

75% · raw 75 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aerospace58.9%high effort
All 1 recorded result & sources

58.9% · raw 58.9 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
Automotive Supplier99%high effort
All 1 recorded result & sources

99% · raw 99 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
Industrial Drive Technology60%high effort
All 1 recorded result & sources

60% · raw 60 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
LongBench Pro64.5%high effort
All 1 recorded result & sources

64.5% · raw 64.5 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101).

https://huggingface.co/Aleph-Alpha/Kolibri-1
AA-LCR68.3%high effort
All 1 recorded result & sources

68.3% · raw 68.3 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor-run AA-LCR, avg@3; not an independent AA measurement.

https://huggingface.co/Aleph-Alpha/Kolibri-1
RGB: invents nothing87.3%high effort
All 1 recorded result & sources

87.3% · raw 87.3 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha grounding evaluation. Release radar visually inspected; exact value from its Show the numbers companion table. Not the RGB Negative (Abstention) metric.

https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/
AA-Omniscience Index (scaled public set)33.6%high effort
All 1 recorded result & sources

33.6% · raw 33.6 %

Headline · high effort · Own vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha public-set evaluation. Published product-page radar companion table gives 33.6% on the scaled display; the raw index is -32.8, preserved separately. Not an independent AA score.

https://aleph-alpha.com/kolibri/

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

No independent evaluator results are recorded for this model.

Read how we select and source scores or the comparison guide.