Qwen3.8 27B Benchmarks, Specifications & Availability
Explore Qwen3.8 27B from Alibaba (Qwen): published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare Qwen3.8 27B with other models →Explore data coverage
Published specifications
- Provider
- Alibaba (Qwen)
- Access
- Open weights
- License
- Apache 2.0
- Context window
- 1M
- Total parameters
- 27B
- Active parameters
- Not published
- Released
- 2026-08-14
- Modalities
- text, image, video
- Family
- Qwen3.8
Model card · Announcement · Website · OpenRouter
Model notes
Dense native VLM; HF reports 27B params / ~28B safetensors size. Native ctx 262K, extensible to 1M via YaRN (QwenCloud hosted defaults to 1M). Ignore OR :free twin for main pricing. Open weights verified on Hugging Face https://huggingface.co/Qwen/Qwen3.8-27B (2026-09-25).
Pricing · OpenRouter
Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.
| Tier | Input / 1M | Output / 1M | Cached input / 1M | Cache write / 1M | Date & source |
|---|---|---|---|---|---|
| Default | [object Object] | [object Object] | [object Object] | — | 2026-10-03 · OpenRouter source |
Recorded pricing notes
OR prompt/completion/input_cache_read ×1e6; ignore :free twin. No -contribute sibling.; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Terminal-Bench 2.1 | 73%unknown effort | All 2 recorded results & sources73% · raw 73 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 Terminal Bench 2.1 (Terminus); HF card HTML table https://huggingface.co/Qwen/Qwen3.8-27B76.8% · raw 76.8 % Alternative · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Terminus 2 agent; max 250 steps; reply cap 80k tokens or half the context; avg@3. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SWE-bench Pro | 61.7%unknown effort | All 1 recorded result & sources61.7% · raw 61.7 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 SWE-bench Pro; Claude Code harness (temp=1.0, top_p=0.95, 256K) except Opus uses official https://huggingface.co/Qwen/Qwen3.8-27B |
| NL2Repo-Bench | 42.3%unknown effort | All 1 recorded result & sources42.3% · raw 42.3 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 NL2Repo-Bench; Claude Code harness https://huggingface.co/Qwen/Qwen3.8-27B |
| DeepSWE v1.1 | 42.2%unknown effort | All 1 recorded result & sources42.2% · raw 42.2 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 DeepSWE 1.1; Claude Code harness https://huggingface.co/Qwen/Qwen3.8-27B |
| QwenSWEBench V2 | 79%unknown effort | All 1 recorded result & sources79% · raw 79 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 Card lists QwenSWEBench (mapped to qwenswebench-v2); Claude Code avg@3 https://huggingface.co/Qwen/Qwen3.8-27B |
| CoWorkBench | 70.7%unknown effort | All 1 recorded result & sources70.7% · raw 70.7 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 CoWorkBench (in-house) https://huggingface.co/Qwen/Qwen3.8-27B |
| JobBench | 33.4%unknown effort | All 1 recorded result & sources33.4% · raw 33.4 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 JobBench https://huggingface.co/Qwen/Qwen3.8-27B |
| Agents' Last Exam | 20.4%unknown effort | All 1 recorded result & sources20.4% · raw 20.4 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 Agents' Last Exam Pass@1 (Score also reported 42.9 on card) https://huggingface.co/Qwen/Qwen3.8-27B |
| IFBench | 79.5%unknown effort | All 1 recorded result & sources79.5% · raw 79.5 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 IFBench https://huggingface.co/Qwen/Qwen3.8-27B |
| GPQA Diamond | 89.2%unknown effort | All 2 recorded results & sources89.2% · raw 89.2 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 GPQA Diamond https://huggingface.co/Qwen/Qwen3.8-27B89.2% · raw 89.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Chain-of-thought, avg@8. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Humanity's Last Exam | 30.8%unknown effort | All 1 recorded result & sources30.8% · raw 30.8 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 Humanity's Last Exam (text; no tools called out) https://huggingface.co/Qwen/Qwen3.8-27B |
| LiveCodeBench v6 | 90.3%unknown effort | All 2 recorded results & sources90.3% · raw 90.3 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 LiveCodeBench v6 https://huggingface.co/Qwen/Qwen3.8-27B93.8% · raw 93.8 % Alternative · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). August 2024-April 2025 problems; avg@3. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| OSWorld-Verified | 84.3%unknown effort | All 1 recorded result & sources84.3% · raw 84.3 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 OSWorld-Verified; from HF VL table https://huggingface.co/Qwen/Qwen3.8-27B |
| WebArena-Verified | 64.8%unknown effort | All 1 recorded result & sources64.8% · raw 64.8 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 WebArena-Verified under OSWorld scaffold https://huggingface.co/Qwen/Qwen3.8-27B |
| AndroidWorld | 81.9%unknown effort | All 1 recorded result & sources81.9% · raw 81.9 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 AndroidWorld https://huggingface.co/Qwen/Qwen3.8-27B |
| RecreationBench | 47.1%unknown effort | All 1 recorded result & sources47.1% · raw 47.1 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 RecreationBench (in-house application recreation) https://huggingface.co/Qwen/Qwen3.8-27B |
| ClawEval-MM | 57.4%unknown effort | All 1 recorded result & sources57.4% · raw 57.4 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 ClawEval-MM Pass@3 (Average also reported 56.9) https://huggingface.co/Qwen/Qwen3.8-27B |
| SWE-MM | 38.6%unknown effort | All 1 recorded result & sources38.6% · raw 38.6 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 SWE-MM; Claude Code harness on public SWE-bench Multimodal dev split https://huggingface.co/Qwen/Qwen3.8-27B |
| Vision2Web | 62.9%unknown effort | All 1 recorded result & sources62.9% · raw 62.9 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 Vision2Web avg across frontend/webpage/website; Claude Code harness https://huggingface.co/Qwen/Qwen3.8-27B |
| MathVision | 94.6%unknown effort | All 1 recorded result & sources94.6% · raw 94.6 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 MathVision With CI (Without CI 90.0 also on card) https://huggingface.co/Qwen/Qwen3.8-27B |
| BabyVision (w/ tools) | 85.6%unknown effort | All 1 recorded result & sources85.6% · raw 85.6 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 BabyVision With CI (Without CI 65.7 also on card) https://huggingface.co/Qwen/Qwen3.8-27B |
| CharXiv RQ | 90.2%unknown effort | All 1 recorded result & sources90.2% · raw 90.2 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 CharXiv (RQ) With CI (Without CI 83.7 also on card) https://huggingface.co/Qwen/Qwen3.8-27B |
| OmniDocBench 1.5 | 91.1%unknown effort | All 1 recorded result & sources91.1% · raw 91.1 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 OmniDocBench 1.5 https://huggingface.co/Qwen/Qwen3.8-27B |
| RealWorldQA | 85.9%unknown effort | All 1 recorded result & sources85.9% · raw 85.9 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 RealWorldQA https://huggingface.co/Qwen/Qwen3.8-27B |
| ERQA | 65.5%unknown effort | All 1 recorded result & sources65.5% · raw 65.5 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 ERQA https://huggingface.co/Qwen/Qwen3.8-27B |
| Aleph Alpha post-training Overall (EN) | 80.2%unknown effort | All 1 recorded result & sources80.2% · raw 80.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha post-training Overall (DE) | 79.9%unknown effort | All 1 recorded result & sources79.9% · raw 79.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Knowledge average (EN) | 56.8%unknown effort | All 1 recorded result & sources56.8% · raw 56.8 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Knowledge average (DE) | 69.2%unknown effort | All 1 recorded result & sources69.2% · raw 69.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| GPQA Diamond (DE) | 88.1%unknown effort | All 1 recorded result & sources88.1% · raw 88.1 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| HLE (text, no tools) | 35.6%unknown effort | All 1 recorded result & sources35.6% · raw 35.6 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Text-only questions, no output cap. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Humanity's Last Exam (DE) | 37.2%unknown effort | All 1 recorded result & sources37.2% · raw 37.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AA-Omniscience Accuracy (public set) | 19%unknown effort | All 2 recorded results & sources19% · raw 19 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. Source conflict: report gives 17.5; card gives 19.0. Card is the featured source; report retained as alternative. https://huggingface.co/Aleph-Alpha/Kolibri-117.5% · raw 17.5 % Alternative · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. Source conflict: report gives 17.5; card gives 19.0. Card is the featured source; report retained as alternative. https://aleph-alpha.com/downloads/tech-report.pdf#page=100 |
| AA-Omniscience Index (public set) | -5.8unknown effort | All 2 recorded results & sources-5.8 · raw -5.8 index Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. Source conflict: report gives -9.5; card gives -5.8. Card is the featured source; report retained as alternative. https://huggingface.co/Aleph-Alpha/Kolibri-1-9.5 · raw -9.5 index Alternative · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. Source conflict: report gives -9.5; card gives -5.8. Card is the featured source; report retained as alternative. https://aleph-alpha.com/downloads/tech-report.pdf#page=100 |
| MMLU-Pro CoT (EN) | 85%unknown effort | All 1 recorded result & sources85% · raw 85 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| MMLU-ProX CoT (DE) | 82.4%unknown effort | All 1 recorded result & sources82.4% · raw 82.4 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Math average (EN) | 97.8%unknown effort | All 1 recorded result & sources97.8% · raw 97.8 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Math average (DE) | 96.7%unknown effort | All 1 recorded result & sources96.7% · raw 96.7 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AIME 2025 (EN) | 97.9%unknown effort | All 1 recorded result & sources97.9% · raw 97.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). avg@16. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AIME 2025 (DE) | 96.5%unknown effort | All 1 recorded result & sources96.5% · raw 96.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. avg@16. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AIME 2026 (EN) | 97.7%unknown effort | All 1 recorded result & sources97.7% · raw 97.7 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). avg@16. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AIME 2026 (DE) | 96.9%unknown effort | All 1 recorded result & sources96.9% · raw 96.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. avg@16. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Agentic average (EN) | 66.7%unknown effort | All 1 recorded result & sources66.7% · raw 66.7 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Tau2-Bench (Telecom) | 82.5%unknown effort | All 1 recorded result & sources82.5% · raw 82.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Domain score, avg@3; not the three-domain aggregate. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Tau2-Bench (Retail) | 68.7%unknown effort | All 1 recorded result & sources68.7% · raw 68.7 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Domain score, avg@3; not the three-domain aggregate. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Tau2-Bench (Airline) | 83.3%unknown effort | All 1 recorded result & sources83.3% · raw 83.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Domain score, avg@3; not the three-domain aggregate. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| τ³-Bench Banking | 50%unknown effort | All 1 recorded result & sources50% · raw 50 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Banking alltools retrieval configuration; avg@4; context overflow scores zero. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v3 (multi-turn) | 42.5%unknown effort | All 1 recorded result & sources42.5% · raw 42.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v4 (overall) | 73.2%unknown effort | All 1 recorded result & sources73.2% · raw 73.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v4 (non-live AST) | 85.3%unknown effort | All 1 recorded result & sources85.3% · raw 85.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v4 (live) | 79.9%unknown effort | All 1 recorded result & sources79.9% · raw 79.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v4 (multi-turn) | 55.5%unknown effort | All 1 recorded result & sources55.5% · raw 55.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v4 (memory) | 79.6%unknown effort | All 1 recorded result & sources79.6% · raw 79.6 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BFCL v4 (web search) | 82%unknown effort | All 1 recorded result & sources82% · raw 82 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| BrowseComp | 46.4%unknown effort | All 1 recorded result & sources46.4% · raw 46.4 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Code average (EN) | 94.2%unknown effort | All 1 recorded result & sources94.2% · raw 94.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| HumanEval+ | 94.7%unknown effort | All 1 recorded result & sources94.7% · raw 94.7 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SWE-bench Verified | 72.6%unknown effort | All 1 recorded result & sources72.6% · raw 72.6 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). OpenCode agent; at most 250 steps; dedicated deployment. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Instruction Following average (EN) | 81.9%unknown effort | All 1 recorded result & sources81.9% · raw 81.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| IFBench (loose-prompt) | 81.9%unknown effort | All 1 recorded result & sources81.9% · raw 81.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Loose-prompt metric, avg@5; distinct from strict IFBench scoring. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Grounding / Hallucinations average (EN) | 67.5%unknown effort | All 1 recorded result & sources67.5% · raw 67.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SQuAD (M/A Grounding Score) | 8.3%unknown effort | All 1 recorded result & sources8.3% · raw 8.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SQuAD (Utility Accuracy) | 85.6%unknown effort | All 1 recorded result & sources85.6% · raw 85.6 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| RGB Closed-Book | 73%unknown effort | All 1 recorded result & sources73% · raw 73 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| RGB Negative (Abstention) | 70.6%unknown effort | All 1 recorded result & sources70.6% · raw 70.6 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AA-Omniscience Non-Hallucination Rate (1 − Hallucination Rate, public set) | 67.3%unknown effort | All 1 recorded result & sources67.3% · raw 67.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| RGB Fact-Check (Error Correction) | 53%unknown effort | All 1 recorded result & sources53% · raw 53 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| FRAMES (<24k) | 78.6%unknown effort | All 1 recorded result & sources78.6% · raw 78.6 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| FRAMES (>24k) | 83.3%unknown effort | All 1 recorded result & sources83.3% · raw 83.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SealQA (no distractors, <24k) | 86.9%unknown effort | All 1 recorded result & sources86.9% · raw 86.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SealQA (12 distractors, <24k) | 84%unknown effort | All 1 recorded result & sources84% · raw 84 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SealQA (no distractors, >24k) | 94.4%unknown effort | All 1 recorded result & sources94.4% · raw 94.4 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| SealQA (12 distractors, >24k) | 82.5%unknown effort | All 1 recorded result & sources82.5% · raw 82.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Agentic Retrieval average (EN) | 83.7%unknown effort | All 1 recorded result & sources83.7% · raw 83.7 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Agentic Retrieval average (DE) | 73.5%unknown effort | All 1 recorded result & sources73.5% · raw 73.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| MuSiQue (EN) | 83.7%unknown effort | All 1 recorded result & sources83.7% · raw 83.7 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Honeypot | 85%unknown effort | All 1 recorded result & sources85% · raw 85 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Agentic Wiki QA (DE) | 73.5%unknown effort | All 1 recorded result & sources73.5% · raw 73.5 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aleph Alpha Industry RAG average (EN) | 93.3%unknown effort | All 2 recorded results & sources93.3% · raw 93.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. Source conflict: report gives 93.2; card gives 93.3. Card is the featured source; report retained as alternative. https://huggingface.co/Aleph-Alpha/Kolibri-193.2% · raw 93.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. Source conflict: report gives 93.2; card gives 93.3. Card is the featured source; report retained as alternative. https://aleph-alpha.com/downloads/tech-report.pdf#page=101 |
| Aleph Alpha Industry RAG average (DE) | 80.2%unknown effort | All 1 recorded result & sources80.2% · raw 80.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Semiconductors | 89.2%unknown effort | All 1 recorded result & sources89.2% · raw 89.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| German Public Sector | 89%unknown effort | All 1 recorded result & sources89% · raw 89 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Aerospace | 73.8%unknown effort | All 1 recorded result & sources73.8% · raw 73.8 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Automotive Supplier | 97.3%unknown effort | All 1 recorded result & sources97.3% · raw 97.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| Industrial Drive Technology | 71.4%unknown effort | All 1 recorded result & sources71.4% · raw 71.4 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| LongBench Pro | 76.9%unknown effort | All 1 recorded result & sources76.9% · raw 76.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
| AA-LCR | 81.3%unknown effort | All 1 recorded result & sources81.3% · raw 81.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-10-03 Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor-run AA-LCR, avg@3; not an independent AA measurement. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. https://huggingface.co/Aleph-Alpha/Kolibri-1 |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Intelligence Index · Artificial Analysis | 34xhigh effort | All 4 recorded results & sources34 · raw 34 index Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/qwen3-8-27b28 · raw 28 index Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/qwen3-8-27b-medium26 · raw 26 index Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/qwen3-8-27b-low20 · raw 20 index Alternative · none effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/qwen3-8-27b-non-reasoning |
| Output speed · Artificial Analysis | 46 tok/sxhigh effort | All 4 recorded results & sources46 tok/s · raw 46 tok/s Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/qwen3-8-27b53 tok/s · raw 53 tok/s Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/qwen3-8-27b-medium52 tok/s · raw 52 tok/s Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/qwen3-8-27b-low51 tok/s · raw 51 tok/s Alternative · none effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/qwen3-8-27b-non-reasoning |
| Vals Index · Vals AI | 48.5%unknown effort | All 1 recorded result & sources48.5% · raw 48.48 % Headline · unknown effort · Independent evaluator Source/record date: 2026-09-23 Accuracy ±1.45; rank noted 32/59 on page extract https://www.vals.ai/models/alibaba_qwen3.8-27b |
| Bugs fixed /105 · Bug Hunt Bench | 15 fixesxhigh effort | All 1 recorded result & sources15 fixes · raw 15 fixes Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Claude Code / Alibaba API; effort xhigh; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench |
| Vibe Code Bench v1.1 · Vals AI | 64.9%unknown effort | All 1 recorded result & sources64.9% · raw 64.85 % Headline · unknown effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $23.46. https://www.vals.ai/benchmarks/vibe-code |
| ARC-AGI-1 · ARC Prize | 87.5%xhigh effort | All 4 recorded results & sources87.5% · raw 87.5 % Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties. https://arcprize.org/results/alibaba-qwen3-8-27b68.7% · raw 68.7 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/alibaba-qwen3-8-27b69.2% · raw 69.2 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/alibaba-qwen3-8-27b34% · raw 34 % Alternative · none effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/alibaba-qwen3-8-27b |
| ARC-AGI-2 · ARC Prize | 42.4%xhigh effort | All 4 recorded results & sources42.4% · raw 42.4 % Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties. https://arcprize.org/results/alibaba-qwen3-8-27b13.2% · raw 13.2 % Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/alibaba-qwen3-8-27b22.8% · raw 22.8 % Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/alibaba-qwen3-8-27b1.5% · raw 1.5 % Alternative · none effort · Independent evaluator Source/record date: 2026-10-03 Semi-Private verified matrix. Published verified configuration. https://arcprize.org/results/alibaba-qwen3-8-27b |
| Cost per Intelligence Index task · Artificial Analysis | $1.01xhigh effort | All 4 recorded results & sources$1.01 · raw 1.01 USD Headline · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/qwen3-8-27b$1.13 · raw 1.13 USD Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/qwen3-8-27b-medium$1.05 · raw 1.05 USD Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/qwen3-8-27b-low$2.49 · raw 2.49 USD Alternative · none effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/qwen3-8-27b-non-reasoning |
Read how we select and source scores or the comparison guide.