Examenos

Qwen3.8 27B Benchmarks, Specifications & Availability

Explore Qwen3.8 27B from Alibaba (Qwen): published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare Qwen3.8 27B with other models →Explore data coverage

Published specifications

Provider
Alibaba (Qwen)
Access
Open weights
License
Apache 2.0
Context window
1M
Total parameters
27B
Active parameters
Not published
Released
2026-08-14
Modalities
text, image, video
Family
Qwen3.8

Model card · Announcement · Website · OpenRouter

Model notes

Dense native VLM; HF reports 27B params / ~28B safetensors size. Native ctx 262K, extensible to 1M via YaRN (QwenCloud hosted defaults to 1M). Ignore OR :free twin for main pricing. Open weights verified on Hugging Face https://huggingface.co/Qwen/Qwen3.8-27B (2026-09-25).

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Default[object Object][object Object][object Object]—2026-10-03 · OpenRouter source
Recorded pricing notes

OR prompt/completion/input_cache_read ×1e6; ignore :free twin. No -contribute sibling.; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Terminal-Bench 2.173%unknown effort
All 2 recorded results & sources

73% · raw 73 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

Terminal Bench 2.1 (Terminus); HF card HTML table

https://huggingface.co/Qwen/Qwen3.8-27B

76.8% · raw 76.8 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Terminus 2 agent; max 250 steps; reply cap 80k tokens or half the context; avg@3. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
SWE-bench Pro61.7%unknown effort
All 1 recorded result & sources

61.7% · raw 61.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

SWE-bench Pro; Claude Code harness (temp=1.0, top_p=0.95, 256K) except Opus uses official

https://huggingface.co/Qwen/Qwen3.8-27B
NL2Repo-Bench42.3%unknown effort
All 1 recorded result & sources

42.3% · raw 42.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

NL2Repo-Bench; Claude Code harness

https://huggingface.co/Qwen/Qwen3.8-27B
DeepSWE v1.142.2%unknown effort
All 1 recorded result & sources

42.2% · raw 42.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

DeepSWE 1.1; Claude Code harness

https://huggingface.co/Qwen/Qwen3.8-27B
QwenSWEBench V279%unknown effort
All 1 recorded result & sources

79% · raw 79 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

Card lists QwenSWEBench (mapped to qwenswebench-v2); Claude Code avg@3

https://huggingface.co/Qwen/Qwen3.8-27B
CoWorkBench70.7%unknown effort
All 1 recorded result & sources

70.7% · raw 70.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

CoWorkBench (in-house)

https://huggingface.co/Qwen/Qwen3.8-27B
JobBench33.4%unknown effort
All 1 recorded result & sources

33.4% · raw 33.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

JobBench

https://huggingface.co/Qwen/Qwen3.8-27B
Agents' Last Exam20.4%unknown effort
All 1 recorded result & sources

20.4% · raw 20.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

Agents' Last Exam Pass@1 (Score also reported 42.9 on card)

https://huggingface.co/Qwen/Qwen3.8-27B
IFBench79.5%unknown effort
All 1 recorded result & sources

79.5% · raw 79.5 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

IFBench

https://huggingface.co/Qwen/Qwen3.8-27B
GPQA Diamond89.2%unknown effort
All 2 recorded results & sources

89.2% · raw 89.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

GPQA Diamond

https://huggingface.co/Qwen/Qwen3.8-27B

89.2% · raw 89.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Chain-of-thought, avg@8. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Humanity's Last Exam30.8%unknown effort
All 1 recorded result & sources

30.8% · raw 30.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

Humanity's Last Exam (text; no tools called out)

https://huggingface.co/Qwen/Qwen3.8-27B
LiveCodeBench v690.3%unknown effort
All 2 recorded results & sources

90.3% · raw 90.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

LiveCodeBench v6

https://huggingface.co/Qwen/Qwen3.8-27B

93.8% · raw 93.8 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). August 2024-April 2025 problems; avg@3. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
OSWorld-Verified84.3%unknown effort
All 1 recorded result & sources

84.3% · raw 84.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

OSWorld-Verified; from HF VL table

https://huggingface.co/Qwen/Qwen3.8-27B
WebArena-Verified64.8%unknown effort
All 1 recorded result & sources

64.8% · raw 64.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

WebArena-Verified under OSWorld scaffold

https://huggingface.co/Qwen/Qwen3.8-27B
AndroidWorld81.9%unknown effort
All 1 recorded result & sources

81.9% · raw 81.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

AndroidWorld

https://huggingface.co/Qwen/Qwen3.8-27B
RecreationBench47.1%unknown effort
All 1 recorded result & sources

47.1% · raw 47.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

RecreationBench (in-house application recreation)

https://huggingface.co/Qwen/Qwen3.8-27B
ClawEval-MM57.4%unknown effort
All 1 recorded result & sources

57.4% · raw 57.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

ClawEval-MM Pass@3 (Average also reported 56.9)

https://huggingface.co/Qwen/Qwen3.8-27B
SWE-MM38.6%unknown effort
All 1 recorded result & sources

38.6% · raw 38.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

SWE-MM; Claude Code harness on public SWE-bench Multimodal dev split

https://huggingface.co/Qwen/Qwen3.8-27B
Vision2Web62.9%unknown effort
All 1 recorded result & sources

62.9% · raw 62.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

Vision2Web avg across frontend/webpage/website; Claude Code harness

https://huggingface.co/Qwen/Qwen3.8-27B
MathVision94.6%unknown effort
All 1 recorded result & sources

94.6% · raw 94.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

MathVision With CI (Without CI 90.0 also on card)

https://huggingface.co/Qwen/Qwen3.8-27B
BabyVision (w/ tools)85.6%unknown effort
All 1 recorded result & sources

85.6% · raw 85.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

BabyVision With CI (Without CI 65.7 also on card)

https://huggingface.co/Qwen/Qwen3.8-27B
CharXiv RQ90.2%unknown effort
All 1 recorded result & sources

90.2% · raw 90.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

CharXiv (RQ) With CI (Without CI 83.7 also on card)

https://huggingface.co/Qwen/Qwen3.8-27B
OmniDocBench 1.591.1%unknown effort
All 1 recorded result & sources

91.1% · raw 91.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

OmniDocBench 1.5

https://huggingface.co/Qwen/Qwen3.8-27B
RealWorldQA85.9%unknown effort
All 1 recorded result & sources

85.9% · raw 85.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

RealWorldQA

https://huggingface.co/Qwen/Qwen3.8-27B
ERQA65.5%unknown effort
All 1 recorded result & sources

65.5% · raw 65.5 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

ERQA

https://huggingface.co/Qwen/Qwen3.8-27B
Aleph Alpha post-training Overall (EN)80.2%unknown effort
All 1 recorded result & sources

80.2% · raw 80.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha post-training Overall (DE)79.9%unknown effort
All 1 recorded result & sources

79.9% · raw 79.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Knowledge average (EN)56.8%unknown effort
All 1 recorded result & sources

56.8% · raw 56.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Knowledge average (DE)69.2%unknown effort
All 1 recorded result & sources

69.2% · raw 69.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
GPQA Diamond (DE)88.1%unknown effort
All 1 recorded result & sources

88.1% · raw 88.1 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
HLE (text, no tools)35.6%unknown effort
All 1 recorded result & sources

35.6% · raw 35.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Text-only questions, no output cap. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Humanity's Last Exam (DE)37.2%unknown effort
All 1 recorded result & sources

37.2% · raw 37.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
AA-Omniscience Accuracy (public set)19%unknown effort
All 2 recorded results & sources

19% · raw 19 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. Source conflict: report gives 17.5; card gives 19.0. Card is the featured source; report retained as alternative.

https://huggingface.co/Aleph-Alpha/Kolibri-1

17.5% · raw 17.5 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. Source conflict: report gives 17.5; card gives 19.0. Card is the featured source; report retained as alternative.

https://aleph-alpha.com/downloads/tech-report.pdf#page=100
AA-Omniscience Index (public set)-5.8unknown effort
All 2 recorded results & sources

-5.8 · raw -5.8 index

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. Source conflict: report gives -9.5; card gives -5.8. Card is the featured source; report retained as alternative.

https://huggingface.co/Aleph-Alpha/Kolibri-1

-9.5 · raw -9.5 index

Alternative · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. Source conflict: report gives -9.5; card gives -5.8. Card is the featured source; report retained as alternative.

https://aleph-alpha.com/downloads/tech-report.pdf#page=100
MMLU-Pro CoT (EN)85%unknown effort
All 1 recorded result & sources

85% · raw 85 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
MMLU-ProX CoT (DE)82.4%unknown effort
All 1 recorded result & sources

82.4% · raw 82.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Math average (EN)97.8%unknown effort
All 1 recorded result & sources

97.8% · raw 97.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Math average (DE)96.7%unknown effort
All 1 recorded result & sources

96.7% · raw 96.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
AIME 2025 (EN)97.9%unknown effort
All 1 recorded result & sources

97.9% · raw 97.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). avg@16. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
AIME 2025 (DE)96.5%unknown effort
All 1 recorded result & sources

96.5% · raw 96.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. avg@16. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
AIME 2026 (EN)97.7%unknown effort
All 1 recorded result & sources

97.7% · raw 97.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). avg@16. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
AIME 2026 (DE)96.9%unknown effort
All 1 recorded result & sources

96.9% · raw 96.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. avg@16. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Agentic average (EN)66.7%unknown effort
All 1 recorded result & sources

66.7% · raw 66.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Tau2-Bench (Telecom)82.5%unknown effort
All 1 recorded result & sources

82.5% · raw 82.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Domain score, avg@3; not the three-domain aggregate. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Tau2-Bench (Retail)68.7%unknown effort
All 1 recorded result & sources

68.7% · raw 68.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Domain score, avg@3; not the three-domain aggregate. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Tau2-Bench (Airline)83.3%unknown effort
All 1 recorded result & sources

83.3% · raw 83.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Domain score, avg@3; not the three-domain aggregate. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
τ³-Bench Banking50%unknown effort
All 1 recorded result & sources

50% · raw 50 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Banking alltools retrieval configuration; avg@4; context overflow scores zero. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v3 (multi-turn)42.5%unknown effort
All 1 recorded result & sources

42.5% · raw 42.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v4 (overall)73.2%unknown effort
All 1 recorded result & sources

73.2% · raw 73.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v4 (non-live AST)85.3%unknown effort
All 1 recorded result & sources

85.3% · raw 85.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v4 (live)79.9%unknown effort
All 1 recorded result & sources

79.9% · raw 79.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v4 (multi-turn)55.5%unknown effort
All 1 recorded result & sources

55.5% · raw 55.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v4 (memory)79.6%unknown effort
All 1 recorded result & sources

79.6% · raw 79.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BFCL v4 (web search)82%unknown effort
All 1 recorded result & sources

82% · raw 82 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). bfcl-eval; different web-search backend and per-model sampling; not comparable to BFCL leaderboard. Overall includes irrelevance detection not separately printed. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
BrowseComp46.4%unknown effort
All 1 recorded result & sources

46.4% · raw 46.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Code average (EN)94.2%unknown effort
All 1 recorded result & sources

94.2% · raw 94.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
HumanEval+94.7%unknown effort
All 1 recorded result & sources

94.7% · raw 94.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
SWE-bench Verified72.6%unknown effort
All 1 recorded result & sources

72.6% · raw 72.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). OpenCode agent; at most 250 steps; dedicated deployment. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Instruction Following average (EN)81.9%unknown effort
All 1 recorded result & sources

81.9% · raw 81.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
IFBench (loose-prompt)81.9%unknown effort
All 1 recorded result & sources

81.9% · raw 81.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Loose-prompt metric, avg@5; distinct from strict IFBench scoring. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Grounding / Hallucinations average (EN)67.5%unknown effort
All 1 recorded result & sources

67.5% · raw 67.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
SQuAD (M/A Grounding Score)8.3%unknown effort
All 1 recorded result & sources

8.3% · raw 8.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
SQuAD (Utility Accuracy)85.6%unknown effort
All 1 recorded result & sources

85.6% · raw 85.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
RGB Closed-Book73%unknown effort
All 1 recorded result & sources

73% · raw 73 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
RGB Negative (Abstention)70.6%unknown effort
All 1 recorded result & sources

70.6% · raw 70.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
AA-Omniscience Non-Hallucination Rate (1 − Hallucination Rate, public set)67.3%unknown effort
All 1 recorded result & sources

67.3% · raw 67.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
RGB Fact-Check (Error Correction)53%unknown effort
All 1 recorded result & sources

53% · raw 53 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
FRAMES (<24k)78.6%unknown effort
All 1 recorded result & sources

78.6% · raw 78.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
FRAMES (>24k)83.3%unknown effort
All 1 recorded result & sources

83.3% · raw 83.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
SealQA (no distractors, <24k)86.9%unknown effort
All 1 recorded result & sources

86.9% · raw 86.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
SealQA (12 distractors, <24k)84%unknown effort
All 1 recorded result & sources

84% · raw 84 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
SealQA (no distractors, >24k)94.4%unknown effort
All 1 recorded result & sources

94.4% · raw 94.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
SealQA (12 distractors, >24k)82.5%unknown effort
All 1 recorded result & sources

82.5% · raw 82.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Agentic Retrieval average (EN)83.7%unknown effort
All 1 recorded result & sources

83.7% · raw 83.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Agentic Retrieval average (DE)73.5%unknown effort
All 1 recorded result & sources

73.5% · raw 73.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
MuSiQue (EN)83.7%unknown effort
All 1 recorded result & sources

83.7% · raw 83.7 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Honeypot85%unknown effort
All 1 recorded result & sources

85% · raw 85 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Agentic Wiki QA (DE)73.5%unknown effort
All 1 recorded result & sources

73.5% · raw 73.5 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aleph Alpha Industry RAG average (EN)93.3%unknown effort
All 2 recorded results & sources

93.3% · raw 93.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. Source conflict: report gives 93.2; card gives 93.3. Card is the featured source; report retained as alternative.

https://huggingface.co/Aleph-Alpha/Kolibri-1

93.2% · raw 93.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present. Source conflict: report gives 93.2; card gives 93.3. Card is the featured source; report retained as alternative.

https://aleph-alpha.com/downloads/tech-report.pdf#page=101
Aleph Alpha Industry RAG average (DE)80.2%unknown effort
All 1 recorded result & sources

80.2% · raw 80.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). German translation (ellamind), except MMLU-ProX and vendor retrieval/proxy tasks. Vendor aggregate excludes incomplete/non-percentage rows; Overall EN uses eight categories, DE four. Long Context excluded. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Semiconductors89.2%unknown effort
All 1 recorded result & sources

89.2% · raw 89.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
German Public Sector89%unknown effort
All 1 recorded result & sources

89% · raw 89 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Aerospace73.8%unknown effort
All 1 recorded result & sources

73.8% · raw 73.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Automotive Supplier97.3%unknown effort
All 1 recorded result & sources

97.3% · raw 97.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
Industrial Drive Technology71.4%unknown effort
All 1 recorded result & sources

71.4% · raw 71.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
LongBench Pro76.9%unknown effort
All 1 recorded result & sources

76.9% · raw 76.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1
AA-LCR81.3%unknown effort
All 1 recorded result & sources

81.3% · raw 81.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-10-03

Harness: Aleph Alpha eval-framework/Harbor, released post-trained checkpoint. Kolibri reasoning effort high; peers use vendor-documented sampling. Technical report Tables 27-29 (pp. 96-101). Vendor-run AA-LCR, avg@3; not an independent AA measurement. As reported by Aleph Alpha for Qwen3.8 27B; preserves existing own-vendor headline when present.

https://huggingface.co/Aleph-Alpha/Kolibri-1

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis34xhigh effort
All 4 recorded results & sources

34 · raw 34 index

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/qwen3-8-27b

28 · raw 28 index

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/qwen3-8-27b-medium

26 · raw 26 index

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/qwen3-8-27b-low

20 · raw 20 index

Alternative · none effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/qwen3-8-27b-non-reasoning
Output speed · Artificial Analysis46 tok/sxhigh effort
All 4 recorded results & sources

46 tok/s · raw 46 tok/s

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/qwen3-8-27b

53 tok/s · raw 53 tok/s

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/qwen3-8-27b-medium

52 tok/s · raw 52 tok/s

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/qwen3-8-27b-low

51 tok/s · raw 51 tok/s

Alternative · none effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/qwen3-8-27b-non-reasoning
Vals Index · Vals AI48.5%unknown effort
All 1 recorded result & sources

48.5% · raw 48.48 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

Accuracy ±1.45; rank noted 32/59 on page extract

https://www.vals.ai/models/alibaba_qwen3.8-27b
Bugs fixed /105 · Bug Hunt Bench15 fixesxhigh effort
All 1 recorded result & sources

15 fixes · raw 15 fixes

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Claude Code / Alibaba API; effort xhigh; 3 runs; evaluation 2026-09-15. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench
Vibe Code Bench v1.1 · Vals AI64.9%unknown effort
All 1 recorded result & sources

64.9% · raw 64.85 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $23.46.

https://www.vals.ai/benchmarks/vibe-code
ARC-AGI-1 · ARC Prize87.5%xhigh effort
All 4 recorded results & sources

87.5% · raw 87.5 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/alibaba-qwen3-8-27b

68.7% · raw 68.7 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/alibaba-qwen3-8-27b

69.2% · raw 69.2 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/alibaba-qwen3-8-27b

34% · raw 34 %

Alternative · none effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/alibaba-qwen3-8-27b
ARC-AGI-2 · ARC Prize42.4%xhigh effort
All 4 recorded results & sources

42.4% · raw 42.4 %

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration. Headline: best verified value; source matrix effort order breaks ties.

https://arcprize.org/results/alibaba-qwen3-8-27b

13.2% · raw 13.2 %

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/alibaba-qwen3-8-27b

22.8% · raw 22.8 %

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/alibaba-qwen3-8-27b

1.5% · raw 1.5 %

Alternative · none effort · Independent evaluator

Source/record date: 2026-10-03

Semi-Private verified matrix. Published verified configuration.

https://arcprize.org/results/alibaba-qwen3-8-27b
Cost per Intelligence Index task · Artificial Analysis$1.01xhigh effort
All 4 recorded results & sources

$1.01 · raw 1.01 USD

Headline · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/qwen3-8-27b

$1.13 · raw 1.13 USD

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/qwen3-8-27b-medium

$1.05 · raw 1.05 USD

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/qwen3-8-27b-low

$2.49 · raw 2.49 USD

Alternative · none effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/qwen3-8-27b-non-reasoning

Read how we select and source scores or the comparison guide.