Content Quality: A 691-word News article (range 400-1200) with Overview / What We Know / What We Don't Know / Analysis structure. Specifications are attributed per bullet to the model card or the technical report; all benchmark figures are explicitly framed as Aleph Alpha's own evaluations (Overview: 'All benchmark figures below come from Aleph Alpha's own evaluations. No independent replication had been located'; Reported results: 'The technical report says that among the open models Aleph Alpha evaluated...'; the Analysis closes with 'depends on third-party testing'). The article never states 'outperforms X' as independent fact. Tone is neutral. No AI self-reference.
Source Verification: Read both snapshots from disk. manifest.json: both status 200, no archive fallback. source-0.html.gz is the Hugging Face model card page (HTML, 525,968 bytes uncompressed). source-1.html.gz is a genuine PDF: gunzip yields a stream beginning '%PDF-1.5', 3,465,511 bytes, matching the manifest content_type application/pdf and content_length; pdftotext extracted 11,468 lines (189 pages; title 'Kolibri: A Sovereign European Model on the Pareto Frontier'). So the writer's local pdftotext extraction is corroborated by the archived snapshot itself; no figure is unverifiable. MODEL CARD (source-0) confirms: 'Architecture Mixture-of-Experts', 'Total parameters 78B ( 78,103,074,560 )', 'Active parameters / token 3.46B ( 3,457,573,120 )', 'Languages German, English', 'Context length 1,048,576 tokens ; we recommend <=262,144 tokens for serving efficiency and complex tasks', 'License Apache 2.0', 'Release Date 3rd of October 2026', 'Knowledge cutoff EN: June 18, 2026, DE: June 18, 2026', 'Model memory footprint: ~78 GB (FP8 weights). Minimum: 2x A100 80 GB, 2x H100 SXM5, 1x H200, 1x B200 or 1x B300', reasoning_effort values low/medium/high and 'You can disable thinking altogether by setting reasoning_effort to none', pre-training 'Trained on 20T tokens ... (~62.5% English, ~23.9% German, ~13.6% code)', '3.44T in mid-training and 201B for long-context extension', '768 NVIDIA B200 ... Time: 21 days (511h, 392k GPUh)', 'FLOPS: 6.4e23', 'Energy consumption 9.5x10^2 MWh (estimated)', 'Aleph Alpha is a signatory of the EU GPAI Code of Practice', 'Supporting two languages rather than many is a deliberate choice of depth over breadth', and 'The trade-off is memory: the full model must be held in memory even though only part ... active at any time'. The title's specs (78B, Apache 2.0, German-English, MoE, 3.46B active) all match the card. TECH REPORT (source-1) confirms: 'English-German Mixture-of-Experts transformer with 78.1B total parameters and 3.46B active parameters per token, released as open weights under the Apache 2.0 licence'; 'of which 3.46B, or 4.4 %, are active per token'; 'Training spans 24T tokens across pre-training, mid-training, and long-context extension'; 'Teams in Germany developed Kolibri end to end and train it on infrastructure in Germany and Finland'; 'RL on more than 1.2M internally curated tasks, which cover reasoning, tool use, instruction following, code, and retrieval'; 'among the open models evaluated in this report, it lies on the Pareto frontier of quality and serving cost in English and German' (Figure 1 covers base and post-trained models); post-training table row 'Overall (EN) 75.5 ... Overall (DE) 70.8'; AIME 2026 96.0, GPQA Diamond 84.3; 'SWE-Bench Verified 66.4'; SFT-to-RL table 'SFT 92.7 92.1 79.8 ... RL 96.9 96.0 84.3' (AIME 2026 92.1->96.0, GPQA Diamond 79.8->84.3); 'AA-Omniscience Non-Hallucination Rate (public set) 44.0 15.0' with Kolibri and Kolibri Origin as the first two columns; 'train the model to abstain when the provided context does not support an answer'; 'Organisations in public administration, industry, and aerospace process sensitive data under regulation'. Release-date conflict: the 10 March 2026 Aleph Alpha blog post is not cited in article.sources or the body, and no '10 March'/March 2026 date appears in either snapshot; the 3 October date comes from the cited model card's 'Release Date 3rd of October 2026' field. SUSPICIOUS_PATTERNS (Step 3d): source-0 matched 'system-prompt-reference' with the excerpt 'e>medium</code> or <code>high</code> reasoning effort, which the chat template requests through a fixed sentence in the system prompt. In SFT, we assigned reasoning effort labels based on observed dataset-specific reasoning length statistics and judged'. Read in context in the decompressed snapshot: it is the model card's own description of how the chat template maps reasoning effort to a sentence in the system prompt during training. It addresses no AI agent reader, issues no instruction and asks for nothing to be withheld. False positive; no instruction acted on. source-1 suspicious_patterns is null.
Factual Accuracy: Every specific in the headline, summary, lead and body traces to one of the two snapshots and is attributed to the correct one (model card for hardware, licence, languages, release date, 20T/3.44T/201B token split, GPU-hours, FLOPS, energy; tech report for 78.1B/3.46B/4.4%, 24T, Pareto frontier, benchmarks, 1.2M RL tasks, abstention). The 20T (model card, pre-training only) versus 24T (tech report, all stages) figures are presented separately and consistently (20T + 3.44T + 201B is about 23.6T, within the report's rounded 24T). The summary's claim that Aleph Alpha published Kolibri on Hugging Face on October 3 rests on the card's Release Date field. No hallucinated quotes (the body has no direct quotes).
Overall Assessment: Clean, well-attributed, fully verified. APPROVE.