Aleph Alpha opens Kolibri-1: 78B MoE, 1M context, Apache 2.0
Just 3.46B params fire per token, so it serves on a single H200—bringing a reasoning dial, tool calling, and benchmarks that edge Qwen3.5 on agentic work.

Copy markdown
3.46B active, 78B total
A sparse MoE fires only 6 of 384 routed experts (plus 1 shared) per token, so you pay roughly 3.5B-model inference cost for 78B-class output. FP8 weights are ~78GB and serve on a single B200/B300/H200 or 2x H100 SXM5 via vLLM.
1M-token context, one-GPU serve
Native 262K extends to 1,048,576 tokens through a hybrid stack of 40 sliding-window and 10 full-attention layers—enough to feed it whole repos or document sets. Aleph Alpha suggests staying at or under 262K for the hardest tasks.
Apache 2.0, reasoning dial, tool calling
A fully commercial license, a per-request reasoning-effort knob (none/low/medium/high), and Hermes-style function schemas. It drops into an agent loop with no license friction and no API dependency to rate-limit you.
Benchmarks punch above the active count
84.3 on GPQA Diamond, 80.0 MMLU-Pro CoT, 66.4 SWE-Bench Verified, and 96.0 AIME 2026; 75.5 overall in English. It edges Qwen3.5 35B-A3B on agentic tasks (63.4 vs 62.1) while activating fewer parameters per token.
An EU-sovereign model that's actually competitive
Trained on 2T+ German tokens with a UniBPE tokenizer that is 11.2% more efficient on German text than GPT-5's. For teams needing EU data residency or strong German, it's the first general-purpose open model from the EU besides Mistral at this tier.