Aleph Alpha opens Kolibri-1: 78B MoE, 1M context, Apache 2.0

Just 3.46B params fire per token, so it serves on a single H200—bringing a reasoning dial, tool calling, and benchmarks that edge Qwen3.5 on agentic work.

Nowline OCT 6 1:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 3.46B active, 78B total

    A sparse MoE fires only 6 of 384 routed experts (plus 1 shared) per token, so you pay roughly 3.5B-model inference cost for 78B-class output. FP8 weights are ~78GB and serve on a single B200/B300/H200 or 2x H100 SXM5 via vLLM.

  • 1M-token context, one-GPU serve

    Native 262K extends to 1,048,576 tokens through a hybrid stack of 40 sliding-window and 10 full-attention layers—enough to feed it whole repos or document sets. Aleph Alpha suggests staying at or under 262K for the hardest tasks.

  • Apache 2.0, reasoning dial, tool calling

    A fully commercial license, a per-request reasoning-effort knob (none/low/medium/high), and Hermes-style function schemas. It drops into an agent loop with no license friction and no API dependency to rate-limit you.

  • Benchmarks punch above the active count

    84.3 on GPQA Diamond, 80.0 MMLU-Pro CoT, 66.4 SWE-Bench Verified, and 96.0 AIME 2026; 75.5 overall in English. It edges Qwen3.5 35B-A3B on agentic tasks (63.4 vs 62.1) while activating fewer parameters per token.

  • An EU-sovereign model that's actually competitive

    Trained on 2T+ German tokens with a UniBPE tokenizer that is 11.2% more efficient on German text than GPT-5's. For teams needing EU data residency or strong German, it's the first general-purpose open model from the EU besides Mistral at this tier.