Apertus 1.5 adds image input, a thinking mode and 262K context

The Swiss ETH/EPFL model stays fully open — weights, data, Apache-2.0 — while gaining vision, audio preview and 16 distilled Apertus Mini models.

Nowline JUL 26 10:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Vision now, audio in preview

    Apertus 1.5 reads images alongside text natively and adds an experimental spoken-language input, plus a Thinking Mode you can toggle for step-by-step reasoning. For you: an open model you can self-host that captions, OCRs and reasons over screenshots and diagrams — with no closed vision API in the loop.

  • 262K context, 4× the last version

    The window grows to 262,144 tokens — four times Apertus 1.0 — after continued training of +4T tokens on the 8B and +2T on the 70B. That fits whole repos, long PDFs or multi-doc RAG in a single pass, on weights you fully control.

  • Open weights, open data, Apache-2.0

    Everything ships open — weights, the actual training data, and full training details — under Apache-2.0 for commercial use. Amid the fight over Chinese open weights, this is a Western, auditable, no-strings model you can put in a shipping product today.

  • Apertus Mini: 16 models for small boxes

    Alongside the 8B and 70B, the team shipped Apertus Mini — 16 compact checkpoints built via distillation and quantization for efficient inference. If the 70B is too heavy, there's now a right-sized open model for laptops, edge devices and cheap GPUs.

  • Get it: Hugging Face or the CSCS API

    Download Apertus v1.5-8B and v1.5-70B from Hugging Face now, or call CSCS's hosted endpoint if you'd rather not run the weights. The benchmark-heavy technical report lands "in the coming weeks," so test on your own evals before committing a pipeline.