Apple's M5 Ultra Mac Studio brings 512GB unified memory for local AI
M5 Ultra hits 1.2TB/s bandwidth; a $899 M6 Mac mini runs agents on-device; cluster four for the biggest open weights. 512GB config slips to late October.

Copy markdown
512GB in a single desktop box
The M5 Ultra tops out at 512GB of unified memory at 1.2TB/s bandwidth — enough to hold a frontier-size open-weight model entirely in RAM on one machine, no multi-GPU rig to wire up. The Ultra starts at $5,499; the M5 Max caps at 128GB for $2,499.
The $899 M6 Mac mini for on-device agents
Apple's first 2nm chip pairs a dual 16-core Neural Engine (2x the AI compute of the prior gen) with up to 32GB unified memory, explicitly pitched for running LLMs and agentic tasks on-device. At $899 it's the cheapest serious local-inference box Apple ships.
Cluster four for frontier-size weights
Chain four Mac Studios and Apple claims up to 3x faster AI inference, enough to run the 'largest and most demanding' open-weight frontier models locally — a datacenter-free path for teams that can't send prompts to a cloud.
Pre-order now, mind the ship dates
Pre-orders open today; standard configs ship September 22. The headline 512GB Ultra doesn't arrive until late October, so budget your local-AI build around that gap rather than the launch date.
Elsewhere: DeepSeek's cheap open vision model
DeepSeek-V4-Flash-Vision (experimental) adds multimodal to the open V4-Flash at the same ~$0.44 in / $1.32 out per Mtok and a 1M-token context, with no separate image surcharge — a low-cost OSS way to drop vision into existing pipelines.