Perplexity and Nvidia launch a local-first agent, no per-token cost

Runs the full agent stack on a DGX Spark or 24GB RTX card; local file, code, and connector work is free, cloud steps billed only on approval.

Nowline AUG 26 5:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • A full agent that runs on your own hardware

    Perplexity and Nvidia's Portable Computer runs Perplexity's agent harness, orchestrator, and models entirely on-device — on an Nvidia DGX Spark (GB10 superchip, 128GB unified memory) or, per MarkTechPost, any Linux box with a 24GB+ RTX GPU. Anything the local models handle costs zero per token; only optional cloud escalation is billed.

  • What you get without touching the cloud

    It ships with Qwen 3.8 27B or PPLX 27B (a post-trained Qwen tuned for the harness), an OS-enforced sandbox that fences off filesystem paths and network, and local connectors for Gmail, Slack, and GitHub. That means local file and code search, sandboxed code execution, and scheduled workflows with nothing leaving the machine.

  • The numbers, and the fine print

    PPLX 27B scores 85.4% on Perplexity's 53-task local knowledge-work bench; Terminal-Bench 2.1 runs 59.6% fully local, climbing to 73.0% once it escalates to frontier models at ~$0.415 per rollout. It's Pro, Max, and Enterprise only — Linux today, Windows in September, and no macOS planned.