Gemini's agentic video hits the API — search hours, skip the frames
Set processing to 'agentic' and Gemini picks what to watch — 88% fewer tokens on long clips. Plus: Fable 5.1 cache reads fell 75%; labs gate offensive cyber.

Copy markdown
Set processing to 'agentic,' and the model decides
Instead of sampling video at a fixed frame rate, Gemini now searches across frames, audio, and transcript to decide what to watch and how fast. It's live in the Gemini API via the Interactions and GenerateContent endpoints on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite — you opt in by setting processing to 'agentic.'
The unlock: needle-in-a-haystack across hours of footage
You get sub-second moment retrieval, action and object counting over time, and anomaly detection across multi-hour videos — the long-form search that used to torch your token budget. Point it at a four-hour recording and have it jump straight to the exact clip you need.
88% fewer tokens, up to 66% cheaper, no surcharge
Google says agentic processing cuts analysis cost by up to 66% and token use by up to 88% on long content, while nudging accuracy up about 7%. It bills at standard token pricing with no extra feature fee, so the savings land straight on your existing video pipelines.
Elsewhere: Fable 5.1's cache reads dropped 75%
Claude Fable 5.1's cache-read price fell from $1.00 to $0.25 per million tokens — roughly 25% cheaper on typical workloads and up to 45% on agentic loops that reread context. Input and output stay at $10/$50 per million, so heavy-context agents are the ones that win.
Elsewhere: labs draw the AI-security line
Anthropic now lets Fable 5.1 flag software vulnerabilities defensively while routing pentesting and exploit generation to Opus, and its new Enterprise Frontier Safeguards pair zero data retention with misuse detection. Google's Gemini 3.8 Flash Cyber and OpenAI's Astra cyber tier stay gated behind trusted-access programs.