Reka Edge 2603: open 7B VLM encodes a 1024px image in 331 tokens
Open weights on Hugging Face, runs on a 24GB Mac, and a 4-bit build drops to phones — at roughly a third the image tokens of Qwen 3.5 or Gemini 3 Pro.

Copy markdown
A third the image tokens of the big VLMs
Reka Edge 2603 encodes a 1024x1024 image in 331 input tokens, versus 1,041 for Qwen 3.5 9B, 1,063 for Cosmos-Reason2 8B and 1,094 for Gemini 3 Pro. For OCR, media tagging or screen-watching agents, that's roughly 3x less you pay or compute per frame.
Open weights, but check the license
The 7B (BF16) weights are live on Hugging Face under a custom reka-edge-2603 license: free commercial use only for organizations under $1M in annual revenue; larger shops need a separate deal. Fine for indie and side projects, a gotcha for funded startups.
Runs on a Mac, quantizes to a phone
At BF16 it fits a 24GB Apple Silicon Mac or a Jetson AGX Orin/Thor. A 4-bit build drops memory from 13GB to 5GB while keeping over 98% of quality — small enough for a Jetson Orin Nano, a Galaxy S25, even an iPhone. Local, offline vision with no per-token bill.
Fast enough for real-time vision
Reka reports 4.69s end-to-end and 5.46 images/sec against 16.67s for Gemini 3 Pro, with about 0.5s to first token. Benchmarks hold up: VQA-V2 88.40, RefCOCO-A 93.13 for grounding, and 74.30 on MLVU for video — enough for live camera and agent loops.
Prototype it on OpenRouter first
Beyond the weights, reka-edge-2603 is listed on OpenRouter behind the standard OpenAI SDK, so you can point a few requests at the hosted endpoint and test your vision pipeline before committing to self-hosting.