Lumen Labs open-sources a 9B model that runs on a phone at conversational speed
The benchmark table is less interesting than the quantization notes buried in section 4.
Lumen Labs released weights for Lumen-9, a 9-billion-parameter model it says holds 28 tokens/sec on a two-year-old flagship phone. License is permissive for companies under 50M users. Independent evals so far land it between last year's mid-size models on reasoning, weaker on long context.
Product SparkOffline-first assistants just became a weekend project
What product does this inspire or accelerate?
A phone-resident model at this speed makes private, no-signal assistants viable: field service, trail guides, clinical note-taking in basements. Expect a wave of "works in airplane mode" apps.
Investment ImpactPressure on per-token inference pricing
Where to jump in, or out? (Not financial advice.)
If good-enough runs locally, the low end of paid APIs gets squeezed. Watch mid-tier inference resellers; edge-chip designers are the likely beneficiaries.
Wade InTry it on your own phone tonight
How can you try this yourself this weekend?
The reference app is in the repo. Budget 5.1 GB of storage and turn off low-power mode.
-
After exploring Lumen-9 hands-on, you might want to look at how often Lumen actually ships model updates:
Lumen's release cadence, chartedEvery Lumen release since v1 and the gaps between them. · Lattice Weekly · Wade · 5 min18 -
How 4-bit quantization actually loses informationA visual walkthrough of what gets rounded away, and why some layers are spared. · Quiet Signal Review · Swim · 30 min1
-
The long history of "small models are enough"Every five years someone is right about this. Here's who, and when. · Fernbank Review of Computing · Plunge · an evening6