A 284B-parameter model now runs on a 128 GB Mac. Who loses?
ds4, an open-source C engine, squeezes DeepSeek V4 Flash onto high-memory Macs and NVIDIA and AMD boxes, with a local API that Claude Code and Codex CLI can point at.
DwarfStar 4, or ds4, is an MIT-licensed inference engine written in C that runs DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next on high-memory Mac, CUDA and ROCm machines. The project's own benchmark table lists an M5 Max with 128 GB of memory generating 39.4 tokens per second at a 2,048-token context, and 27.6 at 65,536 tokens. Those are the project's numbers, not independent measurements, but they describe a model with 284 billion parameters answering at conversational speed on a single workstation.
How a 284-billion-parameter model fits on one machine
DeepSeek V4 Flash is a mixture-of-experts model. Instead of running every parameter on every word, it contains many specialist sub-networks, called experts, and a router picks a few of them per token. Most of the model's weight sits in those routed experts, which makes them the obvious place to save memory.
ds4 uses what it calls asymmetric 2-bit quantization. Quantization means storing each number with fewer bits, trading some precision for a much smaller file. The asymmetry is the point: the routed experts get compressed hard, while the shared paths every token passes through stay precise. The project describes this as compressed, not lobotomized. Whether that holds for your workload is something the source does not demonstrate, since it publishes speed figures but no quality comparisons.
Two other design choices matter. The first is that ds4 saves the KV cache to disk. The KV cache is the model's working memory of your prompt so far, and rebuilding it after a restart (re-prefill) is slow on long contexts. ds4 stores long prefixes on SSD and resumes by prompt hash. The second is that the same engine serves a chat CLI, a local HTTP server and a persistent coding agent, all sharing one model state and cache.
Who gains, and who is exposed
The server exposes /v1/chat/completions, /v1/messages and /v1/responses, and the project lists OpenCode, Claude Code, Codex CLI and Pi as clients. In practical terms, a team can aim its existing coding tools at a machine under a desk instead of a metered cloud endpoint. The beneficiaries are teams with code or documents they cannot send to a third party, and anyone whose token bill scales with agent usage.
The exposed party is anyone whose business is reselling access to these open-weight models. The source does not discuss pricing, so how much money moves is unknown. But the hardware bar is specific: V4 Flash Q2 is the baseline for 128 GB machines, GLM 5.3 Q2 and Qwen Q4 also fit at that size, V4.1 Q2 streams from SSD, and the page flags Qwen as workable on 64 GB.
There is a quieter exposure on the buyer's side. Moving inference in-house moves the operational burden with it: patching, capacity, and the question of who is on call when the engine misbehaves. The source describes ds4 as a narrow engine, which cuts both ways.
What the benchmarks do and don't say
The numbers are uneven by hardware. The DGX Spark, also at 128 GB, shows 825.8 tokens per second of prefill at 2,048 tokens but only 18.1 generating, against the M5 Max's 39.4. Prefill is reading your prompt; generation is writing the answer. For long documents and agent sessions, prefill speed and cache reuse matter as much as the headline generation rate. The page also labels its figures estimates drawn from the project's benchmark table.
Questions You Should Be Asking
- What does 2-bit quantization of the routed experts cost in answer quality on our tasks, and where is the evidence beyond speed tables?
- If we move agent traffic local, who owns uptime, updates and security fixes for a young, narrowly scoped C engine?
- At roughly 28 tokens per second with a 65,536-token context, is a multi-step coding agent fast enough that engineers will actually use it?
- Which of our current cloud spend is really model access, and which is the surrounding service we would have to rebuild?
- Do the licenses on the model weights, not just ds4's MIT license, permit how we plan to use them?
What To Watch Next
The signal is independent quality testing of the Q2 builds against the full-precision models, run on real coding and long-context work rather than the project's own speed tables. If those results hold up, the question shifts from whether local frontier inference works to which workloads should leave the cloud first.
- 1Deploy DeepSeek V4 Flash using DwarfStar 4 on your Mac to run 284B models locally without cloud costs.
- 2Test context length trade-offs: expect 39.4 tokens/sec at 2K context vs 27.6 at 65K on M5 Max 128GB.
- 3Leverage mixture-of-experts architecture to run massive models on consumer hardware by using only activated experts.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
