An Open-Source Engine Tunes Itself to Your Chip, Claiming 92% Faster
Magnitude says it compiles kernels on your own hardware to run open models up to 2x faster than llama.cpp. Here is how to test that before you rely on it.
Magnitude, a Y Combinator S25 company, has released an open-source inference engine that it says runs open-weight models up to 2x faster than llama.cpp. The company's own figures: 92% faster decode on Apple's Metal and 19% on NVIDIA's CUDA. The software is Apache 2.0 licensed and ships as a desktop app for macOS, Windows and Linux. Those are the company's benchmarks, not independent results, and the gap between 92% and 19% is the first thing an operator should notice.
What 'Tuning on Your Device' Actually Means
An inference engine is the program that runs a trained model and turns your prompt into output. Inside it, the heavy lifting is done by kernels: small, highly specialized routines that perform the core math on a GPU or CPU. How fast a kernel runs depends on details of the specific chip, such as how its memory is arranged and how many calculations it can run in parallel.
Magnitude's argument is that popular tools such as llama.cpp, Ollama and LM Studio ship kernels precompiled for broad classes of hardware. That is convenient, but a kernel built for a category of chips is rarely the best one for your particular chip. Magnitude says it instead compiles and tunes its kernels on your actual machine before a model runs, so they fit your exact silicon. It also says it hand-writes optimized kernels for the most popular open-weight model families, which it credits for beating generalist engines.
The trade-off is worth understanding. Precompiled kernels start instantly and behave the same everywhere. A tune-on-your-device approach adds a preparation step and makes performance depend on the machine it runs on, which is partly why the company's gains differ so much between Metal and CUDA.
What Else It Claims
The company lists several other features. It says agents use 27% less memory, with memory freed when agents stop. It says concurrent sessions share prefix caches, meaning the engine reuses work already done on a shared opening portion of a prompt, so running several agents at once does not slow everything down. It connects with one click to agents including Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi and Cline, and anything else can attach through an OpenAI-compatible API. Everything runs locally: prompts, files and models stay on your machine, and no internet is needed once a model is downloaded. It supports Apple Silicon, NVIDIA, AMD, or a CPU alone, with no fixed minimum hardware; smaller machines run smaller models.
A Sensible Month-Long Plan
Treat this as a bake-off, not a migration. Pick one machine your team actually uses, install Magnitude alongside your current setup, and run the same model on the same prompts through both. Measure the thing you care about, whether that is time to first token, sustained output speed, or how many agent sessions fit before things slow down. If your fleet is mostly NVIDIA machines, the company's own CUDA figure of 19% is the one to test against, not the headline 2x.
The mistake would be rewriting infrastructure around a vendor benchmark, or assuming speed on one chip carries to another. Equally, dismissing it is a mistake: an Apache 2.0 license and a local-only design mean trying it carries little lock-in or data risk.
Questions You Should Be Asking
- On which models, quantization levels and prompt lengths were the 92% and 19% figures measured, and does the benchmark match how our agents really behave?
- How long does the on-device tuning step take, and what happens to speed the first time a new model or a new machine is used?
- The 'up to 2x' headline sits well above the 19% CUDA figure. What hardware do we own, and which number applies to it?
- Which of the supported models are covered by hand-optimized kernels, and how much slower is a model outside that list?
- Can anyone outside the company reproduce these results, and has anyone?
What To Watch Next
The signal is independent, reproducible benchmarks on NVIDIA and AMD hardware, where the company's own number is smallest, along with how quickly the supported-model list grows beyond the families the team hand-tunes. If third-party tests land near the claimed figures and the model list keeps pace with new open-weight releases, per-device tuning becomes a serious challenger to precompiled engines; if the gains shrink outside Apple's chips, it stays a niche advantage.
- 1Test Magnitude's inference engine on your specific hardware before switching; benchmark gains vary dramatically (92% on Apple Metal vs 19% on NVIDIA).
- 2Use the open-source Apache 2.0 licensed engine to optimize model inference locally, avoiding cloud costs for compatible open-weight models.
- 3Compare independent benchmarks against company claims, as performance heavily depends on your exact GPU/CPU architecture and kernel optimization.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
