This browser LLM benchmark grades your GPU as much as the model
MicroLLM Lab runs quantized tiny models locally with no server, and hands you a shareable certificate. What it actually measures is the harder question.
A benchmark that tells you its models will fail
MicroLLM Lab, an experiment posted to Hacker News, loads small quantized language models directly into a browser tab and runs a test suite against them. The most honest line in the whole interface is in its own explainer: the checks are "objective checks (regex / exact tokens), not writing quality," and "a 135M model is allowed to fail — that is the measurement."
That is a rarer stance than it sounds. Most model demos are built to make the model look capable. This one is built to find out whether a very small model can hit an exact target, and to time how fast it misses.
What is actually happening in the tab
The page describes itself as initializing a WebGPU engine and model catalog. WebGPU is the browser interface that lets a web page use your graphics hardware for general computation, which is what makes local inference viable without a native app. The models are described as Q4 — four-bit quantization, meaning the model's weights have been compressed from higher-precision numbers down to four bits each. That shrinks a model enough to download and hold in memory on ordinary hardware, at some cost to output quality.
Once loaded, weights are cached in IndexedDB, the browser's local database, so a second visit does not re-download everything. There are controls to load all models, unload all, and pick an active model, plus a token cap per generation.
The measurement side splits speed into two numbers that people routinely conflate. Peak speed is the fastest single test. Sustained speed comes from a continuous 256-token decode — a longer run where thermal throttling, memory pressure and a busy machine start to show. There is also an average sustained figure across tested models and a cumulative suite wall time, charted as runtime per test and accuracy per test. The page states plainly that all of this comes from runs in your browser and that "numbers stay on this machine."
Where the claim and the shipped thing diverge
What is demonstrated: the thing runs, locally, and reports real timings from your own device. That is not a rendered mockup, and the privacy claim follows structurally from the design — if nothing leaves the tab, nothing leaves the tab.
What is a demo, in the sense of being suggestive rather than conclusive:
- Every speed number is a joint measurement of the model and your hardware, browser, drivers and whatever else is running. Two people comparing scores are comparing laptops as much as models.
- The page offers a "Verified Benchmark Certificate" with your handle, device information, peak and sustained tokens per second, downloadable as a PNG and shareable to X and LinkedIn. Nothing in the source explains what verifies it. A self-generated image of a self-reported local run is a screenshot with better typography.
- Accuracy is a pass rate on regex and exact-token checks. That measures format compliance and narrow correctness, not reasoning or usefulness. The tool says as much; the certificate does not carry that caveat.
- There is a custom benchmark editor, and the source states that the editor "is eval()'d in this origin." In plain terms, code you paste in runs as trusted JavaScript inside the page. Fine for your own snippets. Not fine for a benchmark a stranger sends you.
Questions You Should Be Asking
- If two devices produce different tokens-per-second for the same model, what exactly does a shared "score" prove to the person reading it?
- What makes the certificate "verified" — is there any check beyond the browser reporting its own numbers?
- Does a regex pass rate predict anything about how a tiny model behaves on your actual workload, where the correct answer is not a fixed string?
- Before pasting someone else's benchmark into an editor that eval()s in the page origin, what else is in that origin?
- If Q4 quantization is the reason these models fit in a tab, how much accuracy did that compression cost, and against which unquantized baseline?
What To Watch Next
The signal is whether the project publishes the model list and test definitions in a form anyone can re-run and diff, alongside a normalization for hardware. Without that, the certificates will circulate as device bragging rights. With it, this becomes a cheap, honest way to find out which tiny model clears your bar on your own machine — which is the only benchmark that ever mattered for local inference.
- 1Report your GPU, browser, and VRAM alongside any in-browser LLM score, since WebGPU throughput varies more than the model does.
- 2Use regex or exact-token checks to grade small models, and treat a 135M model's failures as valid data rather than a broken test.
- 3Warm up the model with a throwaway prompt before timing, so shader compilation and weight loading don't inflate your first-token latency.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
