About

I am a hobbyist. Not a researcher, not a reviewer with vendor relationships, not a lab. I tinker with AI and computing tools because I find them genuinely interesting, and this site is what came out of doing a lot of that in one place.

The specific thing that started it was an NVIDIA DGX Spark sitting on my desk and a list of questions nobody had answered. Is a mixture-of-experts model really faster than a dense model of similar size, and by how much? Does speculative decoding pay for itself in practice? What actually happens to throughput after a system update? What I could find was vendor slides, forum anecdotes, and numbers with no method attached to them.

So I measured, and then I wrote down what I found.

What this site is

Every figure published here is read out of a file produced by a benchmark harness. Nothing is typed into a table by hand, which is the whole reason the numbers in an article always match the run that produced them. The methodology page describes how those runs are conducted in enough detail to reimplement, and the hardware page is generated from the same data.

The measurements are automated. The conclusions are not — I write the analysis, the recommendations and the caveats myself, and that division is deliberate. Generated tables are reliable in a way generated opinions are not.

What this site is not

I have no vendor relationships, and nothing here was supplied by a manufacturer. That also means I can only test hardware I actually own, so the coverage is narrow and biased toward whatever I happened to buy. A machine's absence from this site says nothing about it.

I get things wrong. Two articles carry published corrections: one where the harness reported a throughput that exceeded what the memory bus can physically deliver, and one where a caching effect inflated a result by 81% before I caught it. Both corrections are printed on the articles themselves rather than quietly edited in, and the superseded data files are kept alongside an explanation of why they were wrong. A site about measurement that hides its bad measurements is not worth reading.

The machine is not a benchmark rig

The Spark does real work for other projects, which makes benchmarking it harder and the results more honest. The harness refuses to measure while another workload is on the GPU, and it has cost me results — one sweep was quarantined mid-run because a scheduled job from another project started at 21:00 UTC. Two measurements lost, two wrong numbers prevented. That is the right trade: a benchmark taken under contention is not a slow measurement, it is a wrong one.

Who owns this

Local LLM Labs is independently published and holds the rights to everything on it, and is the data controller described on the privacy page. Everything written here is my own personal opinion, and nothing on this site is written on behalf of any employer of mine or any company whose products it measures.

How this site makes money

I would like this site to pay for the hardware it measures, and in time to pay for more of it. That is a real motive, and pretending it away would be its own small dishonesty on a site whose whole point is not shading the truth.

The revenue comes from affiliate links — buy something through one and this site may earn a commission, at no extra cost to you — and the site may carry advertising later. Any post containing affiliate links opens with a disclosure saying so, above the article rather than buried under it.

What that money does not buy is a number, and the useful thing is that you do not have to take my word for it:

If those three ever stop being true, this site is worth considerably less than the commission.

Get in touch

Corrections are the most welcome thing you can send, especially if you have measured something that disagrees with what is published here. The contact page describes what makes a report useful, or you can write directly to [email protected].