Machines
Available Models
Ranked by quant score — base benchmark scores adjusted for the quantization on disk, so a Q2 70B and a Q8 8B compare directly. Click any column to sort.
How the quant score is calculated
Each local model is matched to its base modelin the catalog (quantization suffixes stripped), and scored on that base model's external benchmark results — Open LLM Leaderboard, Arena, Artificial Analysis, OpenCompass, LLM Stats and others.
Scores are percentile ranks across the whole catalog rather than raw benchmark averages. The underlying benchmarks are on incompatible scales — Arena is an Elo figure around 900–1500, LiveBench is a 0–1 fraction, most others are 0–100 percentages — so averaging raw values would be dominated by whichever benchmark has the largest numbers. Percentiles put every benchmark on a common 0–100 basis.
That base score is then reduced once by the estimated quality loss for the quant on disk (roughly 0.5% at Q8_0, 3% at Q4_K_M, 10% at Q2_K), giving the quant score used for ranking. Because two quants of the same model resolve to the same base, the only thing separating them is the quantization penalty.
Locally-run benchmark results (lc*) are deliberately excluded from the base score — they were measured on the quantized weights already, so including them would count the same degradation twice.
Throughput is measured where a local speed benchmark has been run, and otherwise estimated from active parameters and quantization — estimates are marked Est. For MoE models the bracketed figure in Params is the active parameter count, which is what drives both memory bandwidth and speed.
The degradation figures are published estimates, not per-model measurements. Treat the ranking as a sound like-for-like ordering rather than a precise quality prediction, and note that models tagged fuzzy were matched to their base by an approximate name match.