small-llm-bench leaderboard

What the score is

Models are ranked on the judged passk rate — the module-weighted share of tasks a model passed on all k independent trials, after the LLM judge's adjustments. The charts show it as a percentage; the Detailed tab's "Judged pass" column is the same number on a 0–1 scale. A model with no judge run falls back to its unjudged passk, is drawn with a hatched bar, and is flagged in its detail card — its score is not strictly comparable to a judged one.

What the tiers are, and are not

Tiers are fixed bands on that score (). They are an editorial cut, chosen to stay put between runs — a tier says a model scored in a given range and nothing more. They are not a claim that the bank can tell one tier from another. That question is answered by the pairs separable chip, which counts how many of the model pairs on this board the bank can actually distinguish, by a paired sign test over the tasks both models ran, Holm-corrected across every pair shown. When that count is low the fix is better tasks, not a longer run — and adding models lowers it, because the correction is over the number of simultaneous comparisons. The per-pair verdicts ship in the sibling .json.

What the By params tab groups on

Buckets cut on total parameters — what you have to load — while the Params column ranks a sparse model by its active parameters, which is what a forward pass costs. The two disagree on purpose: ornith-1.5-35b is a 35B model with ~3B active, so it sorts among the 3B rows and sits in the 24B-and-above bucket, and both are right. A MOE or PLE badge marks a model that activates less than it loads; an unbadged row is dense. Comparing the two halves of a bucket asks the useful question: at this footprint, what does sparsity cost?

Reading the Detailed tab

Module columns show the raw deterministic score, with a +delta/−delta pill when a model has been judged (judged score minus raw). "Det score" is the weighted mean task score; "Pass" is the passk headline rate. "Judged det"/"Judged pass" repeat those two after the LLM judge's adjustments, where available. A ⚠ next to a model means the judge did not cover every trial (a * marks which modules): those modules fall back to their deterministic score, so the judged columns of that row are a blend and are not comparable to a fully-judged run. format and knowledge are never sent to the judge by design, so they carry no delta.