errorbars — statistical power calculator

How many eval questions do I need?

A 1.5-point gain on a 500-question benchmark is often noise. This calculator answers the question before you run the eval: given a baseline accuracy, the effect size you care about, how correlated your models' answers are, and how your questions are clustered, how many questions does it take to tell signal from noise? The same formulas back the errorbars power command in the Python package — move a slider here and you're looking at exactly what that command would print.

Figure 1 — paired power calculator

questions needed

—

Minimum detectable effect vs. sample size

Every point on the curve is the smallest gap you could reliably detect at that many questions. The marker is your current effect size.

Reading the numbers

Baseline accuracy
Roughly where your models score. Sets the per-question variance, p(1-p).
Effect size (δ)
The accuracy gap between two models you want to be able to detect, not just observe.
Correlation (ρ)
How often both models get the same questions right or wrong. Harder questions are harder for everyone, so ρ is usually positive — and pairing on it shrinks the standard error (see compare in the Python package).
Cluster size & ICC
If questions come in groups (e.g. several per reading passage) that share difficulty, effective sample size shrinks by the design effect, 1 + (size − 1) × ICC.