Figure 1 — paired power calculator
questions needed
—
Minimum detectable effect vs. sample size
Every point on the curve is the smallest gap you could reliably detect at that many questions. The marker is your current effect size.
Reading the numbers
- Baseline accuracy
- Roughly where your models score. Sets the per-question variance, p(1-p).
- Effect size (δ)
- The accuracy gap between two models you want to be able to detect, not just observe.
- Correlation (ρ)
-
How often both models get the same questions right or wrong. Harder questions are
harder for everyone, so ρ is usually positive — and pairing on it shrinks the standard
error (see
comparein the Python package). - Cluster size & ICC
- If questions come in groups (e.g. several per reading passage) that share difficulty, effective sample size shrinks by the design effect, 1 + (size − 1) × ICC.
Full derivations: docs/formulas.md §11. This page's TypeScript implementation is checked against the Python one on every commit — see web/src/power.test.ts.