Free tools

Sample size & scorecard health

Two questions every QA lead gets asked: “how many evaluations is enough?” and “is our form any good?”

QA sample-size calculator

How many evaluations per agent per month to know each agent’s average score within a margin you choose.

±5 means “the agent’s true average is within 5 points of what we measured”.

How much one agent’s scores vary between evaluations. 8–12 is typical for a settled scorecard; use 15 for a new one. Your log shows the real figure.

Only matters for low-volume agents (finite-population correction).

Evaluations per agent per month

–

How this is calculated

We treat each agent’s monthly QA average as an estimate of their true average and ask how many evaluations make that estimate precise enough. The margin of error for a mean is

E = t(1 − α/2, n − 1) × σ / √n

where σ is the spread of one agent’s scores and t is Student’s t with n − 1 degrees of freedom (with only a handful of evaluations you’re also estimating σ, so t is larger than the normal z). We start from the normal approximation n₀ = (z·σ/E)² and then take the smallest n whose t-based margin meets your target.

When an agent handles few contacts, we apply the finite-population correction n′ = n / (1 + (n − 1)/N). Example: σ = 10, ±5 points, 95% confidence → z gives 16, the t correction gives 18 per agent per month.

What it doesn’t tell you: whether your evaluations are a random sample (pick contacts at random, not just escalations) or whether reviewers score consistently (that’s what calibration is for).

Scorecard health check

Checks a scorecard for common design problems: too many or too few questions, duplicate or double-barrelled wording, over-weighted questions, weights that don’t add up, and too many (or no) critical items.

…or paste scorecard JSON