Calibration
In Short
Calibration is the recurring practice of having reviewers score the same sample material and comparing the results, so that everyone applying a rubric reads it the same way. It turns "the rubric says so" into "our reviewers agree on what the rubric says", and it catches individual drift before customers do.
Definition
A rubric is only as consistent as the people applying it. Two reviewers can read the same level descriptor — "clear, with minor ambiguity" — and draw the line in different places. Neither is careless; the words simply leave room. Calibration closes that room with evidence rather than exhortation.
The practice has four parts.
Benchmark samples. A curated set of documents with agreed expected scores. Good benchmarks include borderline cases, because easy examples produce agreement that does not survive real work.
Independent scoring. Reviewers score the benchmarks without seeing each other's answers. Discussion comes after, not during, or the exercise measures persuasion instead of interpretation.
Comparison. Each reviewer's scores are compared with the expected scores and with the group. The useful output is not an average but a pattern: a reviewer who is consistently one level harsh on a single dimension, or a dimension where the whole group scatters.
Follow-up. Patterns lead to targeted action — a conversation, a clarified descriptor, a new benchmark. A scattered dimension often means the rubric wording needs work, not the reviewers.
The distinction people miss is between calibration and quality control. Quality control checks finished work; calibration checks the instrument and the people using it, before the work is done.
Why It Matters
Customers compare scorecards over time and across documents. If scores depend on which reviewer picked up the work, those comparisons mean nothing, and a trend in the customer's content becomes indistinguishable from a change in staffing. Calibration is what makes a score a measurement.
How QueryTek Uses It
QueryTek Review runs calibration against benchmark sets for each module rubric, measures how far individual scoring departs from expected values, and flags reviewer drift for follow-up by quality staff. Benchmark contents, tolerance thresholds, and calibration frequency are internal quality controls and are not published.
Related Terms