Validation returns two independent numbers for every datapoint, both 0-100:
Quality Rating — what the panel concluded. Directional: 100 is high quality, 0 is low.
Consensus Strength — how tightly they agreed, regardless of the answer. See
How consensus works.
This page is about the Quality Rating, and how to read it against Consensus Strength.
What the Quality Rating measures
To judge quality, a panel of validators rates each item on a scale — say an AI's report graded from
Poor to Excellent. The Quality Rating is where the panel landed, mapped onto 0-100. The
rubric decides which end is "good", so the rating carries the direction of the answer, not just how
far apart the votes were.
Labels map evenly onto 0–100 (Poor = 0 … Excellent = 100); the Quality Rating is where the panel's picks land — here 80.
How it rolls up
One report is usually scored on several dimensions (accurate, complete, clear), and the rating rolls
up from there:
Per-dimension ratings blend into one score per report (the datapoint); every report's score then averages into the dataset headline.
Quality is not the same as agreement
The two numbers are independent. A panel can agree completely that a report is poor — high
Consensus Strength, low Quality Rating.
Both panels agree just as strongly, so Consensus Strength matches — but they agree on opposite things, so the Quality Rating is far apart.
That independence is why you keep both numbers: Consensus Strength is the one that acts — it
decides whether a datapoint is settled or needs another look — while the Quality Rating is only ever
reported.
Reading the two together: the outcome quadrant
Split each axis at 70 and every datapoint lands in one of four cells:
Consensus Strength (horizontal) says whether the result is reliable; Quality Rating (vertical) says the verdict. Right column = settled; left column = escalate.
Consensus Strength decides whether to trust the result (right column = settled); Quality
Rating decides the verdict (top row = passed). Confidently rejected is a healthy outcome — the
panel reliably caught weak work.
Pushing a datapoint rightward means
escalating to more
validators: it raises consensus, but each added reviewer is paid — so escalate only where more
confidence would change a decision.
Where you'll see it
In the app — on each datapoint, beside its Consensus Strength and outcome.
In the CSV export — a column you can sort and threshold.
In the PoQ Report — the signed record, per datapoint and per dataset.