How quality is measured

Validation returns two independent numbers for every datapoint, both 0-100:

  • Quality Ratingwhat the panel concluded. Directional: 100 is high quality, 0 is low.
  • Consensus Strengthhow tightly they agreed, regardless of the answer. See How consensus works.

This page is about the Quality Rating, and how to read it against Consensus Strength.

What the Quality Rating measures

To judge quality, a panel of validators rates each item on a scale — say an AI's report graded from Poor to Excellent. The Quality Rating is where the panel landed, mapped onto 0-100. The rubric decides which end is "good", so the rating carries the direction of the answer, not just how far apart the votes were.

Labels map evenly onto 0–100 (Poor = 0 … Excellent = 100); the Quality Rating is where the panel's picks land — here 80.

How it rolls up

One report is usually scored on several dimensions (accurate, complete, clear), and the rating rolls up from there:

Per-dimension ratings blend into one score per report (the datapoint); every report's score then averages into the dataset headline.

Quality is not the same as agreement

The two numbers are independent. A panel can agree completely that a report is poor — high Consensus Strength, low Quality Rating.

Both panels agree just as strongly, so Consensus Strength matches — but they agree on opposite things, so the Quality Rating is far apart.

That independence is why you keep both numbers: Consensus Strength is the one that acts — it decides whether a datapoint is settled or needs another look — while the Quality Rating is only ever reported.

Reading the two together: the outcome quadrant

Split each axis at 70 and every datapoint lands in one of four cells:

Consensus Strength (horizontal) says whether the result is reliable; Quality Rating (vertical) says the verdict. Right column = settled; left column = escalate.

Consensus Strength decides whether to trust the result (right column = settled); Quality Rating decides the verdict (top row = passed). Confidently rejected is a healthy outcome — the panel reliably caught weak work.

Pushing a datapoint rightward means escalating to more validators: it raises consensus, but each added reviewer is paid — so escalate only where more confidence would change a decision.

Where you'll see it

  • In the app — on each datapoint, beside its Consensus Strength and outcome.
  • In the CSV export — a column you can sort and threshold.
  • In the PoQ Report — the signed record, per datapoint and per dataset.

See also

PageDescription
How consensus worksThe other rating: how agreement is measured and how split panels escalate
Validation interfaceHow validators score each dimension on the 0-100 scale
PoQ ReportThe record where the Quality Rating and Consensus Strength are certified
The PoQ workflowWhere validation fits: Define → Validate → Attest