Skip to content

Flagged for human review

Every indicator here tripped a confidence, evidence-quality, or cross-pass signal — the calls where a reviewer’s judgment is needed. Everything else, the AI scored on its own.

Curricula scanned

24

latest judge run each

Auto-finalizable

493

71% of 691 — high-confidence, well-evidenced

Awaiting review

198

29% flagged across 24 curricula

Decided

0

0% of flagged reviewed

Active flags

Passes disagreed0 findings

cross-pass

Independent judge passes (K>1) proposed different scores for this indicator. Cross-pass disagreement is the strongest signal that the score is unstable. Only fires on multi-pass runs; the production judge is single-pass.

Insufficient evidence11 findings

self-reported

The judge flagged that the retrieved evidence was insufficient to assess the indicator — it was guessing, not assessing. The judge’s own "I don’t have enough to go on" signal, and a direct route to a human.

Thin evidence180 findings

objective

Retrieval surfaced few chunks, weakly-matched chunks (low rerank relevance), or evidence concentrated in a single source. Confidently scoring on thin evidence is the most insidious failure mode — and confidence alone does not catch it.

Borderline19 findings

positional

The score sits on a rating boundary: moving it by one point would change the EdReports rating (Does Not Meet / Partially Meets / Meets). The calls where a point in either direction matters most.

Low confidence37 findings

self-reported

The judge self-reported low confidence in its score. A useful signal but the weakest of the set — models can be confidently wrong, so confidence alone is not enough.

The Utah Middle School Math Project

Math6-8

19 flagged of 26 · judged by anthropic/claude-opus-4-7

Eureka Math / EngageNY (K-5)

MathK-5

16 flagged of 36 · judged by anthropic/claude-opus-4-7

Eureka Math / EngageNY (6-8)

Math6-8

13 flagged of 26 · judged by anthropic/claude-opus-4-7

OpenSciEd 6-8 Science

ELA6-8

13 flagged of 41 · judged by anthropic/claude-opus-4-7

Open Up Resources 6-8 Math

Math6-8

11 flagged of 26 · judged by anthropic/claude-opus-4-7

OUR Odell HSLP

ELA9-12

11 flagged of 30 · judged by anthropic/claude-opus-4-7

EL Education K-5 Language Arts (K-2)

ELAK-2

10 flagged of 38 · judged by anthropic/claude-opus-4-7

Open Up High School Mathematics Traditional

MathHS

10 flagged of 30 · judged by anthropic/claude-opus-4-7

EngageNY ELA

ELA6-8

9 flagged of 30 · judged by anthropic/claude-opus-4-7

Kendall Hunt's Illustrative Mathematics 6-8 Math

Math6-8

9 flagged of 26 · judged by anthropic/claude-opus-4-7

Open Up Resources 6-8 Mathematics (3rd Edition)

Math6-8

9 flagged of 26 · judged by anthropic/claude-opus-4-7

Imagine Learning EL Education 6-8 ELA

ELA6-8

7 flagged of 30 · judged by anthropic/claude-opus-4-7

EL Education K-5 Language Arts (3-5)

ELA3-5

6 flagged of 41 · judged by anthropic/claude-opus-4-7

EL Education K-8 Language Arts

ELA6-8

6 flagged of 30 · judged by anthropic/claude-opus-4-7

Kendall Hunt's Illustrative Mathematics K-5

MathK-5

6 flagged of 26 · judged by anthropic/claude-opus-4-7

Open Up Resources K-5 Math

MathK-5

6 flagged of 26 · judged by anthropic/claude-opus-4-7

OpenSciEd Chemistry

ELA9-12

6 flagged of 25 · judged by anthropic/claude-opus-4-7

OpenSciEd Physics

ELA9-12

6 flagged of 25 · judged by anthropic/claude-opus-4-7

EL Education 6-8 Language Arts (Open Up Resources edition)

ELA6-8

5 flagged of 30 · judged by anthropic/claude-opus-4-7

Illustrative Mathematics (Kendall Hunt OER)

MathK-8

4 flagged of 29 · judged by anthropic/claude-opus-4-7

Imagine Learning Illustrative Mathematics IM 9-12 Math

MathHS

4 flagged of 23 · judged by anthropic/claude-opus-4-7

Kendall Hunt's Illustrative Mathematics Traditional

MathHS

4 flagged of 23 · judged by anthropic/claude-opus-4-7

McGraw-Hill Illustrative Mathematics AGA

MathHS

4 flagged of 23 · judged by anthropic/claude-opus-4-7

OpenSciEd Biology

ELA9-12

4 flagged of 25 · judged by anthropic/claude-opus-4-7

Flags are computed from the judge’s own signals (confidence, retrieved-evidence quality, cross-pass agreement) — never from a comparison to gold — so the same triage works for new reviews with no published report.