McGraw-Hill Illustrative Mathematics AGA
McGraw-Hill Education
Judge anthropic/claude-opus-4-7 · Signal Studio judge (claude-opus-4-7) · 2026-05-28
% indicator
exact match
Agreement vs gold
Indicator level — gold and the judge as two raters over the rubric criteria.
Exact match
100.0%
Weighted κ
1.00
MAE
0.00
Signed bias
+0.00
Attribution
Where the judge diverges, and whether it's signal: per-gateway bias localizes the gap, the confusion matrix shows over- vs under-rating, and the noise floor is the judge's self-agreement across passes (a gap below it isn't trustworthy).
Indicator confusion · rows = gold, cols = judge
| Meets | Partial | DNM | |
|---|---|---|---|
| Meets | 16 | 0 | 0 |
| Partial | 0 | 0 | 0 |
| DNM | 0 | 0 | 0 |
Diagonal = agreement. Amber = judge rated higher than gold (over-rating).
Gateway rollup
Indicator scores rolled up to EdReports gateway ratings (sequential gating + no-0s cap), gold vs judge.
G1 · Focus & Coherence
18/18 pts
G2 · Rigor & Mathematical Practices
16/16 pts
G3 · Usability
16/16 pts
Gateway-level agreement: exact 100% · κ 1.00 (3 gateways)
Divergences (0)
Indicators where the judge's score differed from gold. Amber = judge under-scored, blue = over-scored.
No divergences
Judge and gold agreed on every scored indicator.