Skip to content
MathGrades 6-8

Eureka Math / EngageNY (6-8)

Great Minds (EngageNY)

Judge anthropic/claude-opus-4-7 · Signal Studio judge (claude-opus-4-7) · 2026-05-28

75

% indicator exact match

Gateways 3/3 Meets

Agreement vs gold

Indicator level — gold and the judge as two raters over the rubric criteria.

Exact match

75.0%

24 indicators vs gold

Weighted κ

0.47

ordinal-weighted agreement

MAE

0.33

lower is better

Signed bias

+0.33

judge over-scores

Attribution

Where the judge diverges, and whether it's signal: per-gateway bias localizes the gap, the confusion matrix shows over- vs under-rating, and the noise floor is the judge's self-agreement across passes (a gap below it isn't trustworthy).

gateway-onen=6100.0%bias +0.00
gateway-threen=955.6%bias +0.67
gateway-twon=977.8%bias +0.22

Indicator confusion · rows = gold, cols = judge

MeetsPartialDNM
Meets1700
Partial410
DNM200

Diagonal = agreement. Amber = judge rated higher than gold (over-rating).

Gateway rollup

Indicator scores rolled up to EdReports gateway ratings (sequential gating + no-0s cap), gold vs judge.

G1 · Focus & Coherence

14/14 pts

GoldMeets
JudgeMeets
agree

G2 · Rigor & Mathematical Practices

22/18 pts

GoldMeets
JudgeMeets
agree

G3 · Usability

17/18 pts

GoldPartially Meets
JudgeMeets
differ

Gateway-level agreement: exact 67% · κ 0.00 (3 gateways)

Divergences (6)

Indicators where the judge's score differed from gold. Amber = judge under-scored, blue = over-scored.

2g-ii
12
gold Partially Meets Expectations · judge 2 pts
2f
12
gold Partially Meets Expectations · judge 2 ptsinsufficient evidence
3f
12
gold Partially Meets Expectations · judge 2 ptsinsufficient evidence
3p-ii
02
gold Does Not Meet Expectations · judge 2 pts
3m
02
gold Does Not Meet Expectations · judge 2 pts
3p-i
12
gold Partially Meets Expectations · judge 2 pts
Live read from the Signal Studio spine (core-data).