Skip to content

Reliability

Inter-rater reliability between gold and the judge — the research-grade view: chance-corrected agreement (Cohen’s κ), ordinal-weighted κ, Krippendorff’s α, and error, per curriculum and pooled.

Pooled exact match

91.2%

589 units across 24 curricula

Pooled weighted κ

0.87

sample-size weighted

Curricula

24

gold vs judge as raters
Inter-rater reliability between gold and the judge, per curriculum
CurriculumExactPoκκwαSelf-αiMAEn
Open Up Resources 6-8 Mathindicator100.0%1.001.001.001.000.0019
Kendall Hunt's Illustrative Mathematics Traditionalindicator100.0%1.001.001.001.000.0017
McGraw-Hill Illustrative Mathematics AGAindicator100.0%1.001.001.001.000.0016
Imagine Learning Illustrative Mathematics IM 9-12 Mathindicator100.0%1.001.001.001.000.0017
Illustrative Mathematics (Kendall Hunt OER)indicator100.0%1.001.001.001.000.0029
Kendall Hunt's Illustrative Mathematics 6-8 Mathindicator100.0%1.001.001.001.000.0019
Open Up High School Mathematics Traditionalindicator96.7%0.970.940.900.900.0730
Kendall Hunt's Illustrative Mathematics K-5indicator96.2%0.960.840.810.810.0826
Open Up Resources K-5 Mathindicator96.2%0.960.840.810.810.0826
OpenSciEd Physicsindicator96.0%0.960.931.001.000.0425
OpenSciEd Biologyindicator96.0%0.960.931.001.000.0425
Open Up Resources 6-8 Mathematics (3rd Edition)indicator95.2%0.950.650.890.890.0521
OpenSciEd Chemistryindicator92.0%0.920.870.990.990.0825
EL Education K-5 Language Arts (K-2)indicator91.7%0.920.830.910.920.1136
OpenSciEd 6-8 Scienceindicator90.2%0.900.790.900.900.1241
EL Education K-5 Language Arts (3-5)indicator87.2%0.870.750.900.900.1539
EL Education 6-8 Language Arts (Open Up Resources edition)indicator85.7%0.860.740.840.840.2128
OUR Odell HSLPindicator83.3%0.830.710.840.850.2330
Eureka Math / EngageNY (K-5)indicator82.4%0.820.290.670.670.1834
EL Education K-8 Language Artsindicator82.1%0.820.670.760.770.2928
Imagine Learning EL Education 6-8 ELAindicator78.6%0.790.590.690.690.3628
Eureka Math / EngageNY (6-8)indicator75.0%0.750.340.470.450.3324
The Utah Middle School Math Projectgateway100.0%1.001.001.001.000.003
EngageNY ELAgateway100.0%1.001.001.001.000.003

Gold and the judge treated as two raters over each curriculum’s criteria. κw (quadratic-weighted Cohen’s κ) is the headline — it credits near-misses on the ordinal 0/1/2/4 scale; κ and α are unweighted/interval cross-checks. Self-αi is the judge’s self-agreement across K self-consistency passes (the noise floor): a judge↔gold gap is trustworthy signal only when the judge agrees with itself more than with gold.

Live read from the Signal Studio spine (core-data).