Reliability
Inter-rater reliability between gold and the judge — the research-grade view: chance-corrected agreement (Cohen’s κ), ordinal-weighted κ, Krippendorff’s α, and error, per curriculum and pooled.
Pooled exact match
91.2%
589 units across 24 curricula
Pooled weighted κ
0.87
sample-size weighted
Curricula
24
gold vs judge as raters
| Curriculum | Exact | Po | κ | κw | α | Self-αi | MAE | n |
|---|---|---|---|---|---|---|---|---|
| Open Up Resources 6-8 Mathindicator | 100.0% | 1.00 | 1.00 | 1.00 | 1.00 | — | 0.00 | 19 |
| Kendall Hunt's Illustrative Mathematics Traditionalindicator | 100.0% | 1.00 | 1.00 | 1.00 | 1.00 | — | 0.00 | 17 |
| McGraw-Hill Illustrative Mathematics AGAindicator | 100.0% | 1.00 | 1.00 | 1.00 | 1.00 | — | 0.00 | 16 |
| Imagine Learning Illustrative Mathematics IM 9-12 Mathindicator | 100.0% | 1.00 | 1.00 | 1.00 | 1.00 | — | 0.00 | 17 |
| Illustrative Mathematics (Kendall Hunt OER)indicator | 100.0% | 1.00 | 1.00 | 1.00 | 1.00 | — | 0.00 | 29 |
| Kendall Hunt's Illustrative Mathematics 6-8 Mathindicator | 100.0% | 1.00 | 1.00 | 1.00 | 1.00 | — | 0.00 | 19 |
| Open Up High School Mathematics Traditionalindicator | 96.7% | 0.97 | 0.94 | 0.90 | 0.90 | — | 0.07 | 30 |
| Kendall Hunt's Illustrative Mathematics K-5indicator | 96.2% | 0.96 | 0.84 | 0.81 | 0.81 | — | 0.08 | 26 |
| Open Up Resources K-5 Mathindicator | 96.2% | 0.96 | 0.84 | 0.81 | 0.81 | — | 0.08 | 26 |
| OpenSciEd Physicsindicator | 96.0% | 0.96 | 0.93 | 1.00 | 1.00 | — | 0.04 | 25 |
| OpenSciEd Biologyindicator | 96.0% | 0.96 | 0.93 | 1.00 | 1.00 | — | 0.04 | 25 |
| Open Up Resources 6-8 Mathematics (3rd Edition)indicator | 95.2% | 0.95 | 0.65 | 0.89 | 0.89 | — | 0.05 | 21 |
| OpenSciEd Chemistryindicator | 92.0% | 0.92 | 0.87 | 0.99 | 0.99 | — | 0.08 | 25 |
| EL Education K-5 Language Arts (K-2)indicator | 91.7% | 0.92 | 0.83 | 0.91 | 0.92 | — | 0.11 | 36 |
| OpenSciEd 6-8 Scienceindicator | 90.2% | 0.90 | 0.79 | 0.90 | 0.90 | — | 0.12 | 41 |
| EL Education K-5 Language Arts (3-5)indicator | 87.2% | 0.87 | 0.75 | 0.90 | 0.90 | — | 0.15 | 39 |
| EL Education 6-8 Language Arts (Open Up Resources edition)indicator | 85.7% | 0.86 | 0.74 | 0.84 | 0.84 | — | 0.21 | 28 |
| OUR Odell HSLPindicator | 83.3% | 0.83 | 0.71 | 0.84 | 0.85 | — | 0.23 | 30 |
| Eureka Math / EngageNY (K-5)indicator | 82.4% | 0.82 | 0.29 | 0.67 | 0.67 | — | 0.18 | 34 |
| EL Education K-8 Language Artsindicator | 82.1% | 0.82 | 0.67 | 0.76 | 0.77 | — | 0.29 | 28 |
| Imagine Learning EL Education 6-8 ELAindicator | 78.6% | 0.79 | 0.59 | 0.69 | 0.69 | — | 0.36 | 28 |
| Eureka Math / EngageNY (6-8)indicator | 75.0% | 0.75 | 0.34 | 0.47 | 0.45 | — | 0.33 | 24 |
| The Utah Middle School Math Projectgateway | 100.0% | 1.00 | 1.00 | 1.00 | 1.00 | — | 0.00 | 3 |
| EngageNY ELAgateway | 100.0% | 1.00 | 1.00 | 1.00 | 1.00 | — | 0.00 | 3 |
Gold and the judge treated as two raters over each curriculum’s criteria. κw (quadratic-weighted Cohen’s κ) is the headline — it credits near-misses on the ordinal 0/1/2/4 scale; κ and α are unweighted/interval cross-checks. Self-αi is the judge’s self-agreement across K self-consistency passes (the noise floor): a judge↔gold gap is trustworthy signal only when the judge agrees with itself more than with gold.