Model Insights
Computed live from the 6 run bundles in /runs (1200 items total) — open-ended track, primary (judge-of-record) and secondary scores.
Key takeaways
kimi-k3 scores highest at 76.5%, 23.5 points ahead of kimi-k2.6_strict.
The two models' 95% confidence intervals don't overlap, so this gap is unlikely to be sampling noise alone.
"Physical mobility" is the hardest category — it's each model's own weakest spot in 5 of 6 trials.
That consistency across otherwise very different models suggests the difficulty sits in the category itself (harder scenes, more ambiguous ground truth) rather than in any one model's training.
Primary and secondary judges agree on 85% of items on average.
Agreement is lowest for gemini-3.1-flash-image_strict (82%) — when two independent judges disagree that often, the score for that model is more sensitive to which judge you trust.
Across all models, "spatial correctness" is the weakest rubric dimension (56% pass rate).
Compared to "intent match" at 87%, models are relatively better at naming the right visual cues than at getting the spatial or directional detail right — the failure mode is usually "close, but misplaced," not "unrelated."
Judge scores are validated against hand-labeled human judgments (4 of 6 trials) — the primary judge is the more human-aligned one in 2 of them.
Average primary-judge reliability is κ ≈ 0.56 (Cohen's κ; Landis & Koch scale). The weakest case is kimi-k2.6_strict, where even its better judge only reaches κ = 0.63 against human labels — accuracy numbers for that trial should be read with real caution, not treated as ground truth.
Accuracy by model
Hover a bar for the secondary judge's score and judge agreement.
Accuracy by persona category
All six models share the same four accessibility personas.
Faded bars mark categories with fewer than 5 items — too small to read as a real rate.
Model comparison
"Judges agree" is how often the two judges scoring this model agree with each other. The last two columns are each judge's κ against a small hand-labeled human sample (Landis & Koch: <0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect) — that's what actually justifies trusting a judge's numbers, not judge-vs-judge agreement alone.
| Model | N | Accuracy | Weakest category | Judges agree | Primary judge vs. human (κ) | Secondary judge vs. human (κ) |
|---|---|---|---|---|---|---|
| kimi-k3 moonshotai | 200 | 76.5% (153/200) | Physical mobility (69.8%) | 85.5% | claude-sonnet-4.5 κ 0.342 (75.0%) | gpt-4o-mini κ 0.634 (85.0%) |
| gemini-3.1-flash-image google | 200 | 72.5% (145/200) | Physical mobility (67.4%) | 89.0% | claude-sonnet-4.5 κ 0.857 (95.0%) | gpt-4o-mini κ 0.483 (85.0%) |
| kimi-k2.6 moonshotai | 200 | 68.5% (137/200) | Physical mobility (58.1%) | 86.0% | — | — |
| gemini-3.1-flash-image_strict google | 200 | 56.5% (113/200) | General (52.1%) | 81.5% | claude-sonnet-4.5 κ 0.419 (75.0%) | gpt-4o-mini κ 0.692 (90.0%) |
| minimax_minimax-m3 unknown | 200 | 56.0% (112/200) | Physical mobility (41.9%) | 86.0% | — | — |
| kimi-k2.6_strict moonshotai | 200 | 53.0% (106/200) | Physical mobility (44.2%) | 83.5% | claude-sonnet-4.5 κ 0.634 (85.0%) | gpt-4o-mini κ 0.571 (85.0%) |