Model Insights

Computed live from the 6 run bundles in /runs (1200 items total) — open-ended track, primary (judge-of-record) and secondary scores.

Best accuracy
76.5%
kimi-k3 · 153/200
Weakest accuracy
53.0%
kimi-k2.6_strict · 106/200
Spread across models
23.5 pts
best minus weakest
Hardest category
Physical mobility
weakest in 5 of 6 models

Key takeaways

01

kimi-k3 scores highest at 76.5%, 23.5 points ahead of kimi-k2.6_strict.

The two models' 95% confidence intervals don't overlap, so this gap is unlikely to be sampling noise alone.

02

"Physical mobility" is the hardest category — it's each model's own weakest spot in 5 of 6 trials.

That consistency across otherwise very different models suggests the difficulty sits in the category itself (harder scenes, more ambiguous ground truth) rather than in any one model's training.

03

Primary and secondary judges agree on 85% of items on average.

Agreement is lowest for gemini-3.1-flash-image_strict (82%) — when two independent judges disagree that often, the score for that model is more sensitive to which judge you trust.

04

Across all models, "spatial correctness" is the weakest rubric dimension (56% pass rate).

Compared to "intent match" at 87%, models are relatively better at naming the right visual cues than at getting the spatial or directional detail right — the failure mode is usually "close, but misplaced," not "unrelated."

05

Judge scores are validated against hand-labeled human judgments (4 of 6 trials) — the primary judge is the more human-aligned one in 2 of them.

Average primary-judge reliability is κ ≈ 0.56 (Cohen's κ; Landis & Koch scale). The weakest case is kimi-k2.6_strict, where even its better judge only reaches κ = 0.63 against human labels — accuracy numbers for that trial should be read with real caution, not treated as ground truth.

Accuracy by model

Hover a bar for the secondary judge's score and judge agreement.

0%25%50%75%100%76.5%kimi-k3moonshotai72.5%gemini-3.1-flash-imagegoogle68.5%kimi-k2.6moonshotai56.5%gemini-3.1-flash-image_strictgoogle56.0%minimax_minimax-m3unknown53.0%kimi-k2.6_strictmoonshotai

Accuracy by persona category

All six models share the same four accessibility personas.

gemini-3.1-flash-image
gemini-3.1-flash-image_strict
kimi-k2.6
kimi-k2.6_strict
kimi-k3
minimax_minimax-m3
0%25%50%75%100%Aged personGeneralPhysical mobilitySensory impairments

Faded bars mark categories with fewer than 5 items — too small to read as a real rate.

Model comparison

"Judges agree" is how often the two judges scoring this model agree with each other. The last two columns are each judge's κ against a small hand-labeled human sample (Landis & Koch: <0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect) — that's what actually justifies trusting a judge's numbers, not judge-vs-judge agreement alone.

ModelNAccuracyWeakest categoryJudges agreePrimary judge vs. human (κ)Secondary judge vs. human (κ)
kimi-k3
moonshotai
20076.5% (153/200)Physical mobility (69.8%)85.5%
claude-sonnet-4.5
κ 0.342 (75.0%)
gpt-4o-mini
κ 0.634 (85.0%)
gemini-3.1-flash-image
google
20072.5% (145/200)Physical mobility (67.4%)89.0%
claude-sonnet-4.5
κ 0.857 (95.0%)
gpt-4o-mini
κ 0.483 (85.0%)
kimi-k2.6
moonshotai
20068.5% (137/200)Physical mobility (58.1%)86.0%
gemini-3.1-flash-image_strict
google
20056.5% (113/200)General (52.1%)81.5%
claude-sonnet-4.5
κ 0.419 (75.0%)
gpt-4o-mini
κ 0.692 (90.0%)
minimax_minimax-m3
unknown
20056.0% (112/200)Physical mobility (41.9%)86.0%
kimi-k2.6_strict
moonshotai
20053.0% (106/200)Physical mobility (44.2%)83.5%
claude-sonnet-4.5
κ 0.634 (85.0%)
gpt-4o-mini
κ 0.571 (85.0%)
Computed directly from the per-item eval files in /json — newest file dated Sep 12, 2026. Sample sizes range from 200 to 200 items per model, so treat small gaps as directional. See Reports for the full narrative write-ups.