Trial report · Open-ended track

Alternative-Route accessibility benchmark: kimi-k3

Model: moonshotai/kimi-k3N = 200Judges: claude-sonnet-4.5 · gpt-4o-miniRubric: 2-of-3Paraphrase allowedGenerated 2026-09-10
Overall accuracy
69.0–76.5%
Judge-dependent · reported 152/200
Weakest category
67.6%
General · 48/71
Primary judge vs. human
κ = 0.342
claude-sonnet-4.5 · 75.0% agree (n=20)
Secondary judge vs. human
κ = 0.634
gpt-4o-mini · 85.0% agree (n=20)

Summary

kimi-k3 was scored on a 200-item open-ended benchmark, judged independently by claude-sonnet-4.5 (primary) and gpt-4o-mini (secondary) against the three-criterion rubric (2-of-3). The two judges agree on 85.5% of items (Cohen's κ = 0.600, “moderate”). Human-label validation gives gpt-4o-mini the higher agreement (κ=0.634 "substantial" vs κ=0.342 "fair" for claude-sonnet-4.5, n=20), so the secondary judge is the judge of record for this model. The reported headline is 76.0% (152/200), with 69.0%–76.5% as the full judge-dependent range.

Accuracy by judge configuration

The same 200 items scored three ways: Primary (claude-sonnet-4.5 alone), Secondary (gpt-4o-mini alone), and AND (both judges must independently mark an item correct — the most conservative reading).

CategorySamplesPrimarySecondaryAND
Aged person5244 (84.6%)46 (88.5%)42 (80.8%)
General7154 (76.1%)48 (67.6%)46 (64.8%)
Physical mobility4330 (69.8%)32 (74.4%)26 (60.5%)
Sensory impairments3425 (73.5%)26 (76.5%)24 (70.6%)
OVERALL200153 (76.5%)152 (76.0%)138 (69.0%)

General shows the largest primary/secondary split of any category — 76.06% vs. 67.61%, a 8.5-point gap. General is the weakest category under the reported configuration, at 67.61% (48/71).

Judge reliability

Every item is scored twice — once by claude-sonnet-4.5 (primary) and once by gpt-4o-mini (secondary), each applying the same rubric independently. Raw agreement is 85.5% and Cohen's κ = 0.600— “moderate” (Landis & Koch: <0 poor, 0–0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect).

Human-label validation (n = 20)

JudgeAgreement w/ humanCohen's κRead
anthropic/claude-sonnet-4.5 (primary)75.0%0.342fair
openai/gpt-4o-mini (secondary)85.0%0.634substantial
Reporting configuration

Human-label validation gives gpt-4o-mini the higher agreement (κ=0.634 "substantial" vs κ=0.342 "fair" for claude-sonnet-4.5, n=20), so the secondary judge is the judge of record for this model.

Failure taxonomy

Every item scored 0 by the secondaryjudge, bucketed into a failure sub-category by keyword match against its ground-truth cues, intent and the model's answer. 48 failures total.

Obstructions/Hazards
10
General/Other
9
Amenities/Accessibility
7
Turns/Intersections
7
Route Planning/Choice
6
Landmarks/Wayfinding
5
Traffic/Signals
4
Main categorySub-categoryFailuresFailed IDs
aged_personAmenities/Accessibility3aged_035, aged_038, aged_049
aged_personGeneral/Other1aged_008
aged_personRoute Planning/Choice1aged_055
aged_personTurns/Intersections1aged_010
generalAmenities/Accessibility2general_014, general_049
generalGeneral/Other5general_040, general_050, general_060, general_063, general_068
generalLandmarks/Wayfinding3general_015, general_030, general_056
generalObstructions/Hazards3general_026, general_037, general_053
generalRoute Planning/Choice4general_004, general_022, general_025, general_027
generalTraffic/Signals3general_019, general_031, sensory_035
generalTurns/Intersections3general_035, general_036, general_058
physical_mobilityAmenities/Accessibility1physical_025
physical_mobilityGeneral/Other1physical_015
physical_mobilityObstructions/Hazards5physical_005, physical_017, physical_027, physical_033, physical_041
physical_mobilityRoute Planning/Choice1physical_036
physical_mobilityTraffic/Signals1physical_031
physical_mobilityTurns/Intersections2physical_028, physical_045
sensory_impairmentsAmenities/Accessibility1sensory_010
sensory_impairmentsGeneral/Other2sensory_016, sensory_032
sensory_impairmentsLandmarks/Wayfinding2sensory_011, sensory_026
sensory_impairmentsObstructions/Hazards2sensory_025, sensory_034
sensory_impairmentsTurns/Intersections1sensory_017

Takeaways

01
Report the secondary judge as final: 76.0% (152/200).

Human-label validation gives gpt-4o-mini the higher agreement (κ=0.634 "substantial" vs κ=0.342 "fair" for claude-sonnet-4.5, n=20), so the secondary judge is the judge of record for this model.

02
General is the category most sensitive to judge choice.

76.06% under the primary judge vs. 67.61% under the secondary judge — a 8.5-point swing, the largest of any category. Worth a manual read of the General disagreements before quoting either extreme.

03
Obstructions/Hazards is the dominant failure mode.

10 of 48 failures (21%), concentrated in: 3 general, 5 physical_mobility, 2 sensory_impairments.

04
claude-sonnet-4.5's per-category numbers should be read as a lower-confidence bound.

At κ = 0.342 against human labels ("fair"), this judge is meaningfully less reliable than the other — treat its category-level figures with more caution, especially where it diverges sharply from the other judge.

Source: Alternative_route_problem_Evaluation.ipynb · open-ended track only · rubric: visual_cue_coverage · intent_match · spatial_correctness (2-of-3).