Trial report · Open-ended track

Alternative-Route accessibility benchmark: kimi-k2.6_strict

Model: moonshotai/kimi-k2.6_strictN = 200Judges: claude-sonnet-4.5 · gpt-4o-miniRubric: strict intent+spatial, visual cue 50% partial creditParaphrase allowedGenerated 2026-09-11
Overall accuracy
52.5–68.5%
Judge-dependent · reported 106/200
Weakest category
44.2%
Physical mobility · 19/43
Primary judge vs. human
κ = 0.634
claude-sonnet-4.5 · 85.0% agree (n=20)
Secondary judge vs. human
κ = 0.571
gpt-4o-mini · 85.0% agree (n=20)

Summary

kimi-k2.6_strict was scored on a 200-item open-ended benchmark, judged independently by claude-sonnet-4.5 (primary) and gpt-4o-mini (secondary) against the three-criterion rubric (strict intent+spatial, visual cue 50% partial credit). The two judges agree on 83.5% of items (Cohen's κ = 0.663, “substantial”). Human-label validation gives claude-sonnet-4.5 the higher agreement (κ=0.634 "substantial" vs κ=0.571 "moderate" for gpt-4o-mini, n=20), so the primary judge is the judge of record for this model. The reported headline is 53.0% (106/200), with 52.5%–68.5% as the full judge-dependent range.

Accuracy by judge configuration

The same 200 items scored three ways: Primary (claude-sonnet-4.5 alone), Secondary (gpt-4o-mini alone), and AND (both judges must independently mark an item correct — the most conservative reading).

CategorySamplesPrimarySecondaryAND
Aged person5231 (59.6%)41 (78.8%)30 (57.7%)
General7137 (52.1%)47 (66.2%)37 (52.1%)
Physical mobility4319 (44.2%)27 (62.8%)19 (44.2%)
Sensory impairments3419 (55.9%)22 (64.7%)19 (55.9%)
OVERALL200106 (53.0%)137 (68.5%)105 (52.5%)

Aged person shows the largest primary/secondary split of any category — 59.62% vs. 78.85%, a 19.2-point gap. Physical mobility is the weakest category under the reported configuration, at 44.19% (19/43).

Judge reliability

Every item is scored twice — once by claude-sonnet-4.5 (primary) and once by gpt-4o-mini (secondary), each applying the same rubric independently. Raw agreement is 83.5% and Cohen's κ = 0.663— “substantial” (Landis & Koch: <0 poor, 0–0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect).

Human-label validation (n = 20)

JudgeAgreement w/ humanCohen's κRead
anthropic/claude-sonnet-4.5 (primary)85.0%0.634substantial
openai/gpt-4o-mini (secondary)85.0%0.571moderate
Reporting configuration

Human-label validation gives claude-sonnet-4.5 the higher agreement (κ=0.634 "substantial" vs κ=0.571 "moderate" for gpt-4o-mini, n=20), so the primary judge is the judge of record for this model.

Failure taxonomy

Every item scored 0 by the primaryjudge, bucketed into a failure sub-category by keyword match against its ground-truth cues, intent and the model's answer. 94 failures total.

Obstructions/Hazards
24
General/Other
19
Route Planning/Choice
12
Amenities/Accessibility
11
Turns/Intersections
10
Landmarks/Wayfinding
9
Traffic/Signals
9
Main categorySub-categoryFailuresFailed IDs
aged_personAmenities/Accessibility4aged_007, aged_035, aged_037, aged_038
aged_personGeneral/Other3aged_008, aged_015, aged_032
aged_personLandmarks/Wayfinding2aged_006, general_017
aged_personObstructions/Hazards6aged_003, aged_014, aged_013, aged_021, aged_023, aged_022
aged_personRoute Planning/Choice4aged_012, aged_042, general_062, aged_055
aged_personTraffic/Signals1aged_041
aged_personTurns/Intersections1aged_005
generalAmenities/Accessibility1general_014
generalGeneral/Other11general_016, general_018, general_006, general_027, general_032, general_033, general_040, general_050, general_049, general_063, general_065
generalLandmarks/Wayfinding4general_015, general_024, general_030, general_067
generalObstructions/Hazards7general_029, general_028, general_026, general_037, general_042, general_053, general_060
generalRoute Planning/Choice5general_013, general_043, general_044, general_048, general_066
generalTraffic/Signals4general_019, general_031, general_045, sensory_035
generalTurns/Intersections2general_035, general_036
physical_mobilityAmenities/Accessibility5physical_003, physical_021, physical_018, physical_024, physical_026
physical_mobilityGeneral/Other1physical_040
physical_mobilityLandmarks/Wayfinding1physical_042
physical_mobilityObstructions/Hazards7physical_014, physical_005, physical_017, physical_027, physical_033, physical_039, physical_041
physical_mobilityRoute Planning/Choice1physical_036
physical_mobilityTraffic/Signals4physical_009, physical_010, physical_012, physical_016
physical_mobilityTurns/Intersections5physical_001, physical_013, physical_028, physical_031, physical_045
sensory_impairmentsAmenities/Accessibility1sensory_031
sensory_impairmentsGeneral/Other4sensory_006, sensory_016, sensory_032, aged_054
sensory_impairmentsLandmarks/Wayfinding2sensory_014, sensory_026
sensory_impairmentsObstructions/Hazards4sensory_009, sensory_013, sensory_034, sensory_037
sensory_impairmentsRoute Planning/Choice2sensory_022, sensory_029
sensory_impairmentsTurns/Intersections2sensory_007, sensory_023

Takeaways

01
Report the primary judge as final: 53.0% (106/200).

Human-label validation gives claude-sonnet-4.5 the higher agreement (κ=0.634 "substantial" vs κ=0.571 "moderate" for gpt-4o-mini, n=20), so the primary judge is the judge of record for this model.

02
Aged person is the category most sensitive to judge choice.

59.62% under the primary judge vs. 78.85% under the secondary judge — a 19.2-point swing, the largest of any category. Worth a manual read of the Aged person disagreements before quoting either extreme.

03
Obstructions/Hazards is the dominant failure mode.

24 of 94 failures (26%), concentrated in: 6 aged_person, 7 general, 7 physical_mobility, 4 sensory_impairments.

04
gpt-4o-mini's per-category numbers should be read as a lower-confidence bound.

At κ = 0.571 against human labels ("moderate"), this judge is meaningfully less reliable than the other — treat its category-level figures with more caution, especially where it diverges sharply from the other judge.

Source: Alternative_route_problem_Evaluation.ipynb · open-ended track only · rubric: visual_cue_coverage · intent_match · spatial_correctness (strict intent+spatial, visual cue 50% partial credit).