Alternative-Route accessibility benchmark: kimi-k2.6_strict
Summary
kimi-k2.6_strict was scored on a 200-item open-ended benchmark, judged independently by claude-sonnet-4.5 (primary) and gpt-4o-mini (secondary) against the three-criterion rubric (strict intent+spatial, visual cue 50% partial credit). The two judges agree on 83.5% of items (Cohen's κ = 0.663, “substantial”). Human-label validation gives claude-sonnet-4.5 the higher agreement (κ=0.634 "substantial" vs κ=0.571 "moderate" for gpt-4o-mini, n=20), so the primary judge is the judge of record for this model. The reported headline is 53.0% (106/200), with 52.5%–68.5% as the full judge-dependent range.
Accuracy by judge configuration
The same 200 items scored three ways: Primary (claude-sonnet-4.5 alone), Secondary (gpt-4o-mini alone), and AND (both judges must independently mark an item correct — the most conservative reading).
| Category | Samples | Primary | Secondary | AND |
|---|---|---|---|---|
| Aged person | 52 | 31 (59.6%) | 41 (78.8%) | 30 (57.7%) |
| General | 71 | 37 (52.1%) | 47 (66.2%) | 37 (52.1%) |
| Physical mobility | 43 | 19 (44.2%) | 27 (62.8%) | 19 (44.2%) |
| Sensory impairments | 34 | 19 (55.9%) | 22 (64.7%) | 19 (55.9%) |
| OVERALL | 200 | 106 (53.0%) | 137 (68.5%) | 105 (52.5%) |
Aged person shows the largest primary/secondary split of any category — 59.62% vs. 78.85%, a 19.2-point gap. Physical mobility is the weakest category under the reported configuration, at 44.19% (19/43).
Judge reliability
Every item is scored twice — once by claude-sonnet-4.5 (primary) and once by gpt-4o-mini (secondary), each applying the same rubric independently. Raw agreement is 83.5% and Cohen's κ = 0.663— “substantial” (Landis & Koch: <0 poor, 0–0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect).
Human-label validation (n = 20)
| Judge | Agreement w/ human | Cohen's κ | Read |
|---|---|---|---|
| anthropic/claude-sonnet-4.5 (primary) | 85.0% | 0.634 | substantial |
| openai/gpt-4o-mini (secondary) | 85.0% | 0.571 | moderate |
Human-label validation gives claude-sonnet-4.5 the higher agreement (κ=0.634 "substantial" vs κ=0.571 "moderate" for gpt-4o-mini, n=20), so the primary judge is the judge of record for this model.
Failure taxonomy
Every item scored 0 by the primaryjudge, bucketed into a failure sub-category by keyword match against its ground-truth cues, intent and the model's answer. 94 failures total.
| Main category | Sub-category | Failures | Failed IDs |
|---|---|---|---|
| aged_person | Amenities/Accessibility | 4 | aged_007, aged_035, aged_037, aged_038 |
| aged_person | General/Other | 3 | aged_008, aged_015, aged_032 |
| aged_person | Landmarks/Wayfinding | 2 | aged_006, general_017 |
| aged_person | Obstructions/Hazards | 6 | aged_003, aged_014, aged_013, aged_021, aged_023, aged_022 |
| aged_person | Route Planning/Choice | 4 | aged_012, aged_042, general_062, aged_055 |
| aged_person | Traffic/Signals | 1 | aged_041 |
| aged_person | Turns/Intersections | 1 | aged_005 |
| general | Amenities/Accessibility | 1 | general_014 |
| general | General/Other | 11 | general_016, general_018, general_006, general_027, general_032, general_033, general_040, general_050, general_049, general_063, general_065 |
| general | Landmarks/Wayfinding | 4 | general_015, general_024, general_030, general_067 |
| general | Obstructions/Hazards | 7 | general_029, general_028, general_026, general_037, general_042, general_053, general_060 |
| general | Route Planning/Choice | 5 | general_013, general_043, general_044, general_048, general_066 |
| general | Traffic/Signals | 4 | general_019, general_031, general_045, sensory_035 |
| general | Turns/Intersections | 2 | general_035, general_036 |
| physical_mobility | Amenities/Accessibility | 5 | physical_003, physical_021, physical_018, physical_024, physical_026 |
| physical_mobility | General/Other | 1 | physical_040 |
| physical_mobility | Landmarks/Wayfinding | 1 | physical_042 |
| physical_mobility | Obstructions/Hazards | 7 | physical_014, physical_005, physical_017, physical_027, physical_033, physical_039, physical_041 |
| physical_mobility | Route Planning/Choice | 1 | physical_036 |
| physical_mobility | Traffic/Signals | 4 | physical_009, physical_010, physical_012, physical_016 |
| physical_mobility | Turns/Intersections | 5 | physical_001, physical_013, physical_028, physical_031, physical_045 |
| sensory_impairments | Amenities/Accessibility | 1 | sensory_031 |
| sensory_impairments | General/Other | 4 | sensory_006, sensory_016, sensory_032, aged_054 |
| sensory_impairments | Landmarks/Wayfinding | 2 | sensory_014, sensory_026 |
| sensory_impairments | Obstructions/Hazards | 4 | sensory_009, sensory_013, sensory_034, sensory_037 |
| sensory_impairments | Route Planning/Choice | 2 | sensory_022, sensory_029 |
| sensory_impairments | Turns/Intersections | 2 | sensory_007, sensory_023 |
Takeaways
Human-label validation gives claude-sonnet-4.5 the higher agreement (κ=0.634 "substantial" vs κ=0.571 "moderate" for gpt-4o-mini, n=20), so the primary judge is the judge of record for this model.
59.62% under the primary judge vs. 78.85% under the secondary judge — a 19.2-point swing, the largest of any category. Worth a manual read of the Aged person disagreements before quoting either extreme.
24 of 94 failures (26%), concentrated in: 6 aged_person, 7 general, 7 physical_mobility, 4 sensory_impairments.
At κ = 0.571 against human labels ("moderate"), this judge is meaningfully less reliable than the other — treat its category-level figures with more caution, especially where it diverges sharply from the other judge.