Alternative-Route accessibility benchmark: kimi-k3
Summary
kimi-k3 was scored on a 200-item open-ended benchmark, judged independently by claude-sonnet-4.5 (primary) and gpt-4o-mini (secondary) against the three-criterion rubric (2-of-3). The two judges agree on 85.5% of items (Cohen's κ = 0.600, “moderate”). Human-label validation gives gpt-4o-mini the higher agreement (κ=0.634 "substantial" vs κ=0.342 "fair" for claude-sonnet-4.5, n=20), so the secondary judge is the judge of record for this model. The reported headline is 76.0% (152/200), with 69.0%–76.5% as the full judge-dependent range.
Accuracy by judge configuration
The same 200 items scored three ways: Primary (claude-sonnet-4.5 alone), Secondary (gpt-4o-mini alone), and AND (both judges must independently mark an item correct — the most conservative reading).
| Category | Samples | Primary | Secondary | AND |
|---|---|---|---|---|
| Aged person | 52 | 44 (84.6%) | 46 (88.5%) | 42 (80.8%) |
| General | 71 | 54 (76.1%) | 48 (67.6%) | 46 (64.8%) |
| Physical mobility | 43 | 30 (69.8%) | 32 (74.4%) | 26 (60.5%) |
| Sensory impairments | 34 | 25 (73.5%) | 26 (76.5%) | 24 (70.6%) |
| OVERALL | 200 | 153 (76.5%) | 152 (76.0%) | 138 (69.0%) |
General shows the largest primary/secondary split of any category — 76.06% vs. 67.61%, a 8.5-point gap. General is the weakest category under the reported configuration, at 67.61% (48/71).
Judge reliability
Every item is scored twice — once by claude-sonnet-4.5 (primary) and once by gpt-4o-mini (secondary), each applying the same rubric independently. Raw agreement is 85.5% and Cohen's κ = 0.600— “moderate” (Landis & Koch: <0 poor, 0–0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect).
Human-label validation (n = 20)
| Judge | Agreement w/ human | Cohen's κ | Read |
|---|---|---|---|
| anthropic/claude-sonnet-4.5 (primary) | 75.0% | 0.342 | fair |
| openai/gpt-4o-mini (secondary) | 85.0% | 0.634 | substantial |
Human-label validation gives gpt-4o-mini the higher agreement (κ=0.634 "substantial" vs κ=0.342 "fair" for claude-sonnet-4.5, n=20), so the secondary judge is the judge of record for this model.
Failure taxonomy
Every item scored 0 by the secondaryjudge, bucketed into a failure sub-category by keyword match against its ground-truth cues, intent and the model's answer. 48 failures total.
| Main category | Sub-category | Failures | Failed IDs |
|---|---|---|---|
| aged_person | Amenities/Accessibility | 3 | aged_035, aged_038, aged_049 |
| aged_person | General/Other | 1 | aged_008 |
| aged_person | Route Planning/Choice | 1 | aged_055 |
| aged_person | Turns/Intersections | 1 | aged_010 |
| general | Amenities/Accessibility | 2 | general_014, general_049 |
| general | General/Other | 5 | general_040, general_050, general_060, general_063, general_068 |
| general | Landmarks/Wayfinding | 3 | general_015, general_030, general_056 |
| general | Obstructions/Hazards | 3 | general_026, general_037, general_053 |
| general | Route Planning/Choice | 4 | general_004, general_022, general_025, general_027 |
| general | Traffic/Signals | 3 | general_019, general_031, sensory_035 |
| general | Turns/Intersections | 3 | general_035, general_036, general_058 |
| physical_mobility | Amenities/Accessibility | 1 | physical_025 |
| physical_mobility | General/Other | 1 | physical_015 |
| physical_mobility | Obstructions/Hazards | 5 | physical_005, physical_017, physical_027, physical_033, physical_041 |
| physical_mobility | Route Planning/Choice | 1 | physical_036 |
| physical_mobility | Traffic/Signals | 1 | physical_031 |
| physical_mobility | Turns/Intersections | 2 | physical_028, physical_045 |
| sensory_impairments | Amenities/Accessibility | 1 | sensory_010 |
| sensory_impairments | General/Other | 2 | sensory_016, sensory_032 |
| sensory_impairments | Landmarks/Wayfinding | 2 | sensory_011, sensory_026 |
| sensory_impairments | Obstructions/Hazards | 2 | sensory_025, sensory_034 |
| sensory_impairments | Turns/Intersections | 1 | sensory_017 |
Takeaways
Human-label validation gives gpt-4o-mini the higher agreement (κ=0.634 "substantial" vs κ=0.342 "fair" for claude-sonnet-4.5, n=20), so the secondary judge is the judge of record for this model.
76.06% under the primary judge vs. 67.61% under the secondary judge — a 8.5-point swing, the largest of any category. Worth a manual read of the General disagreements before quoting either extreme.
10 of 48 failures (21%), concentrated in: 3 general, 5 physical_mobility, 2 sensory_impairments.
At κ = 0.342 against human labels ("fair"), this judge is meaningfully less reliable than the other — treat its category-level figures with more caution, especially where it diverges sharply from the other judge.