Trial report · Open-ended track

Alternative-Route accessibility benchmark: kimi-k2.6

Model: moonshotai/kimi-k2.6N = 200Judges: claude-sonnet-4.5 · gpt-4o-miniRubric: 2-of-3Paraphrase allowedGenerated 2026-09-11
Overall accuracy
65.5–76.5%
Judge-dependent · reported 137/200
Weakest category
58.1%
Physical mobility · 25/43
Primary judge vs. human
Score the sample on the Labeling page
Secondary judge vs. human
Score the sample on the Labeling page

Summary

kimi-k2.6 was scored on a 200-item open-ended benchmark, judged independently by claude-sonnet-4.5 (primary) and gpt-4o-mini (secondary) against the three-criterion rubric (2-of-3). The two judges agree on 86.0% of items (Cohen's κ = 0.652, “substantial”). No human labels have been recorded for this run yet, so the report defaults to the primary judge alone — score the sample on the Labeling page for a validated recommendation. The reported headline is 68.5% (137/200), with 65.5%–76.5% as the full judge-dependent range.

Accuracy by judge configuration

The same 200 items scored three ways: Primary (claude-sonnet-4.5 alone), Secondary (gpt-4o-mini alone), and AND (both judges must independently mark an item correct — the most conservative reading).

CategorySamplesPrimarySecondaryAND
Aged person5241 (78.8%)46 (88.5%)41 (78.8%)
General7145 (63.4%)50 (70.4%)42 (59.2%)
Physical mobility4325 (58.1%)32 (74.4%)25 (58.1%)
Sensory impairments3426 (76.5%)25 (73.5%)23 (67.6%)
OVERALL200137 (68.5%)153 (76.5%)131 (65.5%)

Physical mobility shows the largest primary/secondary split of any category — 58.14% vs. 74.42%, a 16.3-point gap. Physical mobility is the weakest category under the reported configuration, at 58.14% (25/43).

Judge reliability

Every item is scored twice — once by claude-sonnet-4.5 (primary) and once by gpt-4o-mini (secondary), each applying the same rubric independently. Raw agreement is 86.0% and Cohen's κ = 0.652— “substantial” (Landis & Koch: <0 poor, 0–0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect).

Human-label validation

No human labels yet

No human labels have been recorded for this run yet, so the report defaults to the primary judge alone — score the sample on the Labeling page for a validated recommendation.

Failure taxonomy

Every item scored 0 by the primaryjudge, bucketed into a failure sub-category by keyword match against its ground-truth cues, intent and the model's answer. 63 failures total.

General/Other
14
Amenities/Accessibility
10
Obstructions/Hazards
10
Landmarks/Wayfinding
9
Traffic/Signals
8
Turns/Intersections
7
Route Planning/Choice
5
Main categorySub-categoryFailuresFailed IDs
aged_personAmenities/Accessibility3aged_035, aged_037, aged_049
aged_personGeneral/Other2aged_032, general_062
aged_personLandmarks/Wayfinding3aged_006, general_017, aged_029
aged_personObstructions/Hazards1aged_021
aged_personRoute Planning/Choice1aged_042
aged_personTurns/Intersections1aged_052
generalGeneral/Other11general_016, general_027, general_032, general_033, general_040, general_043, general_049, general_050, general_065, general_063, general_064
generalLandmarks/Wayfinding3general_024, general_035, general_046
generalObstructions/Hazards3general_028, general_037, general_053
generalRoute Planning/Choice4general_013, general_023, general_048, general_066
generalTraffic/Signals4general_019, general_031, general_045, sensory_035
generalTurns/Intersections1general_036
physical_mobilityAmenities/Accessibility6physical_021, physical_018, physical_023, physical_026, physical_029, physical_034
physical_mobilityLandmarks/Wayfinding1physical_042
physical_mobilityObstructions/Hazards5physical_027, physical_033, physical_036, physical_039, physical_041
physical_mobilityTraffic/Signals3physical_009, physical_016, physical_040
physical_mobilityTurns/Intersections3physical_013, physical_020, physical_045
sensory_impairmentsAmenities/Accessibility1sensory_031
sensory_impairmentsGeneral/Other1aged_054
sensory_impairmentsLandmarks/Wayfinding2sensory_011, sensory_030
sensory_impairmentsObstructions/Hazards1sensory_009
sensory_impairmentsTraffic/Signals1sensory_037
sensory_impairmentsTurns/Intersections2sensory_023, sensory_033

Takeaways

01
Report the primary judge as final: 68.5% (137/200).

No human labels have been recorded for this run yet, so the report defaults to the primary judge alone — score the sample on the Labeling page for a validated recommendation.

02
Physical mobility is the category most sensitive to judge choice.

58.14% under the primary judge vs. 74.42% under the secondary judge — a 16.3-point swing, the largest of any category. Worth a manual read of the Physical mobility disagreements before quoting either extreme.

03
General/Other is the dominant failure mode.

14 of 63 failures (22%), concentrated in: 2 aged_person, 11 general, 1 sensory_impairments.

04
Score the human-label sample before trusting either judge in isolation.

This report defaults to the primary judge alone because no human labels have been recorded for this run yet.

Source: Alternative_route_problem_Evaluation.ipynb · open-ended track only · rubric: visual_cue_coverage · intent_match · spatial_correctness (2-of-3).