Alternative-Route accessibility benchmark: minimax-m3
Summary
minimax-m3 was scored on a 200-item open-ended benchmark, judged independently by claude-sonnet-4.5 (primary) and gpt-4o-mini (secondary) against the three-criterion rubric (2-of-3). The two judges agree on 86.0% of items (Cohen's κ = 0.715, “substantial”). No human labels have been recorded for this run yet, so the report defaults to the primary judge alone — score the sample on the Labeling page for a validated recommendation. The reported headline is 56.0% (112/200), with 50.0%–58.0% as the full judge-dependent range.
Accuracy by judge configuration
The same 200 items scored three ways: Primary (claude-sonnet-4.5 alone), Secondary (gpt-4o-mini alone), and AND (both judges must independently mark an item correct — the most conservative reading).
| Category | Samples | Primary | Secondary | AND |
|---|---|---|---|---|
| Aged person | 52 | 28 (53.8%) | 31 (59.6%) | 27 (51.9%) |
| General | 71 | 43 (60.6%) | 46 (64.8%) | 39 (54.9%) |
| Physical mobility | 43 | 18 (41.9%) | 18 (41.9%) | 15 (34.9%) |
| Sensory impairments | 34 | 23 (67.6%) | 21 (61.8%) | 19 (55.9%) |
| OVERALL | 200 | 112 (56.0%) | 116 (58.0%) | 100 (50.0%) |
Sensory impairments shows the largest primary/secondary split of any category — 67.65% vs. 61.76%, a 5.9-point gap. Physical mobility is the weakest category under the reported configuration, at 41.86% (18/43).
Judge reliability
Every item is scored twice — once by claude-sonnet-4.5 (primary) and once by gpt-4o-mini (secondary), each applying the same rubric independently. Raw agreement is 86.0% and Cohen's κ = 0.715— “substantial” (Landis & Koch: <0 poor, 0–0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect).
Human-label validation
No human labels have been recorded for this run yet, so the report defaults to the primary judge alone — score the sample on the Labeling page for a validated recommendation.
Failure taxonomy
Every item scored 0 by the primaryjudge, bucketed into a failure sub-category by keyword match against its ground-truth cues, intent and the model's answer. 88 failures total.
| Main category | Sub-category | Failures | Failed IDs |
|---|---|---|---|
| aged_person | Amenities/Accessibility | 4 | aged_007, aged_035, sensory_024, aged_056 |
| aged_person | General/Other | 3 | aged_008, aged_025, aged_047 |
| aged_person | Landmarks/Wayfinding | 4 | aged_006, general_017, aged_026, aged_029 |
| aged_person | Obstructions/Hazards | 3 | aged_003, aged_021, aged_024 |
| aged_person | Route Planning/Choice | 5 | aged_017, aged_042, aged_051, aged_055, general_062 |
| aged_person | Traffic/Signals | 2 | aged_036, aged_041 |
| aged_person | Turns/Intersections | 3 | aged_034, aged_043, aged_052 |
| general | Amenities/Accessibility | 1 | general_014 |
| general | General/Other | 13 | general_016, general_018, general_037, general_040, general_038, general_043, general_050, general_048, general_049, general_057, general_063, general_064, general_068 |
| general | Landmarks/Wayfinding | 3 | general_046, general_056, general_067 |
| general | Obstructions/Hazards | 4 | general_028, general_029, general_053, general_060 |
| general | Route Planning/Choice | 2 | general_013, general_066 |
| general | Traffic/Signals | 3 | general_019, general_031, sensory_035 |
| general | Turns/Intersections | 2 | general_036, general_039 |
| physical_mobility | Amenities/Accessibility | 6 | physical_003, physical_007, physical_021, physical_023, physical_034, physical_038 |
| physical_mobility | General/Other | 2 | physical_015, physical_035 |
| physical_mobility | Landmarks/Wayfinding | 1 | physical_042 |
| physical_mobility | Obstructions/Hazards | 4 | physical_002, physical_033, physical_039, physical_041 |
| physical_mobility | Route Planning/Choice | 7 | physical_011, physical_016, physical_029, physical_030, physical_037, physical_036, physical_040 |
| physical_mobility | Traffic/Signals | 2 | physical_010, physical_031 |
| physical_mobility | Turns/Intersections | 3 | physical_013, physical_019, physical_045 |
| sensory_impairments | General/Other | 1 | sensory_010 |
| sensory_impairments | Landmarks/Wayfinding | 5 | sensory_011, sensory_014, sensory_021, sensory_026, sensory_027 |
| sensory_impairments | Obstructions/Hazards | 2 | sensory_034, sensory_037 |
| sensory_impairments | Route Planning/Choice | 1 | sensory_029 |
| sensory_impairments | Traffic/Signals | 1 | sensory_017 |
| sensory_impairments | Turns/Intersections | 1 | sensory_033 |
Takeaways
No human labels have been recorded for this run yet, so the report defaults to the primary judge alone — score the sample on the Labeling page for a validated recommendation.
67.65% under the primary judge vs. 61.76% under the secondary judge — a 5.9-point swing, the largest of any category. Worth a manual read of the Sensory impairments disagreements before quoting either extreme.
19 of 88 failures (22%), concentrated in: 3 aged_person, 13 general, 2 physical_mobility, 1 sensory_impairments.
This report defaults to the primary judge alone because no human labels have been recorded for this run yet.