Trial report · Open-ended track

Alternative-Route accessibility benchmark: minimax-m3

Model: minimax/minimax-m3N = 200Judges: claude-sonnet-4.5 · gpt-4o-miniRubric: 2-of-3Paraphrase allowedGenerated 2026-09-11
Overall accuracy
50.0–58.0%
Judge-dependent · reported 112/200
Weakest category
41.9%
Physical mobility · 18/43
Primary judge vs. human
Score the sample on the Labeling page
Secondary judge vs. human
Score the sample on the Labeling page

Summary

minimax-m3 was scored on a 200-item open-ended benchmark, judged independently by claude-sonnet-4.5 (primary) and gpt-4o-mini (secondary) against the three-criterion rubric (2-of-3). The two judges agree on 86.0% of items (Cohen's κ = 0.715, “substantial”). No human labels have been recorded for this run yet, so the report defaults to the primary judge alone — score the sample on the Labeling page for a validated recommendation. The reported headline is 56.0% (112/200), with 50.0%–58.0% as the full judge-dependent range.

Accuracy by judge configuration

The same 200 items scored three ways: Primary (claude-sonnet-4.5 alone), Secondary (gpt-4o-mini alone), and AND (both judges must independently mark an item correct — the most conservative reading).

CategorySamplesPrimarySecondaryAND
Aged person5228 (53.8%)31 (59.6%)27 (51.9%)
General7143 (60.6%)46 (64.8%)39 (54.9%)
Physical mobility4318 (41.9%)18 (41.9%)15 (34.9%)
Sensory impairments3423 (67.6%)21 (61.8%)19 (55.9%)
OVERALL200112 (56.0%)116 (58.0%)100 (50.0%)

Sensory impairments shows the largest primary/secondary split of any category — 67.65% vs. 61.76%, a 5.9-point gap. Physical mobility is the weakest category under the reported configuration, at 41.86% (18/43).

Judge reliability

Every item is scored twice — once by claude-sonnet-4.5 (primary) and once by gpt-4o-mini (secondary), each applying the same rubric independently. Raw agreement is 86.0% and Cohen's κ = 0.715— “substantial” (Landis & Koch: <0 poor, 0–0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect).

Human-label validation

No human labels yet

No human labels have been recorded for this run yet, so the report defaults to the primary judge alone — score the sample on the Labeling page for a validated recommendation.

Failure taxonomy

Every item scored 0 by the primaryjudge, bucketed into a failure sub-category by keyword match against its ground-truth cues, intent and the model's answer. 88 failures total.

General/Other
19
Route Planning/Choice
15
Landmarks/Wayfinding
13
Obstructions/Hazards
13
Amenities/Accessibility
11
Turns/Intersections
9
Traffic/Signals
8
Main categorySub-categoryFailuresFailed IDs
aged_personAmenities/Accessibility4aged_007, aged_035, sensory_024, aged_056
aged_personGeneral/Other3aged_008, aged_025, aged_047
aged_personLandmarks/Wayfinding4aged_006, general_017, aged_026, aged_029
aged_personObstructions/Hazards3aged_003, aged_021, aged_024
aged_personRoute Planning/Choice5aged_017, aged_042, aged_051, aged_055, general_062
aged_personTraffic/Signals2aged_036, aged_041
aged_personTurns/Intersections3aged_034, aged_043, aged_052
generalAmenities/Accessibility1general_014
generalGeneral/Other13general_016, general_018, general_037, general_040, general_038, general_043, general_050, general_048, general_049, general_057, general_063, general_064, general_068
generalLandmarks/Wayfinding3general_046, general_056, general_067
generalObstructions/Hazards4general_028, general_029, general_053, general_060
generalRoute Planning/Choice2general_013, general_066
generalTraffic/Signals3general_019, general_031, sensory_035
generalTurns/Intersections2general_036, general_039
physical_mobilityAmenities/Accessibility6physical_003, physical_007, physical_021, physical_023, physical_034, physical_038
physical_mobilityGeneral/Other2physical_015, physical_035
physical_mobilityLandmarks/Wayfinding1physical_042
physical_mobilityObstructions/Hazards4physical_002, physical_033, physical_039, physical_041
physical_mobilityRoute Planning/Choice7physical_011, physical_016, physical_029, physical_030, physical_037, physical_036, physical_040
physical_mobilityTraffic/Signals2physical_010, physical_031
physical_mobilityTurns/Intersections3physical_013, physical_019, physical_045
sensory_impairmentsGeneral/Other1sensory_010
sensory_impairmentsLandmarks/Wayfinding5sensory_011, sensory_014, sensory_021, sensory_026, sensory_027
sensory_impairmentsObstructions/Hazards2sensory_034, sensory_037
sensory_impairmentsRoute Planning/Choice1sensory_029
sensory_impairmentsTraffic/Signals1sensory_017
sensory_impairmentsTurns/Intersections1sensory_033

Takeaways

01
Report the primary judge as final: 56.0% (112/200).

No human labels have been recorded for this run yet, so the report defaults to the primary judge alone — score the sample on the Labeling page for a validated recommendation.

02
Sensory impairments is the category most sensitive to judge choice.

67.65% under the primary judge vs. 61.76% under the secondary judge — a 5.9-point swing, the largest of any category. Worth a manual read of the Sensory impairments disagreements before quoting either extreme.

03
General/Other is the dominant failure mode.

19 of 88 failures (22%), concentrated in: 3 aged_person, 13 general, 2 physical_mobility, 1 sensory_impairments.

04
Score the human-label sample before trusting either judge in isolation.

This report defaults to the primary judge alone because no human labels have been recorded for this run yet.

Source: Alternative_route_problem_Evaluation.ipynb · open-ended track only · rubric: visual_cue_coverage · intent_match · spatial_correctness (2-of-3).