Trial report · Open-ended track

Alternative-Route accessibility benchmark: gemini-3.1-flash-image

Model: google/gemini-3.1-flash-imageN = 200Judges: claude-sonnet-4.5 · gpt-4o-miniRubric: 2-of-3Paraphrase allowedGenerated 2026-09-10
Overall accuracy
68.0–74.5%
Judge-dependent · reported 145/200
Weakest category
67.4%
Physical mobility · 29/43
Primary judge vs. human
κ = 0.857
claude-sonnet-4.5 · 95.0% agree (n=20)
Secondary judge vs. human
κ = 0.483
gpt-4o-mini · 85.0% agree (n=20)

Summary

gemini-3.1-flash-image was scored on a 200-item open-ended benchmark, judged independently by claude-sonnet-4.5 (primary) and gpt-4o-mini (secondary) against the three-criterion rubric (2-of-3). The two judges agree on 89.0% of items (Cohen's κ = 0.718, “substantial”). Human-label validation gives claude-sonnet-4.5 the higher agreement (κ=0.857 "almost perfect" vs κ=0.483 "moderate" for gpt-4o-mini, n=20), so the primary judge is the judge of record for this model. The reported headline is 72.5% (145/200), with 68.0%–74.5% as the full judge-dependent range.

Accuracy by judge configuration

The same 200 items scored three ways: Primary (claude-sonnet-4.5 alone), Secondary (gpt-4o-mini alone), and AND (both judges must independently mark an item correct — the most conservative reading).

CategorySamplesPrimarySecondaryAND
Aged person5243 (82.7%)46 (88.5%)43 (82.7%)
General7150 (70.4%)48 (67.6%)45 (63.4%)
Physical mobility4329 (67.4%)33 (76.7%)29 (67.4%)
Sensory impairments3423 (67.6%)22 (64.7%)19 (55.9%)
OVERALL200145 (72.5%)149 (74.5%)136 (68.0%)

Physical mobility shows the largest primary/secondary split of any category — 67.44% vs. 76.74%, a 9.3-point gap. Physical mobility is the weakest category under the reported configuration, at 67.44% (29/43).

Judge reliability

Every item is scored twice — once by claude-sonnet-4.5 (primary) and once by gpt-4o-mini (secondary), each applying the same rubric independently. Raw agreement is 89.0% and Cohen's κ = 0.718— “substantial” (Landis & Koch: <0 poor, 0–0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect).

Human-label validation (n = 20)

JudgeAgreement w/ humanCohen's κRead
anthropic/claude-sonnet-4.5 (primary)95.0%0.857almost perfect
openai/gpt-4o-mini (secondary)85.0%0.483moderate
Reporting configuration

Human-label validation gives claude-sonnet-4.5 the higher agreement (κ=0.857 "almost perfect" vs κ=0.483 "moderate" for gpt-4o-mini, n=20), so the primary judge is the judge of record for this model.

Failure taxonomy

Every item scored 0 by the primaryjudge, bucketed into a failure sub-category by keyword match against its ground-truth cues, intent and the model's answer. 55 failures total.

Route Planning/Choice
11
General/Other
10
Landmarks/Wayfinding
9
Amenities/Accessibility
7
Turns/Intersections
6
Obstructions/Hazards
6
Traffic/Signals
6
Main categorySub-categoryFailuresFailed IDs
aged_personAmenities/Accessibility1aged_035
aged_personGeneral/Other2aged_008, general_062
aged_personLandmarks/Wayfinding1aged_029
aged_personRoute Planning/Choice3aged_012, aged_027, aged_042
aged_personTurns/Intersections2aged_034, aged_052
generalGeneral/Other5general_010, general_040, general_050, general_063, general_065
generalLandmarks/Wayfinding2general_022, general_034
generalObstructions/Hazards3general_032, general_053, general_060
generalRoute Planning/Choice6general_013, general_023, general_037, general_044, general_048, general_068
generalTraffic/Signals3general_019, general_031, general_045
generalTurns/Intersections2general_035, general_058
physical_mobilityAmenities/Accessibility4physical_018, physical_026, physical_033, physical_034
physical_mobilityGeneral/Other1physical_035
physical_mobilityLandmarks/Wayfinding1physical_042
physical_mobilityObstructions/Hazards2physical_027, physical_039
physical_mobilityRoute Planning/Choice2physical_029, physical_036
physical_mobilityTraffic/Signals2physical_010, physical_031
physical_mobilityTurns/Intersections2physical_013, physical_019
sensory_impairmentsAmenities/Accessibility2sensory_010, sensory_031
sensory_impairmentsGeneral/Other2sensory_006, sensory_032
sensory_impairmentsLandmarks/Wayfinding5sensory_011, sensory_021, sensory_027, sensory_026, sensory_030
sensory_impairmentsObstructions/Hazards1sensory_037
sensory_impairmentsTraffic/Signals1sensory_005

Takeaways

01
Report the primary judge as final: 72.5% (145/200).

Human-label validation gives claude-sonnet-4.5 the higher agreement (κ=0.857 "almost perfect" vs κ=0.483 "moderate" for gpt-4o-mini, n=20), so the primary judge is the judge of record for this model.

02
Physical mobility is the category most sensitive to judge choice.

67.44% under the primary judge vs. 76.74% under the secondary judge — a 9.3-point swing, the largest of any category. Worth a manual read of the Physical mobility disagreements before quoting either extreme.

03
Route Planning/Choice is the dominant failure mode.

11 of 55 failures (20%), concentrated in: 3 aged_person, 6 general, 2 physical_mobility.

04
gpt-4o-mini's per-category numbers should be read as a lower-confidence bound.

At κ = 0.483 against human labels ("moderate"), this judge is meaningfully less reliable than the other — treat its category-level figures with more caution, especially where it diverges sharply from the other judge.

Source: Alternative_route_problem_Evaluation.ipynb · open-ended track only · rubric: visual_cue_coverage · intent_match · spatial_correctness (2-of-3).