Trial report · Open-ended track

Alternative-Route accessibility benchmark: gemini-3.1-flash-image_strict

Model: google/gemini-3.1-flash-image_strictN = 200Judges: claude-sonnet-4.5 · gpt-4o-miniRubric: strict intent+spatial, visual cue 50% partial creditParaphrase allowedGenerated 2026-09-12
Overall accuracy
54.0–70.0%
Judge-dependent · reported 140/200
Weakest category
64.8%
General · 46/71
Primary judge vs. human
κ = 0.419
claude-sonnet-4.5 · 75.0% agree (n=20)
Secondary judge vs. human
κ = 0.692
gpt-4o-mini · 90.0% agree (n=20)

Summary

gemini-3.1-flash-image_strict was scored on a 200-item open-ended benchmark, judged independently by claude-sonnet-4.5 (primary) and gpt-4o-mini (secondary) against the three-criterion rubric (strict intent+spatial, visual cue 50% partial credit). The two judges agree on 81.5% of items (Cohen's κ = 0.610, “substantial”). Human-label validation gives gpt-4o-mini the higher agreement (κ=0.692 "substantial" vs κ=0.419 "moderate" for claude-sonnet-4.5, n=20), so the secondary judge is the judge of record for this model. The reported headline is 70.0% (140/200), with 54.0%–70.0% as the full judge-dependent range.

Accuracy by judge configuration

The same 200 items scored three ways: Primary (claude-sonnet-4.5 alone), Secondary (gpt-4o-mini alone), and AND (both judges must independently mark an item correct — the most conservative reading).

CategorySamplesPrimarySecondaryAND
Aged person5234 (65.4%)38 (73.1%)33 (63.5%)
General7137 (52.1%)46 (64.8%)34 (47.9%)
Physical mobility4323 (53.5%)33 (76.7%)22 (51.2%)
Sensory impairments3419 (55.9%)23 (67.6%)19 (55.9%)
OVERALL200113 (56.5%)140 (70.0%)108 (54.0%)

Physical mobility shows the largest primary/secondary split of any category — 53.49% vs. 76.74%, a 23.3-point gap. General is the weakest category under the reported configuration, at 64.79% (46/71).

Judge reliability

Every item is scored twice — once by claude-sonnet-4.5 (primary) and once by gpt-4o-mini (secondary), each applying the same rubric independently. Raw agreement is 81.5% and Cohen's κ = 0.610— “substantial” (Landis & Koch: <0 poor, 0–0.2 slight, 0.2–0.4 fair, 0.4–0.6 moderate, 0.6–0.8 substantial, 0.8–1.0 almost perfect).

Human-label validation (n = 20)

JudgeAgreement w/ humanCohen's κRead
anthropic/claude-sonnet-4.5 (primary)75.0%0.419moderate
openai/gpt-4o-mini (secondary)90.0%0.692substantial
Reporting configuration

Human-label validation gives gpt-4o-mini the higher agreement (κ=0.692 "substantial" vs κ=0.419 "moderate" for claude-sonnet-4.5, n=20), so the secondary judge is the judge of record for this model.

Failure taxonomy

Every item scored 0 by the secondaryjudge, bucketed into a failure sub-category by keyword match against its ground-truth cues, intent and the model's answer. 60 failures total.

General/Other
12
Route Planning/Choice
10
Amenities/Accessibility
9
Landmarks/Wayfinding
9
Obstructions/Hazards
8
Turns/Intersections
8
Traffic/Signals
4
Main categorySub-categoryFailuresFailed IDs
aged_personAmenities/Accessibility2aged_035, aged_038
aged_personGeneral/Other2aged_008, aged_032
aged_personLandmarks/Wayfinding2aged_006, aged_029
aged_personObstructions/Hazards1aged_003
aged_personRoute Planning/Choice5aged_012, aged_017, aged_027, aged_042, aged_055
aged_personTurns/Intersections2aged_034, aged_052
generalAmenities/Accessibility2general_014, general_016
generalGeneral/Other7general_007, general_010, general_022, general_040, general_060, general_065, general_068
generalLandmarks/Wayfinding2general_015, general_056
generalObstructions/Hazards4general_026, general_032, general_042, general_053
generalRoute Planning/Choice4general_012, general_023, general_044, general_048
generalTraffic/Signals2general_019, general_045
generalTurns/Intersections4general_036, general_035, general_041, general_058
physical_mobilityAmenities/Accessibility3physical_026, physical_033, physical_034
physical_mobilityLandmarks/Wayfinding1physical_042
physical_mobilityObstructions/Hazards3physical_005, physical_014, physical_027
physical_mobilityTraffic/Signals2physical_010, physical_031
physical_mobilityTurns/Intersections1physical_013
sensory_impairmentsAmenities/Accessibility2sensory_031, sensory_032
sensory_impairmentsGeneral/Other3sensory_006, sensory_010, aged_054
sensory_impairmentsLandmarks/Wayfinding4sensory_011, sensory_014, sensory_021, sensory_030
sensory_impairmentsRoute Planning/Choice1sensory_022
sensory_impairmentsTurns/Intersections1sensory_023

Takeaways

01
Report the secondary judge as final: 70.0% (140/200).

Human-label validation gives gpt-4o-mini the higher agreement (κ=0.692 "substantial" vs κ=0.419 "moderate" for claude-sonnet-4.5, n=20), so the secondary judge is the judge of record for this model.

02
Physical mobility is the category most sensitive to judge choice.

53.49% under the primary judge vs. 76.74% under the secondary judge — a 23.3-point swing, the largest of any category. Worth a manual read of the Physical mobility disagreements before quoting either extreme.

03
General/Other is the dominant failure mode.

12 of 60 failures (20%), concentrated in: 2 aged_person, 7 general, 3 sensory_impairments.

04
claude-sonnet-4.5's per-category numbers should be read as a lower-confidence bound.

At κ = 0.419 against human labels ("moderate"), this judge is meaningfully less reliable than the other — treat its category-level figures with more caution, especially where it diverges sharply from the other judge.

Source: Alternative_route_problem_Evaluation.ipynb · open-ended track only · rubric: visual_cue_coverage · intent_match · spatial_correctness (strict intent+spatial, visual cue 50% partial credit).