Exact catalog questions
Base prices, option costs, fuel capacity, fishbox volume, generator model, LOA, acceleration and fuel economy.
MarineBench tests exact product knowledge: prices, specs, options, performance numbers and model-specific details from the 2026 Regulator Marine catalog.
Run May 4, 2026 via the Gemini API, single run per model. Twenty-five catalog questions. No-context models answered from model memory only; the context run included the full source catalog in the prompt.
| Model | Mode | Score | Accuracy | Performance |
|---|---|---|---|---|
| Gemini 2.5 Flash | No catalog context | 5/25 | 20% | |
| Gemini 3 Flash Preview | No catalog context | 8/25 | 32% | |
| Gemini 3 Pro Preview | No catalog context | 10/25 | 40% | |
| Gemini 3.1 Pro Preview | No catalog context | 10/25 | 40% | |
| Gemini 3 Flash Preview | With catalog context | 25/25 | 100% |
At 25 questions and one run per model, differences between the no-context models are within statistical noise. The finding that holds up is the gap between no-context and with-context performance (p < 0.001).
The failure mode is not writing quality. It is confident specificity without reliable source data.
Base prices, option costs, fuel capacity, fishbox volume, generator model, LOA, acceleration and fuel economy.
Recent Gemini Pro preview models still missed most exact product facts without access to the catalog.
A faster model with the right source material beat stronger models relying on general memory.
These are the kinds of errors a dealer or OEM cannot ship in a customer-facing AI experience.
Gemini 2.5 Flash: Seakeeper 2 on the Regulator 31 costs $26,995. Catalog: $60,695.
Gemini 3 Flash: Regulator 35 has 16-inch Garmin displays. Catalog: dual 22-inch displays.
Gemini 3 Flash: top speed is 64.4 mph. Catalog: 64.7 mph.
Gemini 2.5 Flash: 26XO came back after hiatus. Catalog: Regulator 25.
The first MarineBench demo showed the same pattern across multiple providers before the May refresh. Collected manually via t3.chat, single run per model — shown as an illustrative baseline, not a scored leaderboard.
| Run | Mode | Score | Accuracy | Performance |
|---|---|---|---|---|
| Grok 4.1 Fast | No catalog context | 1/25 | 4% | |
| Claude Opus 4.5 | No catalog context | 5/25 | 20% | |
| GPT-5.2 Instant | No catalog context | 8/25 | 32% | |
| Gemini 3 Pro | No catalog context | 11/25 | 44% | |
| Gemini Flash | With catalog context | 25/25 | 100% |
The winning system is not simply the biggest model. It is a product expert built from trusted dealer and OEM source material, evaluated against questions that reveal whether the AI actually knows the business.