AI models know boating. They do not know your catalog.

MarineBench tests exact product knowledge: prices, specs, options, performance numbers and model-specific details from the 2026 Regulator Marine catalog.

Fresh Benchmark Results

Run May 4, 2026 via the Gemini API, single run per model. Twenty-five catalog questions. No-context models answered from model memory only; the context run included the full source catalog in the prompt.

Model Mode Score Accuracy Performance
Gemini 2.5 Flash No catalog context 5/25 20%
Gemini 3 Flash Preview No catalog context 8/25 32%
Gemini 3 Pro Preview No catalog context 10/25 40%
Gemini 3.1 Pro Preview No catalog context 10/25 40%
Gemini 3 Flash Preview With catalog context 25/25 100%

At 25 questions and one run per model, differences between the no-context models are within statistical noise. The finding that holds up is the gap between no-context and with-context performance (p < 0.001).

Why This Matters

The failure mode is not writing quality. It is confident specificity without reliable source data.

25

Exact catalog questions

Base prices, option costs, fuel capacity, fishbox volume, generator model, LOA, acceleration and fuel economy.

40%

Best raw score refreshed

Recent Gemini Pro preview models still missed most exact product facts without access to the catalog.

100%

Context closes the gap

A faster model with the right source material beat stronger models relying on general memory.

Representative Misses

These are the kinds of errors a dealer or OEM cannot ship in a customer-facing AI experience.

Option pricing drift

Gemini 2.5 Flash: Seakeeper 2 on the Regulator 31 costs $26,995. Catalog: $60,695.

Wrong standard equipment

Gemini 3 Flash: Regulator 35 has 16-inch Garmin displays. Catalog: dual 22-inch displays.

Performance near-miss

Gemini 3 Flash: top speed is 64.4 mph. Catalog: 64.7 mph.

Model history miss

Gemini 2.5 Flash: 26XO came back after hiatus. Catalog: Regulator 25.

Original Cross-Model Baseline

The first MarineBench demo showed the same pattern across multiple providers before the May refresh. Collected manually via t3.chat, single run per model — shown as an illustrative baseline, not a scored leaderboard.

Run Mode Score Accuracy Performance
Grok 4.1 Fast No catalog context 1/25 4%
Claude Opus 4.5 No catalog context 5/25 20%
GPT-5.2 Instant No catalog context 8/25 32%
Gemini 3 Pro No catalog context 11/25 44%
Gemini Flash With catalog context 25/25 100%

The Dealer Edge takeaway

The winning system is not simply the biggest model. It is a product expert built from trusted dealer and OEM source material, evaluated against questions that reveal whether the AI actually knows the business.