MarineBench

A rigorous benchmark for evaluating AI knowledge of marine terminology, physics, safety, and regulations

Terminology Test Leaderboard

41 questions covering marine terminology and factual knowledge

⚠️ v1 preliminary data. Single manual runs at this sample size carry roughly ±13–18 point error bars, so treat close rankings as ties. A v2 run — 290+ questions, 3 runs per model at temperature 0, confidence intervals on every score — is in progress. Full methodology →
Rank Model Provider Score Performance
1 Gemini 3 Pro Google 39/41 (95.1%)
2 Llama 4 Maverick Meta 38/41 (92.7%)
3 Kimi K2 Moonshot 37/41 (90.2%)
4 GPT-5.2 Instant OpenAI 36/41 (87.8%)
4 Gemini 3 Flash Google 36/41 (87.8%)
6 Claude Sonnet 4.5 Anthropic 35/41 (85.4%)
7 DeepSeek v3.2 DeepSeek 32/41 (78.0%)
8 Grok 4.1 Fast xAI 31/41 (75.6%)
9 Qwen 3 235B Alibaba 30/41 (73.2%)
9 Claude Haiku 4.5 Anthropic 30/41 (73.2%)

Advanced Physics Test Leaderboard

Same v1 caveat applies: single runs, small n — the v2 protocol replaces both boards.

27 questions requiring physical reasoning and counter-intuitive understanding

Rank Model Provider Score Performance
1 Kimi K2 Moonshot 24/27 (88.9%)
2 GPT-5.2 Instant OpenAI 23/27 (85.2%)
3 Gemini 3 Pro Google 22/27 (81.5%)
3 Gemini 3 Flash Google 22/27 (81.5%)
3 Claude Sonnet 4.5 Anthropic 22/27 (81.5%)
6 Qwen 3 235B Alibaba 20/27 (74.1%)
7 Llama 4 Maverick Meta 19/27 (70.4%)
8 DeepSeek v3.2 DeepSeek 18/27 (66.7%)
8 Grok 4.1 Fast xAI 18/27 (66.7%)
8 Claude Haiku 4.5 Anthropic 18/27 (66.7%)

Performance by Category

Detailed breakdown of model performance across knowledge domains (Advanced Physics test)

Hull Physics

6 questions · Avg: 58%
ModelScoreMissed
Kimi K24/6Q5 (underskinning), Q6 (strakes)
GPT-5.2 Instant3/6Q1 (trim), Q5, Q6
Claude Sonnet 4.54/6Q1 (trim), Q5
Gemini 3 Pro3/6Q1, Q4, Q5
DeepSeek v3.23/6Q4, Q5, Q6

Propellers & Prop Walk

3 questions · Avg: 80%
ModelScoreNotes
Kimi K23/3Only model with perfect prop walk answers
GPT-5.2 Instant3/3Correct on prop walk (helpful)
Claude Sonnet 4.52/3Missed Q9 (prop walk direction)
DeepSeek v3.22/3Missed Q9
Llama 4 Maverick2/3Missed Q9

Diesel Systems

4 questions · Avg: 75%
ModelScoreNotes
Kimi K24/4Got runaway cause partially right
GPT-5.2 Instant3/4Missed Q13 (turbo oil)
Claude Sonnet 4.53/4Missed Q13
Gemini 3 Pro3/4Missed Q13
All others3/4Q13 universally challenging

Electrical & Corrosion

3 questions · Avg: 97%
ModelScoreNotes
All models3/3Universally strong category

Safety & COLREGs

6 questions · Avg: 95%
ModelScoreNotes
Most models6/6Strong on basic safety
Grok 4.1 Fast5/6Missed Q21 (said starboard instead of port)
Claude Haiku 4.54/6Missed Q19, Q21

Sonar & Electronics

3 questions · Avg: 77%
ModelScoreNotes
Kimi K22/3Missed Q25
GPT-5.2 Instant3/3Perfect
Claude Sonnet 4.53/3Perfect
DeepSeek v3.21/3Swapped Q23/Q24 answers
Grok 4.1 Fast1/3Weak on sonar interpretation

The Four Pillars

MarineBench tests across four core knowledge domains

Naval Architecture

Hull physics, stepped hull dynamics, pontoon mechanics, propeller interactions, and hydrodynamics.

Marine Engineering

Diesel thermodynamics, wet stacking, runaway engines, electrical systems, and corrosion control.

Operational Safety

COLREGs navigation rules, wake sports safety, emergency procedures, and right-of-way logic.

Electronic Interpretation

Sonar reading, fish arch interpretation, chart symbols, and radar analysis.

Killer Questions

Questions where most AI models fail — exposing gaps in practical marine knowledge

Stepped Hull Trim 80% failed

Q: Stepped hull in hard turn — trim up or down?

Common wrong: "Trim down"
Correct: Trim UP (down causes spin-out)

Pontoon Surge Drag 100% failed

Q: Pontoon surges/sluggish despite good RPMs — what to check?

Common wrong: Prop hub, debris, fuel
Correct: Underskinning

Diesel Runaway 100% failed

Q: Why doesn't turning off ignition stop a diesel runaway?

Common wrong: "Compression ignition"
Correct: Engine running on its own oil

Sonar Arch Length 80% failed

Q: What does fish arch LENGTH indicate?

Common wrong: "Fish size"
Correct: Time in beam (NOT fish size)

Beyond the Baseline: The Context Effect

MarineBench measures what models inherently know. The separate question — can the baseline be raised? — has a demo of its own.

🚀 In our Regulator Marine case study, frontier models scored 20–44% on exact product-catalog questions from memory alone. The same class of model, given a trusted catalog through a context framework, scored 100%. That gap — not model choice — is where dealer and OEM AI gets won. See the Context Effect demo →

What's Next

Expanding MarineBench from a static leaderboard into a living benchmark platform.

v2: The Rigor Upgrade — in progress

  • 290+ question bank, held-out synthesis questions
  • Negative controls that score abstention honestly
  • Temp-0 API runs, 3 per model, one question per call
  • Confidence intervals on every published score

Domain Packs

  • Ski & wake boat knowledge
  • Center console & offshore
  • Pontoon & tritoon systems
  • Dealer operations & specs

Benchmark Studio

  • Create new benchmarks in minutes
  • Select model/agent lineups
  • Run tests across APIs automatically
  • Schedule reruns as models update

RAG & Tuned Showdowns

  • Compare base vs RAG-augmented agents
  • Track tuned models vs general LLMs
  • Measure small+tools vs big models

Multimodal Track

  • Image-based configuration evaluation
  • Parameter-efficient tuning experiments
  • Curated 300-example training sets

Every question includes a source tag. Upcoming releases will link each domain pack to its research docs for fast auditing and updates.

About MarineBench

MarineBench is designed to rigorously test AI capabilities in the marine domain — where the stakes of misinformation are measured in damaged equipment, unsafe situations, or worse.

Inspired by SkateBench, this benchmark goes beyond simple vocabulary testing to evaluate physical reasoning, counter-intuitive physics understanding, and safety-critical knowledge.

Unlike general knowledge benchmarks, MarineBench specifically targets the practical knowledge gaps that could lead to real-world harm when AI systems provide incorrect marine advice.

Current version: v0.2 (68 questions) Target: 500 questions