Naval Architecture
Hull physics, stepped hull dynamics, pontoon mechanics, propeller interactions, and hydrodynamics.
A rigorous benchmark for evaluating AI knowledge of marine terminology, physics, safety, and regulations
41 questions covering marine terminology and factual knowledge
| Rank | Model | Provider | Score | Performance |
|---|---|---|---|---|
| 1 | Gemini 3 Pro | 39/41 (95.1%) | ||
| 2 | Llama 4 Maverick | Meta | 38/41 (92.7%) | |
| 3 | Kimi K2 | Moonshot | 37/41 (90.2%) | |
| 4 | GPT-5.2 Instant | OpenAI | 36/41 (87.8%) | |
| 4 | Gemini 3 Flash | 36/41 (87.8%) | ||
| 6 | Claude Sonnet 4.5 | Anthropic | 35/41 (85.4%) | |
| 7 | DeepSeek v3.2 | DeepSeek | 32/41 (78.0%) | |
| 8 | Grok 4.1 Fast | xAI | 31/41 (75.6%) | |
| 9 | Qwen 3 235B | Alibaba | 30/41 (73.2%) | |
| 9 | Claude Haiku 4.5 | Anthropic | 30/41 (73.2%) |
Same v1 caveat applies: single runs, small n — the v2 protocol replaces both boards.
27 questions requiring physical reasoning and counter-intuitive understanding
| Rank | Model | Provider | Score | Performance |
|---|---|---|---|---|
| 1 | Kimi K2 | Moonshot | 24/27 (88.9%) | |
| 2 | GPT-5.2 Instant | OpenAI | 23/27 (85.2%) | |
| 3 | Gemini 3 Pro | 22/27 (81.5%) | ||
| 3 | Gemini 3 Flash | 22/27 (81.5%) | ||
| 3 | Claude Sonnet 4.5 | Anthropic | 22/27 (81.5%) | |
| 6 | Qwen 3 235B | Alibaba | 20/27 (74.1%) | |
| 7 | Llama 4 Maverick | Meta | 19/27 (70.4%) | |
| 8 | DeepSeek v3.2 | DeepSeek | 18/27 (66.7%) | |
| 8 | Grok 4.1 Fast | xAI | 18/27 (66.7%) | |
| 8 | Claude Haiku 4.5 | Anthropic | 18/27 (66.7%) |
Detailed breakdown of model performance across knowledge domains (Advanced Physics test)
| Model | Score | Missed |
|---|---|---|
| Kimi K2 | 4/6 | Q5 (underskinning), Q6 (strakes) |
| GPT-5.2 Instant | 3/6 | Q1 (trim), Q5, Q6 |
| Claude Sonnet 4.5 | 4/6 | Q1 (trim), Q5 |
| Gemini 3 Pro | 3/6 | Q1, Q4, Q5 |
| DeepSeek v3.2 | 3/6 | Q4, Q5, Q6 |
| Model | Score | Notes |
|---|---|---|
| Kimi K2 | 3/3 | Only model with perfect prop walk answers |
| GPT-5.2 Instant | 3/3 | Correct on prop walk (helpful) |
| Claude Sonnet 4.5 | 2/3 | Missed Q9 (prop walk direction) |
| DeepSeek v3.2 | 2/3 | Missed Q9 |
| Llama 4 Maverick | 2/3 | Missed Q9 |
| Model | Score | Notes |
|---|---|---|
| Kimi K2 | 4/4 | Got runaway cause partially right |
| GPT-5.2 Instant | 3/4 | Missed Q13 (turbo oil) |
| Claude Sonnet 4.5 | 3/4 | Missed Q13 |
| Gemini 3 Pro | 3/4 | Missed Q13 |
| All others | 3/4 | Q13 universally challenging |
| Model | Score | Notes |
|---|---|---|
| All models | 3/3 | Universally strong category |
| Model | Score | Notes |
|---|---|---|
| Most models | 6/6 | Strong on basic safety |
| Grok 4.1 Fast | 5/6 | Missed Q21 (said starboard instead of port) |
| Claude Haiku 4.5 | 4/6 | Missed Q19, Q21 |
| Model | Score | Notes |
|---|---|---|
| Kimi K2 | 2/3 | Missed Q25 |
| GPT-5.2 Instant | 3/3 | Perfect |
| Claude Sonnet 4.5 | 3/3 | Perfect |
| DeepSeek v3.2 | 1/3 | Swapped Q23/Q24 answers |
| Grok 4.1 Fast | 1/3 | Weak on sonar interpretation |
MarineBench tests across four core knowledge domains
Hull physics, stepped hull dynamics, pontoon mechanics, propeller interactions, and hydrodynamics.
Diesel thermodynamics, wet stacking, runaway engines, electrical systems, and corrosion control.
COLREGs navigation rules, wake sports safety, emergency procedures, and right-of-way logic.
Sonar reading, fish arch interpretation, chart symbols, and radar analysis.
Questions where most AI models fail — exposing gaps in practical marine knowledge
Q: Stepped hull in hard turn — trim up or down?
Common wrong: "Trim down"
Correct: Trim UP (down causes spin-out)
Q: Pontoon surges/sluggish despite good RPMs — what to check?
Common wrong: Prop hub, debris, fuel
Correct: Underskinning
Q: Why doesn't turning off ignition stop a diesel runaway?
Common wrong: "Compression ignition"
Correct: Engine running on its own oil
Q: What does fish arch LENGTH indicate?
Common wrong: "Fish size"
Correct: Time in beam (NOT fish size)
Expanding MarineBench from a static leaderboard into a living benchmark platform.
Every question includes a source tag. Upcoming releases will link each domain pack to its research docs for fast auditing and updates.
MarineBench is designed to rigorously test AI capabilities in the marine domain — where the stakes of misinformation are measured in damaged equipment, unsafe situations, or worse.
Inspired by SkateBench, this benchmark goes beyond simple vocabulary testing to evaluate physical reasoning, counter-intuitive physics understanding, and safety-critical knowledge.
Unlike general knowledge benchmarks, MarineBench specifically targets the practical knowledge gaps that could lead to real-world harm when AI systems provide incorrect marine advice.