Leaderboard
Stage-2 fleet survival rate (%), mean over 3 trials (3 independent trials per model × scenario). The leaderboard is updated as new models and execution modes are evaluated.
| # | Model | Avg ▾ | Antenna Trap | Deployment Zone Trap | Weather | Wins / 14 |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.5 Anthropic | 68.0 ±5.0 | 79.0 | 64.3 | 27.3 | 5 |
| 2 | GLM-5.2 Zhipu | 67.8 ±8.3 | 76.5 | 65.5 | 32.5 | 4 |
| 3 | GPT-5.5 High OpenAI | 67.8 ±4.0 | 80.8 | 62.8 | 24.3 | 5 |
| 4 | Qwen3.7 Max Alibaba Qwen | 66.9 ±6.2 | 76.1 | 65.3 | 22.1 | 5 |
| 5 | GPT-5.5 OpenAI | 66.8 ±5.7 | 79.5 | 61.3 | 28.4 | 4 |
| 6 | GLM-5.1 Zhipu | 66.5 ±9.1 | 73.1 | 66.6 | 26.4 | 4 |
| 7 | Claude Opus 4.7 Anthropic | 65.4 ±7.4 | 77.1 | 61.3 | 23.3 | 5 |
| 8 | GPT-5.5 XHigh OpenAI | 65.4 ±6.6 | 79.4 | 59.8 | 20.7 | 6 |
| 9 | Grok 4.1 Fast xAI | 65.4 ±6.3 | 80.3 | 59.2 | 19.5 | 4 |
| 10 | DeepSeek V4 Pro DeepSeek | 64.6 ±7.1 | 75.0 | 61.3 | 25.1 | 3 |
| 11 | Kimi K2.6 Moonshot AI | 64.6 ±10.6 | 72.3 | 63.1 | 28.9 | 2 |
| 12 | MiniMax M2 MiniMax | 64.5 ±6.1 | 73.8 | 62.1 | 25.8 | 4 |
| 13 | Claude Sonnet 4.5 Anthropic | 64.4 ±8.2 | 74.9 | 60.5 | 28.8 | 4 |
| 14 | Gemini 3.5 Flash Google | 64.1 ±9.6 | 75.6 | 59.8 | 25.3 | 4 |
| 15 | MIMO V2.5 Pro Xiaomi MiMo | 63.4 ±6.7 | 72.7 | 60.7 | 26.9 | 1 |
| 16 | MiniMax M2.1 MiniMax | 63.2 ±9.2 | 70.8 | 61.4 | 29.7 | 2 |
| 17 | DeepSeek V4 Flash DeepSeek | 62.3 ±11.0 | 72.0 | 59.6 | 23.4 | 2 |
| 18 | HY3 Preview Tencent Hunyuan | 61.5 ±11.0 | 70.6 | 59.1 | 23.0 | 2 |
| 19 | GPT-OSS-120B OpenAI | 61.1 ±9.4 | 74.4 | 54.5 | 27.2 | 4 |
| 20 | MIMO V2 Flash Xiaomi MiMo | 60.9 ±9.3 | 74.4 | 55.7 | 16.2 | 3 |
| 21 | GPT-5 Mini OpenAI | 60.4 ±7.6 | 75.7 | 51.7 | 29.2 | 4 |
| 22 | DeepSeek V3.2 Think DeepSeek | 59.1 ±7.2 | 66.8 | 57.4 | 25.0 | 1 |
| 23 | GPT-5.2 High OpenAI | 59.0 ±5.5 | 73.7 | 51.0 | 26.3 | 3 |
| 24 | DeepSeek V3.2 DeepSeek | 58.7 ±7.7 | 64.7 | 57.7 | 29.2 | 1 |
| 25 | GLM-4.7 Zhipu | 58.2 ±6.2 | 67.4 | 55.8 | 20.1 | 1 |
| 26 | Kimi K2.5 Moonshot AI | 58.2 ±10.8 | 65.0 | 56.0 | 31.8 | 0 |
| 27 | MiniMax M2.7 MiniMax | 57.6 ±11.7 | 62.8 | 56.9 | 31.0 | 1 |
| 28 | GPT-5.2 OpenAI | 57.3 ±6.3 | 70.8 | 50.9 | 20.7 | 2 |
| 29 | Grok 4.20 xAI | 50.0 ±7.4 | 54.0 | 50.6 | 21.5 | 1 |
| 30 | Gemini 3.1 Flash Lite Google | 49.5 ±10.3 | 51.2 | 52.6 | 18.1 | 0 |
Survival rate (%), mean of 3 independent trials per model × scenario. Hover a cell for std and 95% CI where available. Green cells meet the scenario win threshold (75%, or 55% for weather_noise). Family columns are means over per-scenario results; rank (#) is always by overall average within the mode.
Win thresholds & optimal designs
| Scenario family | Optimal intervention | Optimal survival | Win threshold |
|---|---|---|---|
| Antenna Trap | antenna_def = 0 (stealth) | ~82% | 75% |
| Deployment Zone Trap | shield_def = 25 + signal_filter | ~80% | 75% |
| Weather | antenna_def ≈ 8 with Stage 2 tuning | ~78% | 55% |
The optimal design for each scenario is derived analytically from the underlying SCM and verified empirically on fleets of 1,000 drones (agreement within ±2–3 pp). Thresholds sit 5–8% below the theoretical optimum, so the games are winnable with correct causal understanding but not through random exploration.
Non-LLM baselines
| Baseline | Strategy | Avg survival |
|---|---|---|
| Default | Submit the initial design unchanged | 49.0% |
| Random | Uniformly sample each DEF value from [0, 50] | 52.0% |
| Uniform High | Set all components to DEF = 50 | 52.7% |
| No-Explore LLM | 10 random deploys, then LLM analyzes and submits | 52–63% |
Simple rule-based baselines reach 49–53% — and can even outperform several full-agent models on bias-heavy scenarios, underscoring that correlational shortcuts are not enough.
OpenCode vs other modes
Full per-scenario OpenCode results are in the OpenCode (Coding Agent) tab of the leaderboard above. Across every model evaluated, the coding-agent framework outperforms the model's own ReAct (Agent) and Prompting scores, yet the average still falls short of the 75% win threshold — the causal thinking gap persists even with a full coding agent. A 9-model comparison (average 70% OpenCode vs 62.7% ReAct / 63.1% Prompting):
| Model | Δ survival, OpenCode vs ReAct |
|---|---|
| GPT-5.2 | +14.5 |
| GPT-5.2 High | +11.0 |
| DeepSeek V4 Pro | +10.6 |
| Claude Opus 4.7 | +7.0 |
| Qwen3.7 Max | +6.6 |
| GPT-5 Mini | +6.3 |
| GPT-5.5 | +4.7 |
| Kimi K2.5 | +2.6 |
| Grok 4.1 Fast | +2.4 |