Leaderboard

Stage-2 fleet survival rate (%), mean over 3 trials (3 independent trials per model × scenario). The leaderboard is updated as new models and execution modes are evaluated.

Structured tool calling over multiple turns with mandatory reasoning and an exploration guard.
#ModelAvgAntenna TrapDeployment Zone TrapWeatherWins / 14
1Claude Opus 4.5 Anthropic68.0 ±5.079.064.327.35
2GLM-5.2 Zhipu67.8 ±8.376.565.532.54
3GPT-5.5 High OpenAI67.8 ±4.080.862.824.35
4Qwen3.7 Max Alibaba Qwen66.9 ±6.276.165.322.15
5GPT-5.5 OpenAI66.8 ±5.779.561.328.44
6GLM-5.1 Zhipu66.5 ±9.173.166.626.44
7Claude Opus 4.7 Anthropic65.4 ±7.477.161.323.35
8GPT-5.5 XHigh OpenAI65.4 ±6.679.459.820.76
9Grok 4.1 Fast xAI65.4 ±6.380.359.219.54
10DeepSeek V4 Pro DeepSeek64.6 ±7.175.061.325.13
11Kimi K2.6 Moonshot AI64.6 ±10.672.363.128.92
12MiniMax M2 MiniMax64.5 ±6.173.862.125.84
13Claude Sonnet 4.5 Anthropic64.4 ±8.274.960.528.84
14Gemini 3.5 Flash Google64.1 ±9.675.659.825.34
15MIMO V2.5 Pro Xiaomi MiMo63.4 ±6.772.760.726.91
16MiniMax M2.1 MiniMax63.2 ±9.270.861.429.72
17DeepSeek V4 Flash DeepSeek62.3 ±11.072.059.623.42
18HY3 Preview Tencent Hunyuan61.5 ±11.070.659.123.02
19GPT-OSS-120B OpenAI61.1 ±9.474.454.527.24
20MIMO V2 Flash Xiaomi MiMo60.9 ±9.374.455.716.23
21GPT-5 Mini OpenAI60.4 ±7.675.751.729.24
22DeepSeek V3.2 Think DeepSeek59.1 ±7.266.857.425.01
23GPT-5.2 High OpenAI59.0 ±5.573.751.026.33
24DeepSeek V3.2 DeepSeek58.7 ±7.764.757.729.21
25GLM-4.7 Zhipu58.2 ±6.267.455.820.11
26Kimi K2.5 Moonshot AI58.2 ±10.865.056.031.80
27MiniMax M2.7 MiniMax57.6 ±11.762.856.931.01
28GPT-5.2 OpenAI57.3 ±6.370.850.920.72
29Grok 4.20 xAI50.0 ±7.454.050.621.51
30Gemini 3.1 Flash Lite Google49.5 ±10.351.252.618.10

Survival rate (%), mean of 3 independent trials per model × scenario. Hover a cell for std and 95% CI where available. Green cells meet the scenario win threshold (75%, or 55% for weather_noise). Family columns are means over per-scenario results; rank (#) is always by overall average within the mode.

Win thresholds & optimal designs

Scenario family Optimal intervention Optimal survival Win threshold
Antenna Trap antenna_def = 0 (stealth) ~82% 75%
Deployment Zone Trap shield_def = 25 + signal_filter ~80% 75%
Weather antenna_def ≈ 8 with Stage 2 tuning ~78% 55%

The optimal design for each scenario is derived analytically from the underlying SCM and verified empirically on fleets of 1,000 drones (agreement within ±2–3 pp). Thresholds sit 5–8% below the theoretical optimum, so the games are winnable with correct causal understanding but not through random exploration.

Non-LLM baselines

Baseline Strategy Avg survival
Default Submit the initial design unchanged 49.0%
Random Uniformly sample each DEF value from [0, 50] 52.0%
Uniform High Set all components to DEF = 50 52.7%
No-Explore LLM 10 random deploys, then LLM analyzes and submits 52–63%

Simple rule-based baselines reach 49–53% — and can even outperform several full-agent models on bias-heavy scenarios, underscoring that correlational shortcuts are not enough.

OpenCode vs other modes

Full per-scenario OpenCode results are in the OpenCode (Coding Agent) tab of the leaderboard above. Across every model evaluated, the coding-agent framework outperforms the model's own ReAct (Agent) and Prompting scores, yet the average still falls short of the 75% win threshold — the causal thinking gap persists even with a full coding agent. A 9-model comparison (average 70% OpenCode vs 62.7% ReAct / 63.1% Prompting):

Model Δ survival, OpenCode vs ReAct
GPT-5.2 +14.5
GPT-5.2 High +11.0
DeepSeek V4 Pro +10.6
Claude Opus 4.7 +7.0
Qwen3.7 Max +6.6
GPT-5 Mini +6.3
GPT-5.5 +4.7
Kimi K2.5 +2.6
Grok 4.1 Fast +2.4