The Spectrum Dispatch News

technology

AI Models Play StarCraft: Codex Astra Dominates, Grok Struggles

A benchmark tested large language models playing StarCraft: Brood War, revealing significant gaps in strategic thinking. Codex Astra achieved a 100% win rate.

AI Models Play StarCraft: Codex Astra Dominates, Grok Struggles

Researchers at Freestyle created Brood War Bench, a real-time strategy game benchmark that pits large language models against each other in competitive matches of StarCraft: Brood War. The experiment evaluated 19 different model configurations across a round-robin tournament format to assess their ability to play one of gaming’s most demanding strategic titles.

AI Models Play StarCraft: Codex Astra Dominates, Grok Struggles

According to the report, Codex Astra emerged as the clear winner, with its “xhigh” configuration achieving an 18-0 record and 100% win rate. The same model’s “medium” setting finished second with a 16-2 record (88.9% win rate). Claude Fable placed third with 15 wins and 3 losses (83.3% win rate). In contrast, Grok models performed poorly, with Grok 4.6 winning only 3 games total across all configurations, and Claude Haiku failed to win any games.

Key observations revealed fundamental differences in how models approached the game. Codex models consistently discovered early-game disruption tactics before developing sustained production strategies. According to the report, Codex “often sent a Probe across the map to attack workers or buildings” in Protoss matchups, which “worked shockingly well as the opposing agents often spent dozens of seconds thinking about what to do about a probe instead of doing anything else.” However, the report notes that Codex models “were much weaker at sustained production,” frequently delaying technology upgrades and trickling units into defended bases.

Codex models also exhibited organizational complexity by creating separate subagents to manage the economy, army production, and army control, though these agents rarely coordinated effectively. The report describes this as “a common beginner mistake: sending units in one at a time instead of waiting for a critical mass and a planned attack timing.”

Grok 4.6 demonstrated fundamentally different failures. According to the benchmark, Grok “spent the game between actions,” with one match generating 11,138 reasoning tokens but only six command batches across 43 minutes without fielding any combat units. The model struggled to maintain observation-action loops, often issuing only a handful of unit commands before failing to follow through.

Claude Fable distinguished itself by attempting conventional gameplay. The report states that Fable “usually tried to build an economy and climb the tech tree instead of stopping at the first unit available,” winning games where it reached advanced technologies like Lair, Spire, and Mutalisks.

Despite these variations, the researchers concluded that “none of the models played beyond a beginner level.” According to the report, even Astra and Fable “were unable to build complex armies, defend simple attacks or play concrete strategies,” and that “a beginner playing photon rush would win every single one of these games.” Nevertheless, the authors expressed optimism about future development, noting that “this benchmark is nowhere near exhausted” and that they look forward to watching models improve.

Key facts

  • Codex Astra achieved a perfect 18-0 record in Brood War Bench matches
  • Grok 4.6 models won only 3 total games across all configurations
  • Claude Fable placed third with 15 wins and 3 losses
  • None of the tested models played beyond beginner level, according to researchers
  • Codex models created separate subagents for economy, army production, and army control that rarely communicated with each other

Sources

← All posts