The Spectrum Dispatch News

technology

DeepSeek V4 Flash 0731 scores 89% on reasoning benchmark at low cost

The model achieves strong performance on ARC-AGI tasks, with three reasoning variants tested at different levels of computational effort and cost.

DeepSeek V4 Flash 0731 scores 89% on reasoning benchmark at low cost

DeepSeek released V4 Flash 0731, a model offering three reasoning variants tested on the ARC-AGI benchmark suite according to results posted on arcprize.org.

DeepSeek V4 Flash 0731 scores 89% on reasoning benchmark at low cost

At maximum effort, the model achieves 89.0% accuracy on ARC-AGI-1 Semi-Private tasks at $0.02 per task and 61.4% on ARC-AGI-2 Semi-Private tasks at $0.04 per task. The model was evaluated across three reasoning intensity levels: Max, High, and Low.

On the ARC-AGI-1 benchmark, which consists of 400 public evaluation tasks, the Max variant achieved 89.0%, the High variant scored 87.0%, and the Low variant scored 84.0%. Performance decreased on the more challenging ARC-AGI-2 benchmark with 120 public tasks, where the Max variant achieved 61.4%, High scored 56.0%, and Low achieved 46.0%.

The results show the tradeoff between computational effort and accuracy across the three variants. The Max variant consistently outperformed the High and Low variants across both benchmarks. On ARC-AGI-1, the Max and High variants showed relatively close performance, differing by 2 percentage points, while the Low variant fell further behind at 84.0%. The gap between variants widened significantly on ARC-AGI-2, with the Max variant leading by 5.4 percentage points over High and 15.4 points over Low.

The cost structure reflects the computational demands, with ARC-AGI-2 tasks priced twice as high as ARC-AGI-1 tasks, suggesting greater complexity in the second benchmark. The public evaluation results show task-level performance across all variants, revealing which specific problems each reasoning level could successfully solve. The model’s ability to maintain high performance on ARC-AGI-1 even at the Low setting suggests efficient reasoning patterns for less complex tasks.

Key facts

  • DeepSeek V4 Flash 0731 achieves 89.0% on ARC-AGI-1 at max effort for $0.02 per task
  • The model scores 61.4% on ARC-AGI-2 at max effort for $0.04 per task
  • Three reasoning variants tested: Max (89.0% on ARC-AGI-1), High (87.0%), and Low (84.0%)
  • ARC-AGI-2 performance: Max 61.4%, High 56.0%, Low 46.0%

Sources

← All posts