The Spectrum Dispatch News

technology

M5 Ultra Mac Studio Shows Promise for Running Local AI Models

Apple's latest Mac Studio can run sophisticated AI models locally without cloud costs, according to tests by MacStories.

M5 Ultra Mac Studio Shows Promise for Running Local AI Models

Apple’s M5 Ultra Mac Studio is capable of running advanced AI models locally with strong performance and no cloud expenses, according to a review by MacStories editor Viticci. Over four days of testing, the top-of-the-line model with 256 GB of RAM demonstrated significant improvements over its M3 Ultra predecessor for local AI workloads.

M5 Ultra Mac Studio Shows Promise for Running Local AI Models

The M5 Ultra features a redesigned architecture using UltraFusion to connect two dual-die M5 Max chips into a quad-die configuration. For AI tasks, this delivers up to 4.5× the peak GPU compute compared to the M3 Ultra, with 80 GPU cores each equipped with a Neural Accelerator. Memory bandwidth increased 50%, jumping from 819 GB/s on the M3 Ultra to 1.2 TB/s on the M5 Ultra.

In practical testing, the M5 Ultra demonstrated approximately 70% faster response generation on average compared to the M3 Ultra when running the Qwen 3.8-Flash-Next model. According to the review, token prefill speed—how quickly a model processes an input prompt—and token generation speed both improved noticeably. These improvements enable faster iteration in agentic workflows that use large context windows.

Viticci used the M5 Ultra to run local AI agents as personal assistants, replacing cloud-based alternatives like Siri AI for daily tasks. The machine proved capable of running models continuously without performance degradation over extended multi-turn conversations.

For comparison, the reviewer tested the M5 Ultra against an RTX 5090 graphics card in a desktop gaming PC. While the RTX 5090 maintains higher memory bandwidth, the reviewer noted preferring the M5 Ultra for its compact size, thermal efficiency, lower noise levels, and macOS ecosystem.

A significant use case emerged during research for an iOS and iPadOS 27 review. Viticci created a custom app called Desk that ran local AI agents based on DeepSeek V4 Flash continuously for 99 days to transcribe WWDC sessions, extract features from webpages and PDFs, cross-reference information, and organize data via the Notion API. Running this persistent background task on local hardware cost $0, whereas cloud-based APIs would have been prohibitively expensive.

Key facts

  • M5 Ultra Mac Studio delivers 4.5× the peak GPU compute for AI compared to M3 Ultra
  • Memory bandwidth increased 50% to 1.2 TB/s from 819 GB/s
  • M5 Ultra was ~70% faster on average at generating responses versus M3 Ultra in testing
  • The reviewer ran local AI agents continuously for 99 days during iOS 27 research with zero cloud API costs
  • M5 Ultra can run models like Qwen 3.8-Flash-Next as personal assistants with fast iteration suitable for agentic loops

Sources

← All posts