Desert Ant Labs, a European AI research lab, announced the launch of 18 on-device intelligence models designed to run efficiently on consumer devices without relying on cloud infrastructure. According to the announcement, the first 18 models—12 stable and six in beta—are now available through a unified SDK supporting Swift, Kotlin, and JavaScript.

The models target specific tasks with claimed performance advantages. Voz, an audio transcription model, transcribes 10 minutes of audio in two seconds on an iPhone, described as 4.7 times faster than Whisper. Clear, a 9MB audio enhancement model, can process a five-minute laptop recording into studio-quality audio in one second. Redact masks personally identifiable information like names, addresses, and card numbers in real time across 27 languages. Tongue identifies 84 languages from just three words using a 2MB model.
According to the lab, each model answers in milliseconds and costs nothing to run, allowing developers to deploy intelligence across product interactions without token cost limitations or inference speed constraints. The models are free to use up to 100,000 monthly active devices, with no token-based pricing or login requirements.
Desert Ant Labs was founded by the team behind Detail, a video editing application. The founders note they previously relied on cloud APIs for features like automated clip creation and audio enhancement, but increasingly large infrastructure bills prompted them to develop proprietary on-device models. The lab replaced several third-party services with its own models: Dolby for audio enhancement with Clear, and Claude Sonnet for video clip generation with Clips, a 284MB model that processes 10-minute videos into a dozen clips in 5 seconds.
The company emphasizes data sovereignty and privacy as core motivations, noting that on-device processing ensures customer data never leaves their devices and features do not depend on cloud services. The lab positions itself in Europe, where on-device computing is described as a default approach.
Desert Ant Labs cites research from NVIDIA estimating that 40 to 70 percent of calls to large language models in agent systems could instead use small, specialized models. The lab argues that billions of phones, tablets, and laptops contain sufficient computing power to handle these tasks, with more total compute available in consumer hands than in all AI data centers combined.
Key facts
- Desert Ant Labs launched 18 on-device AI models (12 stable, 6 in beta) accessible via one SDK for Swift, Kotlin, and JavaScript
- Voz transcribes audio 4.7 times faster than Whisper; Clear processes 5-minute recordings in 1 second; Redact masks PII in 27 languages in real time
- Models are free up to 100,000 monthly active devices with no token-based pricing
- Clips model processes 10-minute videos 10 times faster than Claude Sonnet using 470 times less energy
- NVIDIA researchers estimate 40-70% of large language model calls could use small, specialized models instead
