The Spectrum Dispatch News

technology

Whistle: 16.9 MB speech-to-text model runs on-device across 7 languages

Cactus Compute releases a compact speech recognition model for mobile, wearable, and IoT devices that performs transcription, word timing, and speech embedding without cloud

Whistle: 16.9 MB speech-to-text model runs on-device across 7 languages

Cactus Compute has released Whistle, a speech recognition model packaged as a single 16.9 MB file that runs locally on device hardware including mobiles, wearables, robots, smart home systems, and microcontrollers. The model operates on CPU without external dependencies and integrates with the same C++ engine used by the company’s Needle model.

Whistle: 16.9 MB speech-to-text model runs on-device across 7 languages

Whistle performs three tasks entirely on-device: transcription of 16 kHz mono audio up to 30 seconds in English, German, French, Spanish, Italian, Dutch, and Polish with automatic language detection; word-level timestamps including start, end, and probability scores; and speech embedding generation at 80 millisecond frame intervals without requiring full transcript decoding.

The model’s architecture uses eight Simple Attention blocks in its encoder with mHC residual connections and Monarch Hadamard MLP layers. The decoder employs eight Laddered Simple Attention blocks at 512-channel width with gated cross-attention to the encoder. Decoding uses five-beam search with length-normalized log probability scoring and supports keyword biasing through Aho-Corasick automaton matching to prioritize specified phrases.

According to the source, Whistle achieves better word error rates than OpenAI’s Whisper base model (145.3 MB) on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22, and FLEURS average benchmarks. Whisper base performs better on TED-LIUM, AMI, and MLS average tests. Whistle’s time to first token scales with clip length—5.9 milliseconds for 5-second audio, 11.1 milliseconds for 10 seconds, and 36.3 milliseconds for 30 seconds.

The engine supports multiple deployment targets including macOS, Linux, Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, browsers, and WASI components. Model weights are available on Hugging Face, with the engine and source code in the Cactus-Compute/needle3 repository on GitHub. The API includes three core functions: needle_load for loading models, needle_transcribe for transcription, and needle_embed for speech embeddings. Configuration options include forcing a specific language, enabling word timestamps, and providing keyword lists for biasing.

Key facts

  • Whistle is a 16.9 MB speech recognition model that runs on-device CPU with no external dependencies
  • Supports 7 languages: English, German, French, Spanish, Italian, Dutch, and Polish with automatic language detection
  • Performs three tasks: transcription (up to 30 seconds), word-level timestamps, and speech embeddings
  • Outperforms OpenAI’s Whisper base model (145.3 MB) on multiple benchmarks including LibriSpeech and SPGISpeech
  • Deploys to 17 platforms including mobile (iOS, Android), wearables (watchOS), and edge devices (RISC-V, MIPS)

Sources

← All posts