DwarfStar 4 (ds4) is a specialized inference engine designed to run large language models locally on machines with substantial RAM, according to the project documentation. Created by the Redis team, the engine targets high-memory Mac, CUDA, and ROCm systems rather than attempting generic model deployment.

The engine supports specific model families including DeepSeek V4 and V4.1 Flash, GLM 5.x, and Qwen 3.8 Flash, with capabilities for both text and vision models. According to the documentation, ds4 is not a generic GGUF runner but instead “follows a small, opportunistic set of model families and validates each supported layout end to end.”
The core innovation is asymmetric 2-bit quantization, which compresses routed experts while preserving critical shared paths in mixture-of-expertise models. This approach allows large models like DeepSeek V4 Flash to fit within the constraints of high-memory machines. The system also implements SSD-based prefix caching, allowing the engine to save long context prefixes and resume by prompt hash, avoiding full re-prefill on restarts.
According to benchmark data provided, on an M5 Max with 128GB RAM running q2 quantization at 2,048 tokens context, the engine achieves 790.2 tokens/second for prefill and 39.4 tokens/second for generation. At longer context lengths of 65,536 tokens, performance drops to 398.5 tokens/second prefill and 27.6 tokens/second generation on the same hardware.
The project offers three interfaces: a CLI for chat, a local API server compatible with OpenAI and Anthropic-style APIs, and a native agent for persistent coding sessions, all sharing the same model state and cache. The software is released under the MIT license and requires users to download model weights, build for their specific backend, and run locally.
According to the documentation, DeepSeek V4 Flash Q2 is listed as the baseline supported configuration, with GLM 5.3 Q2 and Qwen Q4 also fitting on 128GB systems, while V4.1 Q2 can stream from SSD on the same hardware.
Key facts
- DwarfStar 4 is a C-based inference engine optimized for running large language models locally on high-memory machines
- Supported models include DeepSeek V4/V4.1 Flash, GLM 5.x, and Qwen 3.8 Flash with text and vision capabilities
- Uses asymmetric 2-bit quantization to compress routed experts while preserving critical shared paths
- On M5 Max 128GB hardware, achieves 790.2 tokens/second prefill and 39.4 tokens/second generation at 2K context
- Includes CLI, HTTP APIs, and native agent interfaces all sharing the same model state and cache
- SSD-based prefix caching avoids full re-prefill on restarts
- Compatible with Mac (Metal backend), CUDA, and ROCm platforms
- Released under MIT license
