Magnitude, a Y Combinator S25 startup, has launched an open-source inference engine designed to run large language models locally on users’ devices with significant performance improvements.

According to the project’s GitHub repository, Magnitude achieves up to 2x faster inference than llama.cpp, with benchmarks showing 92% faster decode performance on Apple Silicon Metal and 19% faster on NVIDIA CUDA hardware. The engine uses a hardware-specific optimization approach: rather than relying solely on precompiled kernels, it compiles and tunes kernels directly on the user’s device before running models.
The tool ships as a desktop application available for macOS, Windows, and Linux. Users download the app, select a recommended model from the Discover section, and connect it to their existing AI agent—supporting Pi, OpenCode, Hermes, Codex, and other tools through a single-click integration. The OpenAI-compatible API enables compatibility with additional agents.
Magnitude is designed to work across a wide range of hardware configurations. It supports Apple Silicon, NVIDIA GPUs, AMD GPUs, and CPU-only systems. According to the documentation, there is no fixed minimum hardware requirement; smaller machines can run smaller models, with available memory determining maximum model size.
The engine includes several optimization features. It reduces memory usage by 27% per agent and frees memory when agents stop running. Multiple concurrent sessions can share prefix caches to prevent performance degradation. Magnitude uses hand-optimized kernels for popular open-weight model families rather than taking a generalist approach.
All processing occurs locally on the user’s machine. Prompts, files, and models remain private and never leave the device. The tool requires an internet connection only for the initial model download.
Magnitude is released under the Apache 2.0 open-source license and is free to use with no token costs. The project encourages community support through GitHub stars and provides access to a full list of compatible models on magnitude.dev/models.
Key facts
- Magnitude claims up to 2x faster inference than llama.cpp, with 92% speedup on Apple Silicon and 19% on NVIDIA CUDA
- The engine optimizes kernels directly on users’ hardware before running models
- Reduces memory usage by 27% per agent and supports Apple Silicon, NVIDIA, AMD GPUs, and CPU-only systems
- All processing is local and private; compatible with Pi, OpenCode, Hermes, Codex, and other agents
- Released as free, open-source software under Apache 2.0 license
