The Spectrum Dispatch News

technology

Running Qwen 3.8 Flash Next on Consumer GPUs with Strata

The open‑source Strata framework lets users run the 125‑billion‑parameter Qwen3.8‑Flash‑Next model on RTX 5070, RTX 3090 and similar cards, reporting token‑generation rates of 60‑1

Running Qwen 3.8 Flash Next on Consumer GPUs with Strata

Strata is an open‑source installer that enables a 125‑billion‑parameter AI model, Qwen3.8‑Flash‑Next, to run on a typical gaming PC. According to the project’s documentation, the software works with NVIDIA or AMD graphics cards that have at least 12 GB of VRAM, 32 GB or more of system RAM, about 80 GB of free disk space (preferably on an SSD), and Windows 10/11 or a current Linux distribution with up‑to‑date graphics drivers.

Running Qwen 3.8 Flash Next on Consumer GPUs with Strata

Once installed, Strata provides a local API endpoint at http://127.0.0.1:8080 that behaves like an OpenAI‑compatible service, allowing chat, code generation, image understanding and integration with coding agents such as Claude Code, Cursor or GitHub Copilot. All processing stays on the user’s machine; no data leaves the PC.

Performance numbers shared by the project give a sense of what to expect on consumer hardware. For a short chat, the model’s reply generation speed is described as “60 tokens per second is faster than you can read.” When measuring prompt ingestion, the source notes that a 32 K‑token document is processed at a rate that allows follow‑up messages to start in seconds after the initial load.

The documentation also provides a concrete example of scaling with VRAM: “A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100‑140 tokens per second.” This indicates that models with larger memory budgets can achieve higher token‑throughput. The same tables show that an RTX 5070 (12 GB) paired with a Ryzen 5 7600 and 64 GB of RAM delivers comparable results, though exact figures are not listed in the excerpt.

Strata supports multi‑GPU setups, allowing two or three cards to share the model workload. The installer automatically detects the graphics card, selects an appropriate engine version, and downloads the required model files (approximately 70 GB for the full‑size version). Users can choose among various quantized sizes—such as Coder, IQ2_XS, IQ3_S or Unsloth UD‑IQ4_XS—to match their available RAM and VRAM; smaller sizes run faster while larger ones retain more capability.

Setup is straightforward: on Windows, double‑click START‑HERE.bat; on Linux, run ./setup.sh inside the Strata folder. The script prompts for model size, context length and image support, then downloads the model (resuming if interrupted) and launches the web interface. Subsequent starts are near‑instant because the model files are already present.

The project emphasizes that the first launch may occupy 35‑55 GB of RAM and temporarily lock part of it for the GPU, which can make the system sluggish for one to three minutes. After that, the Strata window shows real‑time GPU, CPU and RAM usage, and users can stop the model by closing the window or run update scripts to pull newer versions without restarting the service.

By bringing a model that normally requires server‑grade hardware to a consumer‑grade PC, Strata demonstrates how recent quantization and engine optimizations make large language models accessible to enthusiasts, developers and researchers without relying on cloud services.

Key facts

  • Strata enables local execution of Qwen3.8‑Flash‑Next (125 B parameters) on PCs with ≥12 GB VRAM and ≥32 GB RAM.
  • Measured token‑generation rates: ~60 tokens/s for chat replies; RTX 3090 (24 GB) reported 100‑140 tokens/s.
  • Installation supports Windows 10/11 and Linux, with optional multi‑GPU sharing and SSD‑accelerated startup.
  • All processing remains on the user’s machine; data does not leave the PC.

Sources

← All posts