Strata is an open-source MIT-licensed inference engine, built partly on llama.cpp/ggml, that runs the 125B-parameter Qwen3.8-Flash-Next model on a gaming PC with just 12 GB VRAM and 64 GB RAM, hitting 80–94 tokens/s on an RTX 5070. It achieves this through a mixture-of-experts architecture (24,576 small experts, 10 active per token), tiered GPU/RAM/CPU/SSD memory, aggressive 2-4 bit GGUF quantization, and speculative decoding. It exposes OpenAI- and Anthropic-compatible local APIs, a browser chat app, and an MCP server, and has passed 16,000 GitHub stars. Limitations include slow model loading, single-request serving by default, and large download sizes, making it best suited for individual offline development rather than team deployment.

•4m read time•From fireup.pro
Post cover image
Table of contents
What is Strata?How does a 125B model fit on a 12 GB GPU?Hardware requirementsHow fast is it?Which variant should you pick?Limitations worth knowingWhy developers will care?

Questions this post answers

How can I run a 125B parameter model like Qwen3.8-Flash-Next on a 12 GB GPU?

Strata, an MIT-licensed inference engine built partly on llama.cpp/ggml, makes this possible through a mixture-of-experts architecture where only 10 of 24,576 experts activate per token, tiered memory across GPU, RAM, CPU and SSD, 2-4 bit GGUF quantization variants, and speculative decoding. On an RTX 5070 with 12 GB VRAM and 64 GB RAM, it reaches 80-94 tokens per second. daily.dev surfaces practical setups like this for developers testing local LLM inference on consumer hardware.

Which Strata quantization variant should I use for running Qwen3.8-Flash-Next locally?

The choice depends on available RAM: 32 GB RAM suits the Coder variant (half the experts removed but keeps 91% of the full model's SWE-bench Verified score), 48 GB suits IQ2_XS or Q2_0, 64 GB suits IQ2_XS as the sweet spot or IQ3_XXS/IQ3_S, and 96 GB+ suits IQ3_S or UD-IQ4_XS for results closest to the full model. Developers weighing local model variants against hardware constraints can track these comparisons on daily.dev.

What are the limitations of running Strata's local LLM setup on a gaming PC?

Loading the model takes 1-3 minutes and 35-55 GB of RAM, potentially freezing the PC, the download is about 70 GB, and by default only one request is served at a time, making it suited for a personal workstation rather than a team server. The first message in a long chat processes slowly, around 1 minute per 30,000 tokens, and AMD cards can't read images yet on Windows. daily.dev helps developers weighing local-first AI setups stay aware of real-world constraints like these before committing.

Share this post