FreeToken Shows Why Local AI Is a Serving Problem, Not Just a VRAM Problem
FreeToken turns CPUs, GPUs, host memory and PCIe into one adaptive inference system, arguing that frontier-scale open models can run locally when the serving stack is designed around heterogeneous consumer hardware.

Open-weight AI has a strange accessibility problem. The model files can be free to download while the hardware required to run them still looks like a datacenter purchase. FreeToken is an attempt to close that gap by changing the serving system rather than shrinking the model.
The Apache-2.0 project, released by a research team that includes Berkeley systems researchers alongside collaborators from MIT and UT Austin, is designed for frontier-scale Mixture-of-Experts models on personal hardware. Its repository has already passed 8,000 GitHub stars, a sign that the idea has moved quickly beyond a paper-only audience.
The headline numbers are striking: the paper reports Qwen3.6-35B-A3B at 39.3 tokens per second on an 8GB RTX 4060 laptop, DeepSeek-V4-Flash at interactive speed on a gaming desktop, and the 753B-parameter GLM-5.2 on a single RTX PRO 6000 workstation GPU. But the important idea is not that FreeToken somehow fits hundreds of billions of parameters into one graphics card. It is that it stops treating VRAM as the boundary of the inference system.
Local AI has been framed as a memory-capacity problem
Most local-model discussions start with a simple question: does the checkpoint fit in GPU memory? That framing works reasonably well for dense models, where most parameters participate in every token. It becomes less useful for large Mixture-of-Experts models.
An MoE model may contain hundreds of experts while activating only a small subset for each token. FreeToken uses DeepSeek-V4-Flash as an example: only a fraction of its 284B parameters participate in a single decoding step. The compute needed per token can therefore fit inside a consumer GPU even when the complete expert pool cannot.
The difficulty is moving the right expert weights to the right processor quickly enough. If every cache miss stalls while weights cross PCIe, the GPU sits idle. If all missing experts run on the CPU, host-memory bandwidth becomes the bottleneck. Static placement also breaks down because expert routing changes from token to token and from workload to workload.
FreeToken’s thesis is that consumer hardware already contains enough aggregate resources to serve much larger models than VRAM capacity suggests. The missing layer is software that can coordinate GPU compute, CPU compute, host RAM and the CPU-GPU interconnect as one elastic system.
The GPU becomes a cache, not the whole machine
FreeToken keeps the complete routed-expert pool in host memory as the source of truth. Non-expert weights remain on the GPU, while the remaining VRAM becomes a shared expert cache. Recently useful experts can stay resident; misses are handled dynamically.
The key mechanism is a bandwidth-adaptive policy the paper calls q*. FreeToken measures the machine’s PCIe bandwidth and host-memory bandwidth, then divides expert-cache misses between two paths: transfer some experts to the GPU and compute others directly on the CPU. The split is chosen for the machine actually running the model rather than hard-coded for a generic desktop.
That matters because a laptop with an RTX 4060, a desktop with an RTX 5090 and a workstation with an RTX PRO 6000 have very different balances between VRAM, PCIe throughput, CPU memory bandwidth and available system RAM. FreeToken treats those differences as scheduling inputs instead of compatibility failures.
Prefill and decode need different strategies
MoE sparsity helps during token generation, but long prompts create a different problem. Across thousands of prompt tokens, routing can touch nearly every expert in a layer, making the effective working set much denser during prefill.
FreeToken responds with full-layer double buffering. While the GPU computes the current layer, the next layer’s experts stream over PCIe in the background. The goal is to overlap data movement with useful computation instead of exposing every transfer as idle time.
It also adds semantic-aware state caching for agent workloads. Tool calls and reasoning systems frequently edit or truncate context between turns, which can invalidate ordinary cached states and force expensive re-prefill. FreeToken anchors checkpoints around semantic boundaries so the engine can reuse surviving state and recompute only the changed suffix when possible.
This is one reason the project is more interesting than a benchmark-oriented local inference demo. Its design assumes the model will sit behind real coding and tool-using agents, where context changes repeatedly and latency compounds across many turns.
FreeToken is trying to make frontier open weights operational
The project currently supports more than 20 MoE models and model families, including DeepSeek-V4, GLM-5.2, Qwen3.6 MoE, GPT-OSS, MiniMax M2.5 and others. It exposes both OpenAI-compatible and Anthropic-compatible APIs, which means existing clients can point at a local FreeToken server instead of being rewritten around a new interface.
The CLI goes further by providing launch paths for coding and agent clients including Claude Code, Codex, OpenCode, OpenClaw, Hermes and DeepSeek Harness. That makes the project’s target clear: not merely proving that a giant checkpoint can produce tokens on a desktop, but making that model usable as local infrastructure for existing agent workflows.
In that sense, FreeToken addresses a different layer of openness than model licensing. Open weights determine whether users can obtain the model. A serving engine determines whether they can practically run it. The gap between those two has become increasingly important as frontier open models grow into hundreds of billions of parameters.
The 753B headline comes with a large asterisk: host RAM
FreeToken does not eliminate memory requirements. It relocates much of them. The complete MoE expert pool still lives in system memory, so running a large model requires enough free RAM to hold those weights at the chosen precision.
The maintainers give a concrete example: Qwen3.6-35B-A3B in BF16 needs roughly 70GB of free host RAM, although lower-precision checkpoints can reduce that requirement substantially. A workstation serving a 753B model is therefore still a serious machine even if it uses only one GPU.
Current compatibility is also narrower than the excitement around the project may imply. Official support is x86_64 with NVIDIA Ampere-or-newer GPUs, an r580-or-newer driver and CUDA 13. Windows and Linux are supported through the desktop application, while the Python installation path is Linux-focused. macOS, AMD GPUs and aarch64 systems such as DGX Spark remain roadmap items.
Model-format support is evolving too. FreeToken loads Hugging Face safetensors directly, while native GGUF support is currently limited rather than universal. Multimodal checkpoints are served text-only. These are normal constraints for a young systems project, but they matter if FreeToken is compared with mature local runtimes such as llama.cpp or Ollama.
The more important shift is from local models to local systems
Local AI has often been treated as a model-compression race: quantize harder, shrink the checkpoint and accept that the strongest models remain cloud-only. FreeToken points toward a second path. Sparse model architectures can remain large if the serving layer becomes good enough at exploiting every resource already present in the machine.
That does not make a gaming PC equivalent to a datacenter. It changes the boundary of what a personal machine can do interactively. A model that was previously impractical because its weights exceeded VRAM may become usable if most of those weights can live in RAM while the system moves or computes only what each token needs.
For open-source AI, that distinction matters. The next accessibility breakthrough may not come from another smaller model. It may come from better infrastructure that lets users run larger open models on hardware they already own.
FreeToken’s strongest claim is therefore not that everyone can run a 753B model at home. It is that the local inference stack has been leaving performance on the table by treating the GPU as the machine instead of treating the machine as the system.
Sources and further reading
- FreeToken repository and README - FlashML / GitHub
- FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution - arXiv
- FreeToken supported models and MoE backends - FlashML / GitHub
- FreeToken quick start - FlashML / GitHub
- FreeToken installation requirements - FlashML / GitHub
- FreeToken maintainer FAQ - FlashML / GitHub