Radar
← All projectsTuesday, August 25, 2026
FreeToken
Run frontier-scale MoE models on a gaming PC — edge-native serving with agent-aware cache so tool edits do not recompute the whole context.
Why it matters
FreeToken is an open edge MoE engine from FlashML aimed at 290B+ class models on laptops and gaming desktops. It treats GPU, CPU, and host memory as one elastic pool, streams experts with bandwidth-adaptive co-execution, and exposes Anthropic/OpenAI-compatible APIs so tools like Claude Code, Codex, and OpenCode can point at your machine. There is a desktop app plus a CLI, and a 2026 paper on the runtime design.
Coding agents burn tokens rewriting long contexts after every tool call. FreeToken is built for that loop: it serves huge open MoE weights on consumer GPUs and keeps semantic checkpoints so agentic edits can reuse work instead of restarting the prefill tax.
How it works
Install the desktop build from flashml.ai or `uv pip install "freetoken[accel]"`, load a supported MoE checkpoint (DeepSeek, Qwen, GLM and friends across MXFP4/NVFP4/FP8/BF16), and hit the local OpenAI- or Anthropic-shaped endpoint from your agent. Expert caches and KV memory can rebalance at runtime without a full reload; semantic anchor checkpoints cut redundant recomputation when the agent mutates tool or thinking blocks mid-run.
Most local stacks stop at “fit the model.” FreeToken is a full MoE serving layer with elastic VRAM policy and caches tuned for agent turns, not just chat demos — so the win is interactive frontier weights on hardware you already own, wired to real coding agents.
Capabilities
- Public demo available
- API / SDK surface
Similar tools