FreeToken Reworks Local MoE Inference for Consumer GPUs
Researchers from UC Berkeley and MIT have open-sourced FreeToken, an inference engine designed to reduce the cost of moving sparse MoE weights between system memory and consumer GPUs. The paper reports about 39 tokens per second for Qwen3.6-35B on a laptop RTX 4060.
Read more