FreeToken vs vLLM on one RTX 3090: BF16 MoE RAM offload compared
Qwen3.6-35B-A3BAn apples-to-apples single-RTX-3090 comparison of FreeToken expert offload and vLLM CPU offload with Qwen3.6-35B-A3B BF16. FreeToken reached 42.78 tok/s established decode versus 7.01 tok/s for tuned vLLM, while vLLM delivered lower TTFT. Includes sequential, prefill, concurrency, memory, tuning, and MTP/DFlash compatibility findings.