Alibaba released Qwen3.8-Flash-Next, a 125B sparse model that activates only 6B parameters per token. Computing all parameters for every word wastes energy. Sparse routing delivers top coding performance at $0.16 per million tokens with 8.6x faster throughput.