Shortlist MTP Mostly Worked Around a Bad Kernel Choice
Cutting Qwen's draft vocabulary from 248,320 tokens to 16,000 looked like a big throughput win. After I fixed the kernel selection, most of that win disappeared.
Read article ->Research & Projects
A personal notebook for research, technical explorations, and the projects I'm building.
Writing
Cutting Qwen's draft vocabulary from 248,320 tokens to 16,000 looked like a big throughput win. After I fixed the kernel selection, most of that win disappeared.
Read article ->I found several ways to skip a meaningful amount of Qwen3.5's work without changing its next token very often. The actual speculative decoder was still MUCH slower than normal decoding.
Read article ->In progress
Loading the latest stream entries...
Elsewhere