Shortlist MTP Mostly Worked Around a Bad Kernel Choice
Cutting Qwen's draft vocabulary from 248,320 tokens to 16,000 looked like a big throughput win. After I fixed the kernel selection, most of that win disappeared.
Read article ->Writing
Long-form research, technical writing, and detailed project notes.
Cutting Qwen's draft vocabulary from 248,320 tokens to 16,000 looked like a big throughput win. After I fixed the kernel selection, most of that win disappeared.
Read article ->I found several ways to skip a meaningful amount of Qwen3.5's work without changing its next token very often. The actual speculative decoder was still MUCH slower than normal decoding.
Read article ->