Your Million-Token Pipeline Is Bleeding Speed You Don’t Know You’re Losing
Speculative decoding was supposed to make long-context inference fast and cheap. At a million tokens, it’s quietly doing the opposite — and the models you’re already shipping have this bug baked in. One researcher just found a training-free fix that cuts per-decode-step cost by up to 44%. Read that number twice.
