Skip to content
LowResearchPreprintLLM-specific

Speedbumps: Rejection Attacks on Speculative Decoding

Published
Record updated
View JSON

Summary

Researchers study Speculative Rejection Attacks (SRAs), which append an adversarial suffix to attacker-controlled content so that draft and target models disagree more often in speculative decoding. Two attacks, Speedbump-P and Speedbump-D, optimise the expected length of the accepted speculative prefix, and in some cases slow inference below autoregressive decoding. The suffixes stay effective under sampling and transfer across drafters or target models sharing a drafter, showing the draft-target interaction is a realistic attack surface for inflating inference costs.