LowResearchPreprintLLM-specific
Speedbumps: Rejection Attacks on Speculative Decoding
- Published
- Record updated
Summary
Researchers study Speculative Rejection Attacks (SRAs), which append an adversarial suffix to attacker-controlled content so that draft and target models disagree more often in speculative decoding. Two attacks, Speedbump-P and Speedbump-D, optimise the expected length of the accepted speculative prefix, and in some cases slow inference below autoregressive decoding. The suffixes stay effective under sampling and transfer across drafters or target models sharing a drafter, showing the draft-target interaction is a realistic attack surface for inflating inference costs.
Related items
- HighHermes Agent - Pre-Authentication Disk ConsumptionSimilar attack · Tenable Research Advisories
- MediumHermes Agent - Pre-Authentication Memory ExhaustionSimilar attack · Tenable Research Advisories
- MediumGHSA-v36g-jcw9-x7cw: Pydantic AI: Excessive resource use when local web fetching converts nested HTMLSimilar attack · GitHub Advisory Database
- MediumGHSA-v2xh-2vp8-57h8: Pydantic AI: Unbounded memory use when downloading remote content via web_fetch or FileUrlSimilar attack · GitHub Advisory Database
- MediumGHSA-fpf4-vwcp-v4hp: Pydantic AI: Event loop blocked by quadratic title extraction in `web_fetch`Similar attack · GitHub Advisory Database