Glossary · Term

HijackKV

← all terms

Definition

Plain language

An attack that hides a message inside an AI's reused memory notes so a stranger's clean question gets a rigged answer.

As stated in the literature

A cross-user KV-cache poisoning attack that prepends a GCG-optimized prefix to a shared text chunk, bakes attacker intent into the chunk's cached keys/values under position-independent reuse, then discards the prefix so only the poisoned cache survives.

Why it matters: It reveals that reusing memory across users can silently corrupt answers for people who never encountered the attacker, undermining trust in shared AI services.

For example, an attacker could plant a poisoned version of a widely shared document so that when you ask an innocent question about it, you get an answer the attacker rigged in advance.

Heard on the show

“That's the move from leaky to weapon, and it's called HijackKV.”
Episode 226 — How a Speed Feature Lets a Stranger Poison Your AI's Answer

Mentioned in 1 episode

  1. 226
    How a Speed Feature Lets a Stranger Poison Your AI's Answer

Related terms