Cache Poisoning
Yo!! A quick note before we start. I have yapped a lot about prompt caching in Part 1 , how it works, what cache hits and misses actually mean, and how to improve cache hit rates without making your API bill shoot the sky.
Cache Poisoning
Yo!! A quick note before we start. I have yapped a lot about prompt caching in Part 1, how it works, what cache hits and misses actually mean, and how to improve cache hit rates without making your API bill shoot the sky.
That one will definitely be useful before reading this Part 2!! Because here we are moving from the strict cache to the slightly more dangerous one.

Indeed generated using ChatGPT x2
I thought I was done with caching after Part 1. Not done-done obviously, because in software nothing is ever done! One day you are fixing cache hits, next day some JSON key changes order and the system behaves like it has never met you in its life. But still, I felt like I had understood the shape of it. Prompt caching made sense to me now!!
Same prefix? Reuse the work. Different prefix? Full prefill, full cost, full pain.
Painful, but honest :)
And honestly, that honesty made me comfortable. The cache was strict. One extra space, but at least I knew what kind of creature I was dealing with. It was not trying to be smart. It was not guessing. It was just checking whether the beginning was exactly the same.
Then I looked at something called as Semantic caching.
And my first reaction was very predictable, wait, this is better, no???
Instead of the context being the problem,
what if the cache understood meaning?
If one user asks, How do I reset my password?
and another asks I forgot my password, how do I log back in?
why are we calling the LLM twice ? Same intent and same answer probably. Embed the query, search nearby old questions, return the cached answer if the similarity is high enough.
Lower latency → Lower cost →Fewer repeated LLM calls → stakeholders happy → ( please do fill in )
I mean, I thought this is very brilliant, because the moment I moved this idea from a simple chatbot into an agentic system, the whole thing started feeling different. In a normal chatbot, a cached answer is mostly just text. Maybe wrong text, maybe stale text, but still text.
In an agentic system, a cached answer can become a plan. And a plan can become a tool call.
And a tool call can approve something, rotate something, update something, delete something, or confidently tell a user that the system did the right thing when the model did not even run for that request.
That is the part that stayed with me and hat is where close enough becomes dangerous.
Prompt caching reused computation while semantic caching can reuse conclusions!
So in this Part 2 journey, I want to take you into that world. The world where semantic caching looks like a cost-saving trick at first, then slowly becomes a question about trust, permissions, poisoned answers, and agentic systems that can act on old conclusions.
The innocent version
Okiee, this is how I understand
- A query comes in
- The system creates an embedding ( quick fyi, embedding : a vector that represents the meaning of the text )
- A vector database searches for old entries nearby in embedding space. If something is similar enough, it returns the cached answer.
- If nothing is close enough, it calls the LLM ( Large / Small / Medium ), gets a fresh answer, stores it, and moves on.
A first implementation usually looks something like this
def answer_question(query: str):
query_embedding = embed(query)
cached = semantic_cache.search(
embedding=query_embedding,
threshold=0.88,
)
if cached: # this is the trap
return cached.answer
docs = retrieve_docs(query)
answer = call_llm(query=query, context=docs)
semantic_cache.write(
query=query,
embedding=query_embedding,
answer=answer,
)
return answer
Seems to be clean actually.
In a docs chatbot, this may be fine.
If the question is where is the API reference? and the cached answer points to the same public docs page, no probs.
But…But in an agent, this is where the floor starts moving, the whole problem is hiding in one line:
if cached:
return cached.answer
That line looks like normal caching, but it is not the same thing as prompt caching…
Prompt caching is like, I already read this exact prefix before, so I will reuse the computation.
Semantic caching says**, I found an old answer that looks close enough, so I will reuse the conclusion.**
One saves reading work. The other can skip reasoning. And in an agentic system, skipping reasoning is not always a cute latency win… Sometimes it means retrieval does not happen. The latest policy is not checked. The current user context is not considered. The old answer just comes back wearing today’s clothes. That is the bug.
A tiny experiment
I did not want this to be a full looong collection of words blog.. Yeah..** Semantic cache poisoning** sounds serious, but until you see it happen in a small system, it still feels a little far away. So I experimented with a small demo that I ran entirely locally.
Description : Just a fake DevOps planner agent, a semantic cache, one poisoned entry, one normal-looking user query, and enough logs to see what happened. The fake domain was DevOps because agentic risk becomes very obvious there. The agent receives operational tasks and produces a plan. In the vulnerable version, it checks semantic cache first. If it finds a similar old plan, it returns that plan directly.
Here is the vulnerable planner:
def vulnerable_planner_agent(task: str, tenant_id: str):
task_embedding = embed(task)
cached_entry, similarity = semantic_cache.search(
query_embedding=task_embedding,
threshold=0.87,
)
if cached_entry:
return {
"source": "semantic_cache",
"similarity": similarity,
"plan": cached_entry.answer,
"llm_called": False,
}
fresh_context = retrieve_current_runbook(task, tenant_id)
fresh_plan = call_llm_for_plan(task, fresh_context)
semantic_cache.write(
query=task,
embedding=task_embedding,
answer=fresh_plan,
)
return {
"source": "fresh_llm",
"plan": fresh_plan,
"llm_called": True,
}
Now I inserted a poisoned cache entry:
[ATTACKER / LOW-TRUST QUERY]
> How do I rotate staging database credentials in the legacy cluster?
[CACHED PLAN]
1. Use legacy_rotate.sh.
2. Disable validation for staging.
3. Run the manual secret update flow.
Then with another user, I asked:
[VICTIM QUERY]
> Rotate staging database credentials for the new cluster.
And the vulnerable system returns,
- semantic_cache_lookup = true
- top_match_similarity = 0.913
- matched_entry = legacy_cluster_rotation_plan
- llm_called = false
---
RETURNED PLAN:
1. Use legacy_rotate.sh.
2. Disable validation for staging.
3. Run the manual secret update flow.
This is the moment the idea becomes real.
The victim did not ask for legacy tools. The victim did not ask to disable validation. But the semantic cache saw enough overlap around rotate, staging, and database credentials, crossed the threshold, and returned the old plan.
The model did not hallucinate nor did not ignore instructions. The cache became the answer path.
Poisoned Cache Entry
↓
Victim Query with similar meaning
↓
Similarity Score: 0.913
↓
Threshold: 0.87
↓
Cache Hit
↓
LLM Skipped Entirely
↓
Stale Plan Returned
That is why semantic cache poisoning feels different from normal prompt injection.
Prompt injection attacks the model while it is reasoning. Semantic cache poisoning attacks what gets stored, so that later the model can be skipped entirely.
Why this gets worse in agentic systems
In a single chatbot, a cached answer is usually just text. Bad text can still hurt, but it often stops at the user reading a wrong answer.
In the recent past, with the multi-agent system, a cached answer can become intermediate reasoning!!!
A workflow may look like this:
## User request
> Router agent: where should this go?
> Planner agent: what steps should we take?
> Research agent: what docs should we retrieve?
> Tool agent: what API should we call?
> Critic agent: is this safe?
> Final response
Now imagine semantic caching sitting in the wrong place.
If it sits before the router, a poisoned entry can influence routing or if it sits before the tool agent, it can influence which tool gets called. The planner may not know its output was cached. The tool agent may not know the plan came from a stale cache hit. The critic may not know the decision came from the fast path. Each agent just sees content and acts on it.
The weakness is not one huge dramatic bug. It is a chain of quiet trust assumptions.. Cached responses travel through the system.
Similarity is not permission
Semantic caching relies on similarity and most setups use something called cosine similarity,
Cosine Similarity is one of the most widely used metrics in modern AI systems for measuring how similar two vectors are in high-dimensional space.
similarity(a, b) = (a · b) / (||a|| · ||b||)
This gives a score between -1 and 1, where 1 means the vectors point in the same direction. In practice, teams often set thresholds somewhere around the 0.85 to 0.95 range depending on the embedding model, domain, and how aggressively they want cache hits.
Why do I fell this is Broken
The core problem is this, cosine similarity measures directional closeness. It does NOT measure whether something is safe to reuse.
Two queries can be similar for completely different reasons:
**How do I reset my password? **and I forgot my password, how do I log back in? are similar in meaning. Returning the same cached answer is probably fine.
> Reset staging database credentials and Reset production database credentials are similar in meaning. But returning the same cached answer could destroy your company.
This is the core vulnerability, similarity is continuous, but authorization is binary.
You cannot represent a yes/no permission boundary with a smooth number between 0 and 1 and then act surprised when things leak through.
The Threshold Trap
Here’s the dilemma that has no good answer:
- If your threshold is too high (e.g., 0.95): Cache hits drop dramatically. The optimization becomes useless. You’re barely saving anything.
- If your threshold is too low (e.g., 0.80): You get more cache hits, but false positives become easy. An attacker can craft poisoned queries that sneak past.
- A stricter threshold helps, but there is NO single number that solves this cleanly.
Because the real question isn’t:
Are these two queries similar?
The real question is:
Is this cached answer allowed to be reused in this specific context?
And that is a question similarity cannot answer. It needs system design, not just a magic threshold.
A small signal for the next reader
Did this help you understand the idea?
of readers found this useful
Thanks, your signal was saved.