Prompt Caching

Okay so real talk before we start….

Prompt Caching

Indeed generated using ChatGPT

Okay so real talk before we start….

There was this phase I went through, and honestly I am not ashamed of it because I think most people who got into building with LLMs went through the exact same thing, where I was just asking the model everything. Like everything.

Bro can you code this → bro this function is broken can you debug → cool now write the tests → now write the docs → now refactor this whole file while you are at it!!

I asked it to do git commits for me. I asked it to explain stack traces I did not feel like reading. I was asking it to write emails I was too tired to write ( lazy is the right word ). There was one point, I am not making this up, where I literally typed can you open Netflix inside VS Code so I do not have to switch windows..

Anyway, the point is I was using these chat UIs constantly. Left right centre. No hesitation. And it felt free; it genuinely felt like this thing that just existed and you could ask it whatever and it would answer and there was no meter running anywhere.

Then I started Building. Not just chatting → actually building. Agents, workflows, a support bot, some RAG pipelines. Started calling the API properly. Deployed things. Got a little too confident.

Came back a week later and looked at the API bill.

….yeah.

It was not great. I had been sending the same 8,000-token context on every single request; the system prompt, the tool definitions, the reference docs…. fresh. From scratch. Every time. The model was re-reading the same instructions on every request while I was blissfully assuming ~ yeah the API probably handles that.~

Somehow prompt became a household word before most of us understood what it actually costs. And that is where the trap starts….

In the chat interface a prompt feels free. You type more context. Paste a document. Add a few examples. Throw in one more instruction because the model was acting slightly possessed yesterday and you are not taking chances. So you just… keep adding stuff. Because why not → it is not like it costs anything, right?

wait what.

In the API world, the world where you are building actual products, actual agents, actual workflows; every extra token has a price tag.

Your system prompt + conversation history has a price tag.

And the part I genuinely find humourous is this → if you send the same 8,000-token system prompt 5,000 times a day, the model does not go:

relax dood. I already read this yesterday.

So that is the whole setup → the user asked 6 tokens. The API charged 8,006 input tokens. And it will do it again on the next request, and the one after that; same system prompt, same tools, full price, every single time. This is not a bug. It is just… how it works. And nobody really talks about it in plain language because it gets buried under documentation and pricing pages that assume you already know what prefill means.

The Hidden tax

There are things about LLM inference that changes how you think about cost.

The model does two very different jobs on every single request.

**Job 1 → Reading (Prefill) : **Before the model writes a single word, it has to process your entire input; all of it. The system prompt + the tools + the docs + the history + the question. This is called prefill. It runs in parallel across all your input tokens, it is compute-heavy, and it scales directly with how long your input is.

Job 2 → Writing (Decode) : Then it generates your response, one token at a time. This is called decode. In a sequential manner btw, each token depends on the one before it. Cheaper per token than prefill.

TTFT — time to first token — is that cursor-blinking pause between you hitting send and the first word appearing. That gap is almost entirely prefill. The model is not being dramatic; it genuinely cannot start writing until it has finished reading. The moment prefill ends, decode kicks in and tokens start flowing.

So a slow TTFT usually means one thing: a long prompt. And a high API bill usually means one thing: a long prompt being sent repeatedly.

Here is the kicker….

The expensive part is not answering the question. The expensive part is reading everything before the question.

So when your 6,000-token system prompt gets processed fresh on every request,

6,000 tokens × 5,000 requests × $3.00 per million × every single day

Just because nobody told it that it already did this work yesterday. And the day before. And every day since you deployed this thing. And here is the thing that should make you slightly annoyed when you realise it….

The model is not learning anything new by re-reading your system prompt on request number 4,999. It is just doing the same computation it did on request number 1. Computing the same internal numbers, storing them in memory for the duration of that one request, then throwing them away when the response is done. So the next request starts again from zero. Full read. Full compute. Full price.

That is the waste. Right there. That is all it is.

So the obvious question is → can the provider just… not throw those numbers away? Can it keep them around and reuse them when request 4,999 arrives with the same prefix as request 4,998?

Yes. That is prompt caching.