Architechural Masterclass in Kimi K3

A 2.8T parameter model is impressive!! But what interested me more was, how do you even run this thing? This is my small attempt at unpacking the architecture choices that make Kimi K3 actually survive in production.

In the ocean full of AI updates, its difficult to swim!! And I am barely surviving everyday with a life jacket. With even a jacket one day of a holiday, when I jus come back its like history to me.. I have heard the phrase time flies fast.. but seems its flying really faster!!

So.. all bout the blog.. post last week other than the word agent and token, one other word that kinda heard was Kimi K3.  I was like yet another model, nowadays there is no excitement like the 2024 era where hearing a new gpt model was exciting ( LOL… 2024 is jus 2 years ago !! nvm )

Kimi and its Logo
Kimi and its Logo

For an marketful of audience going behind giants like laude and open ai models.. To hear Kimi like open weight models is very astonishing

fyi

An open-weight model makes its trained model parameters available for people to download, host, fine-tune and experiment with. But open-weight and open-source are not always the same thing.The weights may be available while the complete training dataset, training code, filtering process and exact recipe used to build the model remain private.

Source :Kimi’s website
Source :Kimi’s website

Kimi K3 is a Mixture-of-Experts model with 2.8 trillion total parameters, out of which roughly 104 billion parameters are activated for each token. It also supports multimodal inputs and a context window of up to one million tokens.

Two point eight trillion parametersss!!!!

I am more interested in not how the model was created, but to the question

How do you actually run it?

Because making the weights publicly available is one thing. Finding somewhere to keep several terabytes of model weights — and then serving them without destroying your memory, network bandwidth, packets, lossess, all my computer network 101 and inference budget → is a completely different problem.

Kimi K3, in regular BF16 precision, would need more than 5 terabytes just to hold the raw parameters. I didnot include all the intricate data science terms like hidden layers, activations, attention states, communication buffers or the memory required while generating tokens!!!

So the architecture is not filled with complicated terms simply because researchers wanted to make the paper sound impressive, most of these choices exist because, without them, a model of this size would be painfully difficult to deploy!!

In the below dissection, I have made sure this entire stuff doesnot look like a research paper and would really give you a light feel but with a full flavour of tech depth!!

Do we really need all 2.8 trillion parameters every time?

I feel this is the first question everyone has, I should answer, and its thankfully, NO

Kimi K3 uses a Mixture-of-Experts, or MoE, architecture.

You can imagine like a company with hundreds of specialised teams. You may have people specialising in coding, mathematics, language, legal documents, visual understanding and several other areas. But every time a request arrives, you do not pull all 619 teams into the same meeting. A router looks at the token and chooses only a small subset of experts that are likely to be useful. In Kimi K3, only around 16 experts are selected for each token.

This is how the model can contain 2.8 trillion parameters while activating only about 104 billion of them at a time. It gives the model access to a massive overall capacity without paying the complete computation cost for every token.

But there is a problem with this as well, reducing computation, i.e using only these 104B does not automatically remove the infrastructure problem. These experts may be distributed across different GPUs or even different machines, once the router selects an expert, the token representation has to be moved to wherever that expert is running!!! When thousands of tokens are being processed together, this creates a huge amount of communication across the cluster. Its like taking your order from n different hotels..

So MoE saves computation but introduces another problem, which is TRAFFIC

Before we jump into the architechural dissection, I want to present a quick visualization on the whole picture, so that you guys who are reading can follow along in each detailed component!!

Okay! Fix to TRAFFIC —> LatentMoE

This line actually comes straight from the brilliant minds at NVIDIA Research, who introduced LatentMoe

“Sparsity ≠ cheap inference”

And that one

Source : Nvidia's Latent Moe architechure

line perfectly explains the problem we just landed on.

MoE saves FLOPs by activating only a few experts, but in real serving, the pain can simply move somewhere else, memory movement and communication between GPUs.

fyi FLOPs

FLOPs = Floating Point Operations. Basically, how much mathematical computation the model is doing. Fewer FLOPs usually means less compute work for the GPU.

So LatentMoE asks a pretty simple question —> Why send the entire huge token representation to the experts?

Instead we do Hidden representation → Compress → Smaller latent → Experts → Expand back

fyi Latent — You might hear this often data sciencyy term

A latent representation is basically a compressed internal representation of the same information. Maybe assume like, same meaning, fewer dimensions. Because data is represented in the form of. vectors and each vector as my physics mam says has magnitude and direction. And this vectors are visualized in n dimension.

The experts now work on this smaller representation, which means fewer bytes need to travel between GPUs.

So basically,

  1. MoE reduced the compute.
  2. LatentMoE reduced the traffic created by MoE.

And I really like this pattern in Kimi K3 where in when solving one bottleneck, expose the next one, then design around that too.

Okayyy, traffic somewhat handled unlike bangalore, but Kimi still claims 1 million tokens of context… my cliff hanger to the next section!!!

1M context…

Cool, 1 bottleneck somewhat handled. Now looking at the next number Kimi casually threw at us -> 1 million tokens of context.

And 1M context is starting to show up on model cards everywhere. Damn everywhere…so its not very amusing or surprising…

While generating text, a Transformer ( underneath Data sciency term ) normally stores the Key and Value representations of previous tokens so that it does not have to redo the attention computation from scratch every single time.

That’s the KV cache.

fyi: KV Cache

Think of it as the model keeping its previous attention notes open while writing the next token. More tokens means more notes, which means more memory.

The problem is that this cache does not sit still. As the context grows, the cache grows with it. Add more layers and more concurrent users, and suddenly this becomes a proper server-side memory problem.

At 1 million tokens, that problem gets very real. And Mr./Ms. Kimi K3’s answer is not one trick, It uses two different attention mechanisms together

Kimi Delta Attention (KDA) + Gated Multi-Head Latent Attention (MLA)

Kimi Delta Attention aka KDA

Kimi Delta Attention, aka KDA, takes a pretty different approach to memory.

Normal attention is basically a giant filing cabinet of every previous token. Each one files away its Key and Value, and when a new query shows up, the model goes rummaging back through the whole cabinet.

Useful, yes…but the cabinet just keeps growing. Forever.

KDA’s whole pitch —> what if we don’t keep the cabinet at all?

Instead, each KDA layer holds one fixed-size memory state. A new token doesn’t get filed away as a new entry — it uses its Key to figure out where in that memory to update, and its Value to figure out what to write there.

fyi Delta

This is where the name comes from (emdash) KDA checks what the memory already thinks for that key, compares it to the new value, and writes only the difference, the delta, back in. Not “append forever,” more like read what I know → spot what changed → update just that.

Normal attention: every page of the book, open on the table, forever.

KDA: a fixed-size whiteboard next to you instead. You keep reading, you erase what’s stopped mattering, you rewrite the important stuff over the same board, the book can hit a million pages whereas the whiteboard stays exactly the same size.

And nowww, obviously, the doubt shows up.If the whiteboard never gets bigger… what happens the one time I actually need some tiny detail from page 14,382?

Yeah. That’s the trade-off.

A fixed-size state is great for memory and speed. It just can’t behave exactly like having every single previous token sitting there individually, on demand.

Which is exactly why Kimi does not go full KDA.

MLA

KDA swaps the growing cache for a whiteboard. Multi-Head Latent Attention, or MLA, keeps the cabinet, just makes the cabinet way smaller.

MLA isn’t Moonshot’s own invention, it’s from DeepSeek. What it actually does

> instead of caching a separate Key and Value for every attention head, it compresses the Key and Value info for a token down into one shared latent vector. One small thing gets cached, not a pile of bigger ones per head.

When attention actually happens, the model expands that latent back out and does real, full attention with it -( em-dash intended ) any token can properly look at any other token, not through some approximated summary.

That is the actual difference from KDA. KDA compresses history into a fixed-size running state, which is an approximation by design, MLA compresses each entry into a smaller shape, but doesn’t approximate what is inside it, so it can still do exact global lookups when it’s asked to.

Kimi’s version specifically is Gated MLA ( em-dash intended ) it adds a learned gate on top that controls how much of the attention output actually passes through. Same spirit!!! as the gating KDA uses on its memory, just applied on the heavier, precise side of the pair instead of the lightweight side. One is extremely efficient for carrying information forward. The other is useful when the model actually needs to look back into the context with more precision.

Kimi mostly mixes them in a 3:1 pattern:

KDA → KDA → KDA → Gated MLA

Then again:

KDA → KDA → KDA → Gated MLA

The final architecture has 69 KDA layers and 24 Gated MLA layers.

And this 3:1 mix was not just pulled out of thin air either. Moonshot had already experimented with hybrid KDA and MLA ratios in their earlier Kimi Linear work before bringing that idea into K3 at a much bigger scale.

So now MoE helped with compute, LatentMoE helped with traffic, and KDA + MLA help with long-context memory!!!

Isnt this tooo nice…except..the model itself still weighs multiple terabytes 😭

Where do we keep 2.8T weights?

At 2.8 trillion parameters, storing everything in 16-bit precision already means we’re talking about

5+ terabytes of weights alone

before generating even one token. So the next obvious weapon is quantization.

>> Quantization basically means representing those numerical values using fewer bits. Instead of keeping every weight in 16-bit precision, we can move to something smaller like 8-bit or 4-bit. Fewer bits means less memory, less data movement, and potentially faster inference too. Of course, there’s a catch. There is ofcourseee a catch!!! The more aggressively you reduce precision, the more numerical information you potentially lose (em-dash intended) that s given anyway. If you train the entire model happily in high precision and suddenly squash everything down to 4-bit at the end, do not be surprised if the model gets a little upset.

Kimi K3 handles this using Quantization-Aware Training, or QAT. Instead of treating quantization as one final deployment hack, Moonshot introduces the target low-precision behaviour from the SFT stage onward. Concretely, this means the model actually trains while simulating the rounding it’ll face later weights get faked into low-precision buckets during the forward pass, so gradients start pushing the model toward values that survive that rounding cleanly instead of values that only work at full 16-bit precision. It’s not “train normally, then compress and hope,” the compression damage is happening during training too, so the model gets thousands of steps to route around it instead of eating it all at once at the end.

QAT lets you practise on the actual ground and Kimi K3 eventually uses MXFP4 weights, along with MXFP8 activations.

fyi MX weights

MX stands for microscaling —> an actual industry standard (OCP, backed by AMD, NVIDIA, Intel, Microsoft, Arm, Meta, and Qualcomm), not a Kimi-specific trick. The core idea is smtg like instead of giving an entire tensor one single scale factor, which gets wrecked the moment one weight in there is a huge outlier, MX chops the tensor into small blocks of 32 values and gives each block its own scale, so one outlier doesn’t ruin precision for everything around it. MXFP4 means each of those 32 values is stored in 4 bits, MXFP8 means 8 bits, small enough to actually shrink the model, structured enough to not fall apart numerically.

NVIDIA’s Blackwell chips and AMD’s newest accelerators have native hardware support for exactly this format, running MXFP4 at meaningfully higher throughput than MXFP8, so Moonshot isn’t picking a precision level in a vacuum, it’s picking the one number format the actual GPUs it’s serving on are built to run fast.

Training and deployment are no longer two completely separate worlds!!!!!!! The hardware a model will eventually run on is starting to influence how that model is trained in the first place.

But with all this compression… are we slowly deleting the brain?

At this point we have compressed something almost everywhere everywhere. Its been the word I used in every para

Experts are sparse. Representations move through latent spaces. Attention uses compact states. Weights are down to 4-bit.

At some point, who have made it this far!!!

Are we making the model efficient, or just making it efficiently confused?

This is where a couple of smaller architecture choices become interesting.

Attention Residuals: Don’t blindly carry everything forward

Some data science incoming!!!!

In a normal Transformer, residual connections basically take information from the previous layer and add it back into the current representation. Very very useful, but also pretty blunt..every layer just gets summed into one running total, no matter how relevant it actually is, and deeper in the network that sum can grow kind of unbounded and messy

Kimi K3 uses Attention Residuals, or AttnRes, and it is a genuinely different move, not just a tweak. Instead of summing, each layer gets its own learned “pseudo-query”, basically a little learned probe and uses it to run actual softmax attention over the outputs of earlier layers, deciding on the fly what’s actually worth pulling forward instead of getting everything by default.

Now here’s the catch, doing this fully where every layer attends to every single earlier layer, is brutally expensive to keep in memory at K3’s depth. So K3 uses the block version : : layers get chunked into 8 blocks of 12, normal residuals happen inside a block, and the fancy cross-attention-across-depth only happens between blocks. Basically the same “don’t check everything, check a cached summary” trick KDA and LatentMoE were already doing, just applied to depth this time instead of tokens or experts.

One detail I actually really liked, those pseudo-queries start at zero. Which means at the very start of training, AttnRes is just… doing normal averaging, behaving exactly like a boring residual connection. It only gradually learns to get picky about which layers to actually listen to as training goes on. Nothing breaks on day one, it just slowly earns the right to be selective.

NoPE: Wait… no positions?

Quick one. Because this one can spiral into another whole blog!!!!!

Most modern Transformers use something like RoPE, or Rotary Position Embeddings, to help the model understand where tokens occur inside a sequence.

Because these two sentences contain almost the same words:

The dog chased the cat.

The cat chased the dog.

But obviously, they do not mean the same thing. Order matters.

The interesting bit in Kimi K3 is that the MLA layers specifically use NoPE, no positional embeddings, not the whole model.

Now, that doesn’t mean Kimi suddenly has no clue where anything appears in the sequence 👀. Remember how KDA writes to its memory —> that gating/decay thing we talked about? That process is inherently order-sensitive: information written more recently has had less decay applied to it than something written way earlier. So by the time a representation reaches an MLA layer, it’s already been passed through KDA layers that quietly baked positional and recency info into it. RoPE in MLA would basically be redoing work that’s already done.

There’s a bonus too bro :: RoPE-based models often get weird once you push them past the context length they were trained on, needing rescaling hacks to cope. Since KDA’s position signal isn’t tied to any fixed formula the way RoPE’s rotation angles are, K3 sidesteps a chunk of that extrapolation pain by construction.

Basically, one mechanism is already quietly doing the position work. So the other doesn’t need to repeat it.

And I am like definitely stopping there before this accidentally becomes,

Architectural Masterclass in Positional Embeddings!!!!

So… Kimi K3 is basically a collection of hardware problems

Okay real talk, why did I even start writing about this?? Tens and thousands of models are dropping every single week now, I genuinely cannot keep up, and somehow this one made me stop scrolling anyway!!!!

Because once you stop reading Kimi K3 as a pile of fancy architecture names, it is basically just a conversation between the researchers and the infrastructure, going something like:

  1. 2.8T parameters every token? Sparse MoE.
  2. Expert routing flooding the network now? LatentMoE.
  3. Cool, want 1M tokens in memory too? KDA + Gated MLA.
  4. Where are we keeping 5+ TB of weights?? Quantization + QAT.
  5. Are we destroying the brain with all this compression AttnRes + hybrid attention.

None of it invented from scratch either, honestly…NVIDIA, DeepSeek, Moonshot’s own earlier Kimi Linear work, all stitched together around one question: how do you make a model this absurd actually survive in production.

2.8 trillion parameters is impressive, sure.

But the crazier part? They figured out how to run the damn thing. Iam still here trying to run my 7B model :(

Now the actual scary question (em-dash) what does “running” it cost you, specifically, if you tried to self-host this on your own GPUs instead of just hitting an API? That math is uglier than the architecture, maybe something on another post, probably!

References