Kimi K3 — Playable Architecture

one screen · four moves

How Kimi K3 stays runnable

Pick a pressure point to see the architecture choice that answers it.

Compute

Move compute, not every expert

Sparse MoE activates only the useful experts. LatentMoE then makes the traffic between those experts smaller.

Token → route a few experts → compressed traffic

2.8T parameters~104B active/token1M context