Halooo, I’m Akilesh.
I am usually building something, breaking something..... or poking at interesting model architectures just to see what happens. Right now I’m deep into inference serving, distributed training and MLOps. I write about what I learn so I don’t forget it lol, and so AI feels a little less confusing for everyone else.
Where latency, memory and batching collide.
Notes from inside the machine
Things I needed to understand properly.
Ideas, experiments and technical rabbit holes, written down before I forget what made them interesting.
- 01Architechural Masterclass in Kimi K3
- 02What a benchmark cannot tell you
- 03Building the smallest version that can teach you something
- 04Cache Poisoning
- 05Prompt Caching
- 06AI Does Multiplication Underneath. So Why Did Older Models Break at School Maths?
- 07The Smallest Thing in PyTorch Opens Half the GPU Stack
- 08What Happens When You Put “n” Billion Weights in Your RAM
- 09Reading Minds Is the New Logging
- 10The Physics of Prompting: Why Words Don’t Matter, But Structure Does
A few loose pages