Topic
Builds
6 notes growing here.
Building the smallest version that can teach you something
A practical way to separate an experiment from a disguised production system.
Cache Poisoning
Yo!! A quick note before we start. I have yapped a lot about prompt caching in Part 1 , how it works, what cache hits and misses actually mean, and how to improve cache hit rates without making your API bill shoot the sky.
AI Does Multiplication Underneath. So Why Did Older Models Break at School Maths?
Like anyone who got pulled into the AI wave, I went through the whole thing. The amazement phase, the over reliance phase, the okay this thing just hallucinated phase.., and then the slow ohh okay I think I am starting to get what this actu
The Smallest Thing in PyTorch Opens Half the GPU Stack
We are living in a time when AI systems are introduced with the kind of language people used to reserve for moon missions.
What Happens When You Put “n” Billion Weights in Your RAM
I was in full vibe-coding mode with Headphones on. Letting Copilot autocomplete half my thoughts. Prompt here, tab there. Confidence at an all-time high. It honestly felt like I had rented extra IQ from the cloud ( having this kind of feel
Latency Diet for LLMs: Cutting Millisecond Fat Without Losing IQ
When I first started building my own model ( a language model for my experimental need ), I went through a lot of videos and built it from scratch. While doing so I noticed something interesting: Speed feels like intelligence. A model that