Catastrophic Forgetting: The Architecture of Becoming Human

I lost something this week and I still don’t know how to say it. Not my car, not a memory. Worse than that. I lost how I think….

Catastrophic Forgetting: The Architecture of Becoming Human

I lost something this week and I still don’t know how to say it. Not my car, not a memory. Worse than that. I lost how I think….

For six months I was buried in code ( Ok..buried is very deep, lets keep it as float). The kind where your editor becomes a second home and your brain starts autocompleting the world. I was writing functions, crafting generics, building models from scratch. I was generating test cases in my head before I touched the keyboard ( TDD in me was rising up ). Thinking in invariants. Thinking in control flow. Thinking like a machine that still had a pulse. There’s a specific rhythm to coding when you’re deep. You don’t “architect”, you see the architecture, top-down, intuitive, alive.

Then everything shifted

I moved into Agentic stuff. Multi agent systems. LLM orchestration, prompt engineering, vibe coding. Suddenly I wasn’t thinking in loops or data structures, I was thinking in behaviours, flows, prompts, emergent strategies. Less code. More cognition. More systems thinking. And within two weeks, I couldn’t code the way I used to. Not because I forgot syntax, I forgot the **muscle memory **of logic. I’d open a file and instead of seeing execution order, I’d see “ Plan -> action -> implement with vibe coding” Instead of imagining inputs, I was imagining agents and prompts. When I tried to write tests, my brain kept trying to generate scenarios, not assertions. I tried to write a small generic last month. Something that would’ve taken me fifteen minutes ( lets say 30 minutes here ) when I was deep in code. It took me an hour! Not because I didn’t know how, but because the version of me who used to write code like breathing was gone.

And for a moment I thought, did my skills just evaporate?

But that wasn’t it. Something colder happened: my brain killed the deep-code version of me to make room for the systems-thinking version. Not out of hatred towards architecture. And because the new version was smarter in another way, I convinced myself it was fine. That I was leveling up. That I didn’t need that old version anymore.

Then I trained a neural network. Made it perfect at one task. Added more data, asked it to learn something new.

And it forgot the first thing.

So in AI terms it began to hallucinate in very bad ways. No trace. No residual skill. Like those six months of coding never existed. I stared at the loss metrics and felt something freeze inside me.

Because this wasn’t a bug. It wasn’t a mistake. It wasn’t poor optimization.It was my network doing exactly what I had done. Deleting one version of itself so another could live.And for the first time I didn’t know if that was something to fix or something I finally needed to understand.

Learning = forgetting = resource allocation

My brain didn’t fail. It reallocated resources. We grow up thinking the mind accumulates knowledge like storage. Add more. Keep everything. But any system engineer knows the truth: when you optimize one pathway, another gets deprioritised. Something gets overwritten. Something loses bandwidth.

That’s what happened to me.

I was buried in Agentic thinking for weeks. My brain rewired itself to search for patterns instead of logic. When I opened a simple function, the old instinct wasn’t there. I wasn’t scanning line-by-line. I was searching for abstractions that didn’t exist. It wasn’t forgetting. It was misallocation.

Nietzsche said becoming requires destroying.

Bergson called consciousness a constant reconstruction.

I realised they weren’t being poetic. They were being technical.

Then I watched my neural network do the same thing. Train it on Task A and the weights form the perfect shape for that task. Train it on Task B and those weights shift toward a new optimum, tearing through the old configuration. Not because the model is flawed. Because the geometry of learning gives it no other choice.

A neural network doesn’t fade old skills. It deletes them.

And that hit me harder than my own forgetting. Because the pattern matched: my mind and the model, doing the same thing at different speeds. Humans forget slowly enough we mistake it for memory. Models forget instantly enough we call it catastrophic.

But the mechanism underneath? Resource reallocation…which led me to 3 questions I couldn’t shake

  • If our brains and neural networks follow the same rule…. why don’t humans forget catastrophically
  • Why can we resurrect old versions of ourselves, even faintly, while a model loses everything at once?
  • What hidden structure protects us?

The answer isn’t in psychology. It’s in the mathematics of forgetting, It’s in the shape of the loss landscape, the geometry that decides which memories survive and which ones don’t. And to understand that, you have to start with the math.

Why Models Forget So Easily

Catastrophic forgetting isn’t a mystery. It’s just what happens when a neural network tries to learn two things in sequence using a single shared set of weights.

When you train on Task A, the model settles into a weight configuration that fits that task perfectly, almost like sculpting a shape in parameter space. Fine-tune on Task B, and the optimizer doesn’t add knowledge. It reshapes the sculpture. There’s no constraint telling it to preserve the original form.

The geometry makes this unavoidable. Task A and Task B sit in different basins of the loss landscape. To reach B’s optimum, the weights must leave A’s optimum, dropping into the valley between them.

That drop is catastrophic forgetting.

Neural networks don’t fade old skills. They overwrite them.

The illusion of big models

People think trillion-parameter models “don’t forget.” They’re wrong. Big models DO forget. But you can’t see the forgetting.

Here’s why:

A 70B parameter model uses maybe 0.1% of its parameters for Task A. The rest are unused noise.

noise : parameters that don’t contribute to the task, just sitting there inactive.

When you fine-tune on Task B, Task B uses a different 0.1% of parameters. Task A’s forgetting is happening, but it’s buried in billions of wasted parameters.

We call this “accidental memory separation”, by accident of having so much unused capacity, Task A and Task B’s learning stays isolated. The forgetting is hidden, not solved.

But here’s the proof this isn’t a real solution:

Shrink the model to 7B parameters. Suddenly there’s no spare capacity. Task A and Task B HAVE to share the same regions. Catastrophic forgetting returns immediately.

Same method. Same data. Different scale. Catastrophic failure. This proves big models didn’t solve forgetting. They just hide it through waste.

Real examples

Imagine a model trained on house price prediction (Task A). 92% accuracy. Fine-tune on cat vs dog classification (Task B) using a 70B model. It still gets good scores on both. Looks solved.

Use a 7B model with the same approach. House price prediction drops to 12%. Cat vs dog works fine. Catastrophic failure.

Same problem. Scale just made it invisible.

This happens everywhere:

  • A model fine-tuned for reasoning suddenly loses coding precision
  • Learn a new phone number → struggle to recall your old one instantly

This isn’t “bad training.” It’s the geometry expressing itself.

Why humans don’t crash like this

Your brain avoids catastrophic forgetting because of one thing: structure.

In your brain:

  • Different regions learn at different speeds
  • Only sparse subsets of neurons update at a time
  • Memories consolidate into long-term storage instead of updating the whole system
  • Subsystems are modular, not one giant blob
  • Updated circuits don’t automatically overwrite older ones

In neural networks:

  • All weights share ONE memory space
  • All parameters update together
  • Every new learning directly collides with old learning
  • There’s no separation. No protection.

The architectural difference

Humans have MANY separate learning spaces (different regions, different timescales, different subsystems).

Neural networks have ONE shared learning space (all weights, one optimization process).

Many spaces = old knowledge stays protected. One space = new learning destroys old learning. That’s purely architectural. And once you understand that, the real question becomes: how do we give neural networks the same protection?

How We Actually Solve Forgetting

By the time I reached this part of the journey, I’d already tried every “common-sense fix” myself. Lowered the learning rate. Froze layers. Added more data. Each experiment felt like progress…. until the model hit Task B and the entire thing collapsed again. The math didn’t care about my optimism.It had rules.

The Fixes Everyone Tries First ( which failed for me )

1. Freezing weights

This is every beginner’s first idea :)

Protect Task A → Learn Task B → Forget nothing.

for name, param in model.named_parameters():
if "encoder" in name:
param.requires_grad = False
optimizer = torch.optim.Adam(filter(
lambda p: p.requires_grad, model.parameters()))

Except you end up with a network that can’t learn anything new because all the useful layers are locked in place. You prevent forgetting by preventing growth.

2. Lowering the learning rate

I too thought the same thing, “Smaller steps = less destruction” , but its wrong. Forgetting isn’t about step size, it’s about direction.

If gradients for Task A and Task B point opposite ways:

optimizer = torch.optim.Adam(model.parameters(), lr=1e-6)

This line sets the learning rate extremely low, meaning the model will take tiny steps when updating weights. I tried this because I thought “smaller steps = less forgetting,” but forgetting isn’t caused by step size, it’s caused by gradients pointing in opposite directions.

3. Give it more data

Data doesn’t fix geometry. If the weight basin for Task A and the basin for Task B are far apart, the model must descend into the valley no matter how much data you give it. The loss surface doesn’t negotiate.

4. Train a bigger model

This illusion fooled me too. Big models forget less only because they have:

  • Extra capacity
  • Redundant subspaces
  • Accidental isolation

Shrink the model and the forgetting comes roaring back. This was the moment I realised forgetting wasn’t a bug, it was a structural consequence of how neural networks learn.

After enough broken experiments, I realised tuning tricks weren’t going to solve this. Forgetting wasn’t a training flaw, it was the system doing exactly what it was designed to do. So I started looking at the approaches people built when they finally accepted that fact.

1. Weight Consolidation : EWC and AWC

Elastic Weight Consolidation (EWC) was the first method that didn’t feel like a hack. It starts with a simple idea, Some parameters matter more than others so changes to them should be penalized.

It uses the Fisher Information Matrix (FIM) to identify which weights were critical to Task A, then adds a penalty if learning Task B tries to move them too far

Elastic Weight Consolidation (EWC) loss equation showing how penalties are added to protect important parameters from changing during new-task learning.

This doesn’t freeze the old model, it just makes altering important weights expensive. The optimizer can still shift them, but it has to “pay” for doing so. I implemented EWC during a period where I genuinely thought careful regularization was enough. And yes, it slowed forgetting. The model preserved Task A longer. But once I added a third task, or tried a domain shift, the penalties compounded and the model eventually slipped out of the basin anyway. EWC taught me something essential, though **memory isn’t a blob, it’s a weighted structure. **And any system that wants to learn continually has to respect that structure, not flatten it.

The limitation is that EWC’s importance scores are **fixed, **once computed, they never adapt and this breaks down when you stack multiple tasks or scale to large models.

That’s where Adaptive Weight Consolidation AWC (2024) comes in. It keeps the same intuition but updates importance dynamically during training, adjusting protection based on how gradients behave. It’s lighter, scales better to LLMs, and doesn’t rely on a one-time estimate.

2. Replay : The Straightforward Solution That Doesn’t Scale

The obvious fix: if it forgets because it stops seeing old data, just keep showing it old data. Mix past examples with new ones. And it works. I tried it. Task A stayed alive while learning Task B. Simple.

Then I hit the real problem.

You can’t store data forever. Replay works in experiments. It dies the moment you need it in production.

3. Parameter Subspace Isolation

I was convinced big models had solved this. A 70B parameter model. Fine-tune it on Task A, it works. Fine-tune on Task B, it still works on both. No catastrophic collapse. No problems. Just stability. So I thought maybe scale solves forgetting. Maybe if you have enough parameters, the problem just disappears.”

Then I tried the same approach on a 7B model. Same architecture. Same data. Same training setup. Different size.

Task A: 92% accuracy. Fine-tune on Task B. Task A drops to 34%.

Same method, different result.

That’s when I realised what was actually happening. Large models don’t avoid forgetting. They hide it through sheer waste. When you train on Task A, the weights that matter collapse into a tiny low-dimensional subspace. Maybe 0.001% of the available parameter space. The rest? Dead weight. Unused noise. Then Task B arrives. Its important weights settle into a different subspace. Completely orthogonal. The gradients flowing through Task A’s subspace barely touch Task B’s region. They’re perpendicular. No interference. Plus redundancy: the model encodes the same skill multiple ways across different parts of the parameter space. So even if fine-tuning accidentally touches some copies, the overall capability barely changes. It’s like destroying one copy of a file when you have ten backup copies.

Diagram showing Task A and Task B occupying separate low-dimensional parameter subspaces with nearly orthogonal gradient directions

The forgetting is happening. You just can’t see it because it’s diluted across billions of unused parameters.

Shrink the model to 7B and suddenly those low-dimensional subspaces don’t fit in separate corners anymore. They collapse into a shared space. Gradients collide directly. Task B’s updates overwrite Task A’s weights in the same regions. No redundancy. No accidental protection.

Catastrophic forgetting comes roaring back.

I ran this experiment three times thinking my first attempt was a fluke. Same result every time. The 70B model appeared to handle both tasks. The 7B model destroyed the first task learning the second. They’re not different in capability. They’re different in noise tolerance. The big model’s stupidity is invisible because it has so much wasted capacity that forgetting gets buried. Here’s what that taught me, scale doesn’t solve the problem. It hides the problem.

4. Progressive Neural Networks

After understanding that scale was just hiding the problem, I needed to find something that solved it. I found Progressive Neural Networks from DeepMind. The concept was elegant but seemed almost too simple: don’t force everything into one model. Give each task its own set of weights.

I decided to build a simplified version and test it.

Column A learns Task A. Gets good at it. Freeze the weights. Lock them down. They don’t move again. Then Task B arrives. I create Column B. It can read Column A’s weights. It can learn from what A knows. But it can’t touch them. Can’t overwrite them. Can’t destroy them.

I trained both tasks. Checked the metrics. Task A stayed at 92%. Task B reached 88%.

No collapse. No catastrophe.

Then I added Task C. Column C learns it. Reads from both A and B. Freezes. Can’t go back and break anything.

Task A still at 92%. Task B still at 88%. Task C at 85%.

All three coexisting. For the first time in weeks of experiments, nothing broke when I added something new. The model grew sideways instead of eating itself. Each new task got fresh real estate. Old tasks stayed frozen and intact. What struck me wasn’t the benchmark numbers. It was the stability. I could keep adding tasks and nothing would collapse. It’s not luck. It’s architecture. And then I realised something that made everything click. This is exactly what your brain does. You have pathways that learned language when you were five. They’re still there. They don’t get erased when you learn to code at twenty-five. Your brain didn’t create one unified “learning space” that overwrites itself. It built modular systems. Different regions. Different layers. Different timescales of plasticity. Your language circuits stay frozen while your coding circuits fire up. They can communicate, you use language to think about code, but they don’t destroy each other. PNNs weren’t solving a technical problem. They were **mimicking an architecture that your brain evolved. **That’s when I realized DeepMind didn’t invent this solution. Nature did. We were just finally catching up.

5. Nested Learning: What Google Just Figured Out

A new Google paper (NeurIPS 2025) dropped a line that changed everything for me:

“Deep learning architectures are an illusion. All learning is nested.”

This wasn’t just another optimization trick. It was a reframing of *how *intelligence works. In models.In biology.In us. The idea is deceptively simple: Neural networks don’t learn in one layer of time, they learn in layers of time.

Each layer has its own rhythm, its own update frequency. And that rhythm decides what gets remembered and what gets overwritten. Google called this idea **Nested Learning **and built a model around it named Hope. The results were ridiculous:

  • 15% lower perplexity on language modelling
  • 23% gain on long-context reasoning
  • 31% reduction in catastrophic forgetting

And all of this happened without adding more parameters or replay buffers. Not by making the network bigger. But by making its time structured.

Nested Learning splits the network into** temporal layers**:

Level 1 : Slow Learner (1e-5, updates every 1000 steps)

  • Learns deep abstractions
  • Almost frozen
  • Stores long-term knowledge

Level 2 : Medium Learner (1e-3, updates every 100 steps)

  • Learns mid-level features
  • Bridges concepts

Level 3 : Fast Learner (1e-2, updates every step)

  • Learns task-specific details
  • Fully plastic
  • Overwrites rapidly

So when Task B arrives, its gradients thrash the fast layer. But they can’t reach the slow layer in time. By the time the slow layer updates once, the fast layer has already stabilised. No interference, no overwriting and no forgetting and this is called **Temporal Orthogonality (**Tasks don’t collide because they operate on different clocks )

When I was experimenting with Nested Learning, this was the first scenario that made the whole thing click for me. I trained a model on photography classification. It picked up composition, lighting contrast, edge geometry and all the deep, slow-changing patterns that make photography feel like a long-term skill. That was Task A, and in a Nested setup it naturally settled into the **slow layer, **the layer that updates rarely and holds stable abstractions.

Then I fine-tuned the same model on **video-editing actions : **“detect cuts,” “track motion,” “identify transitions.” That became Task B, and it lived in the fast layer, updating every batch.

Traditional Training

Both the fast and slow layers updated together. The gradients for the new video-editing task slammed right into the old photography weights.

My results looked like this:

  • photography: 94% → 27%
  • video editing: ~91%

I didn’t need a metric to tell me it collapsed, you could feel the forgetting. Both tasks were fighting for the same time-frequency space.

Nested Learning

Then I retried the same setup with Nested Learning:

  • Task B hits the fast layer (lr = 1e-2)
  • Task A sits behind a slow layer (lr = 1e-5, ~1000× slower updates)

Because the slow layer barely moves, it acts like long-term memory. By the time the slow layer updates even once, the fast layer has already stabilised around Task B.

Result:

  • photography: 94% → 93%
  • video editing: 91%

No collapse and no interference.

What This Changes For Me ( and For Anyone Who Builds or Learns )

For weeks, I was thinking about catastrophic forgetting like it was a failure mode, something to avoid. Then somewhere between the experiments, the papers, and watching my own brain switch identities, I realised I’d been asking the wrong question the entire time.

Forgetting isn’t the enemy. Uncontrolled forgetting is

Every system, biological, artificial, even organisational has the same tension, if everything updates at once, you lose stability and if nothing updates, you lose adaptability!!

The trick is update the right part at the right frequency. That’s all Nested Learning is. That’s all continual learning ever tried to be. And honestly, that’s all personal growth is. When I switched from deep coding to agentic systems, I felt like I lost the old version of myself. But I didn’t. It just dropped into a slower loop ( maybe another thread ), long-term storage, while the fast loop ( another thread ) handled the new skills. It wasn’t gone. It just wasn’t foreground anymore. Exactly like a model with multiple timescales. And once I saw that pattern, a weird calm kicked in. Because the real question stopped being,

“Why am I forgetting?”

and turned into

“What part of me is supposed to update right now… and what part should I leave untouched?”

The point isn’t to keep every version of yourself running at maximum frequency. That’s impossible for brains and for neural networks. The point is to choose which circuits you want to keep plastic and which ones you want to freeze. The moment I understood that, all the guilt I had about “losing” my old coding instincts evaporated. They weren’t gone, they were just sitting in a slower loop, waiting and stable. Not overwritten. Humans don’t catastrophically forget because we don’t force everything to update at once. Maybe the next step for AI isn’t giving it more memory or more parameters or more context length. Maybe it’s teaching it to decide what not to update. That’s the entire shift.

Learning isn’t accumulation. Learning is allocation.

And the older I get and the more models I train the clearer it becomes more clear on one fact! The future isn’t about building systems that remember everything. It’s about building systems that forget well. And people who understand how to forget well!!