01Introduction

Feature Reuse in Transfer Learning

An interactive companion that builds the idea from the ground up — reuse what a model already knows, and train only what is new.

Group 23 Deep Learning · Transfer Learning
Team members
Laishram Khumanleima ChanuG25AIT2056
Suresh Babu GandlaG25AIT2119
Abhishek KumarG25AIT2004
Mahesh VG25AIT2058
02Transfer Learning

What is transfer learning?

Deep networks normally learn every weight from scratch, which demands huge labelled datasets and heavy compute. Transfer learning breaks that assumption: knowledge gained on a data-rich source task is carried to a related, data-poor target task — so the new model starts from competence instead of randomness.

\[ \mathcal{D}=\{\mathcal{X},\,P(X)\}, \qquad \mathcal{T}=\{\mathcal{Y},\,P(Y\mid X)\} \]
In plain words: a domain \(\mathcal{D}\) is the kind of inputs you see; a task \(\mathcal{T}\) is the mapping from inputs to answers. Transfer learning reuses knowledge across a change in the domain, the task, or both.

Domain & task

Transfer applies when source and target differ in their data, in what they predict, or in both — yet still share useful structure.

Why it matters

It slashes the labelled data and compute a new task needs, converges faster, and often lifts accuracy — decisive when target data is scarce.

Where feature reuse fits

The most practical form of transfer learning is feature reuse: keep a pretrained network’s features and adapt only what the new task requires — the focus of the rest of this companion.
03Overview

Feature reuse at a glance

A model trained once at great cost has already learned to see. Feature reuse transplants that learned perception into a new task — so you reuse what is known and train only what is new.

generic features (early) mid features specific features (late) new trainable head
A source model is pretrained on a large dataset (e.g. ImageNet, 1.2M images).

Figure 1. The pretrained extractor (left) is reused by a small target task (right); only the new head is trained.

Real-world example

A hospital wants to flag pneumonia in chest X-rays but has only ~800 labelled scans. Instead of training a network from scratch, it reuses a ResNet pretrained on ImageNet — the reused layers already detect edges, textures, and shapes — and trains only a small classifier on top. Strong accuracy, tiny dataset.

04The Reuse Flow

How a reused network turns input into a prediction

Pick a real task, then run a forward pass to watch an input become reused features and a prediction. Run backprop to watch the gradient update the head and stop at the frozen extractor.

Choose a task, then press Forward Pass.

Figure 2. The same architecture serves three real tasks — only the input, the meaning of each layer, and the output change.

\[ h(x)=g_\varphi\!\big(f_\theta(x)\big) \]
In plain words: a network is a feature extractor \(f_\theta\) (the “eyes’’ that turn an input into features) followed by a small head \(g_\varphi\) (the “verdict’’). Feature reuse keeps the pretrained eyes and trains mostly the verdict.

Prediction

Run a forward pass to see the model’s output for this example.
Illustrative outputs for the selected example — this page demonstrates the idea and does not run a live model.
\[ \text{frozen reuse: } \min_{\varphi}\ \tfrac1N\textstyle\sum_i \mathcal{L}\big(g_\varphi(f_{\theta_0}(x_i)),y_i\big) \qquad\quad \text{fine-tune: } \min_{\theta,\varphi}\ \tfrac1N\textstyle\sum_i \mathcal{L}(\cdot) + \tfrac{\lambda}{2}\lVert\theta-\theta_0\rVert^2 \]
In plain words: with the extractor frozen, only the head’s weights \(\varphi\) change — the reused features stay exactly as they were. When you fine-tune, all weights adapt, but the \(\lVert\theta-\theta_0\rVert^2\) term (enforced by a small learning rate) keeps them close to the pretrained values so the model doesn’t forget what it knew.
05Strategies

Five ways to reuse — choose what trains

Feature reuse is a spectrum, from reuse everything, train almost nothing to reuse only the starting point, retrain it all. Pick a strategy; press Run training to see gradients reach only the trainable layers.

Gradient updates flow only into the trainable (highlighted) layers.

Figure 3. Selecting a strategy re-freezes the stack; each comes with its objective, a plain-language gloss, and a real deployment.

06Implementation

The code, for both approaches

The same two ideas, in real PyTorch. Frozen feature extraction trains only a fresh head on top of a fixed backbone; fine-tuning adapts the whole backbone with a small learning rate. Switch between them and copy the snippet.


  
07Why It Works

Why reuse helps

Two animated intuitions: features run generic → specific with depth (so early layers transfer best), and pretrained weights start optimisation inside a good basin (so training converges far faster).

generic → specific
random vs pretrained init
\[ \varepsilon_{\text{gen}}\ \lesssim\ \sqrt{\mathcal{C}/N} \]
In plain words: the gap between training and test error shrinks when the trainable part is simpler (smaller \(\mathcal{C}\)) and the dataset is larger (bigger \(N\)). Reusing a frozen extractor slashes \(\mathcal{C}\), so even a tiny dataset generalises — that is why reuse fights overfitting.

Universal early features

First-layer filters resemble Gabor edges and colour blobs across datasets — a generic visual alphabet that transfers almost perfectly.

Better starting point

Pretrained weights begin inside a favourable basin of the loss surface, so convergence needs far fewer steps.

Built-in regulariser

Fixing the extractor shrinks the model’s freedom to overfit, turning the prior into a powerful regulariser on small data.
08Training Dynamics

From scratch vs. transfer

Press Run training to race two models on the same task: one from random weights, one reusing a pretrained extractor. Watch the transfer model start lower and converge faster, epoch by epoch.

epoch 0 · ready

Figure 4. A schematic of the 800-X-ray pneumonia task: transfer (teal) starts lower and converges faster than training from scratch (grey).

09Limitations

When reuse hurts

Reuse is a bet on relatedness. Toggle the source domain and press Transfer: a related source dips the error below the from-scratch baseline (positive transfer); an unrelated one can rise above it (negative transfer).

Lower is better. Compare the transfer curve to the baseline.

Figure 5. Lower is better. Reuse wins only when source and target genuinely share structure.

Real-world example of negative transfer

Reusing an ImageNet photo model for audio spectrograms or highly abstract medical scans can backfire: the source’s natural-image statistics don’t match, so the reused features mislead the model. Studies of medical imaging (e.g. Transfusion, Raghu et al.) found that very large ImageNet models often gave little or no benefit on domain-distant tasks — sometimes a small model trained from scratch did just as well.

Negative transferAn unrelated source can do worse than random initialisation.
Domain mismatchFeatures tuned to source statistics may not fit the target.
Architecture lock-inYou inherit the input size, depth, and design.
Inherited biasSource-dataset blind spots carry into your task.
Resource overheadLarge backbones are heavy to run at inference.
Catastrophic forgettingToo high a learning rate erases the reused features.