An interactive companion that builds the idea from the ground up — reuse what a model already knows, and train only what is new.
Deep networks normally learn every weight from scratch, which demands huge labelled datasets and heavy compute. Transfer learning breaks that assumption: knowledge gained on a data-rich source task is carried to a related, data-poor target task — so the new model starts from competence instead of randomness.
A model trained once at great cost has already learned to see. Feature reuse transplants that learned perception into a new task — so you reuse what is known and train only what is new.
Figure 1. The pretrained extractor (left) is reused by a small target task (right); only the new head is trained.
A hospital wants to flag pneumonia in chest X-rays but has only ~800 labelled scans. Instead of training a network from scratch, it reuses a ResNet pretrained on ImageNet — the reused layers already detect edges, textures, and shapes — and trains only a small classifier on top. Strong accuracy, tiny dataset.
Pick a real task, then run a forward pass to watch an input become reused features and a prediction. Run backprop to watch the gradient update the head and stop at the frozen extractor.
Figure 2. The same architecture serves three real tasks — only the input, the meaning of each layer, and the output change.
Feature reuse is a spectrum, from reuse everything, train almost nothing to reuse only the starting point, retrain it all. Pick a strategy; press Run training to see gradients reach only the trainable layers.
Figure 3. Selecting a strategy re-freezes the stack; each comes with its objective, a plain-language gloss, and a real deployment.
The same two ideas, in real PyTorch. Frozen feature extraction trains only a fresh head on top of a fixed backbone; fine-tuning adapts the whole backbone with a small learning rate. Switch between them and copy the snippet.
Two animated intuitions: features run generic → specific with depth (so early layers transfer best), and pretrained weights start optimisation inside a good basin (so training converges far faster).
Press Run training to race two models on the same task: one from random weights, one reusing a pretrained extractor. Watch the transfer model start lower and converge faster, epoch by epoch.
Figure 4. A schematic of the 800-X-ray pneumonia task: transfer (teal) starts lower and converges faster than training from scratch (grey).
Reuse is a bet on relatedness. Toggle the source domain and press Transfer: a related source dips the error below the from-scratch baseline (positive transfer); an unrelated one can rise above it (negative transfer).
Figure 5. Lower is better. Reuse wins only when source and target genuinely share structure.
Reusing an ImageNet photo model for audio spectrograms or highly abstract medical scans can backfire: the source’s natural-image statistics don’t match, so the reused features mislead the model. Studies of medical imaging (e.g. Transfusion, Raghu et al.) found that very large ImageNet models often gave little or no benefit on domain-distant tasks — sometimes a small model trained from scratch did just as well.