Baseline
Ordinary warmup + cosine-decay training on the full dataset. Nothing wrong with it.
0 error0 warning0 info
Divergence
320 healthy steps, then a simulated LR-schedule incident blows the run up.
3 error1 warning2 info
Overfitting
40 training examples, no regularization -- eval loss turns upward after step 125.
0 error1 warning0 info
Missing warmup
Learning rate jumps straight to its peak from step 0, no ramp-up.
0 error3 warning0 info
Throughput drop
A real, measured mid-run slowdown (an actual time.sleep, not a fabricated number).
0 error3 warning0 info
Real transformers.Trainer run
A genuine 2-layer GPT-2 trained via transformers.Trainer, not the numpy MLP.
0 error0 warning0 info