TRAINSPOTTER run diagnostic
source
examples/missing_warmup.trainer_state.json
format
hf
steps
0–249
generated
2026-10-01 17:38 UTC
0error
3warning
0info

train/loss / eval/loss

train/losseval/loss012050100150200

lr

00.20.4050100150200

grad_norm

00.511.52050100150200

eval/accuracy

00.51050100150200

step_time

00.00010.00020.00030501001502000.000442

Findings log

WARN steps 0 lr_schedule

No LR warmup detected

lr starts at 0.5, already 100% of the run's peak (0.5) -- no ramp-up phase visible in the log.

initial_lr=0.5   peak_lr=0.5

  • Add LR warmup (e.g. linear warmup over the first 1-10% of steps) -- it's a common fix for early spikes/instability.
  • If warmup is configured but not logged until after it finishes, this may be a logging artifact -- check the logging interval.
WARN steps 203 spikes

Loss spike

train/loss jumped to 0.5146 at step 203, 7.4x the local median-absolute-deviation scale.

peak_step=203   peak_value=0.5146   robust_z=7.419   window=21   threshold=6

  • Lower the learning rate or add/extend LR warmup.
  • Clip gradients (or lower the existing max-norm).
  • Check the batch at this step for a data or tokenization bug (e.g. an unmasked pad token, a corrupted shard).
  • If it self-recovers within a few steps and doesn't recur, it may be benign -- confirm against the divergence detector's verdict.
WARN steps 230–249 overfitting

Overfitting onset

Best eval/loss was 0.146 at step 230; it's since risen to 0.1941 (32.9%, slope p=0.0245).

best_step=230   best_value=0.146   current_value=0.1941   relative_rise=0.3293   slope_p_value=0.02446

  • Use the checkpoint at step 230 (best eval/loss), not the last one.
  • Add or increase regularization: weight decay, dropout, data augmentation.
  • Reduce model capacity or train for fewer steps / add early stopping.
  • Get more training data, or de-duplicate against the eval set if overlap is possible.