TRAINSPOTTER run diagnostic
source
examples/divergence.trainer_state.json
format
hf
steps
0–322
generated
2026-10-01 17:38 UTC
3error
1warning
2info

train/loss / eval/loss

train/losseval/loss012340100200300✕54.88 / 55.49

lr

00.20.40.6010020030030

grad_norm

01230100200300✕23.02

eval/accuracy

00.510100200300

step_time

05e-050.00010.000150100200300

Findings log

ERR steps 320 spikes

Loss spike

eval/loss jumped to 54.88 at step 320, 4273.9x the local median-absolute-deviation scale.

peak_step=320   peak_value=54.88   robust_z=4.27e+03   window=21   threshold=6

  • Lower the learning rate or add/extend LR warmup.
  • Clip gradients (or lower the existing max-norm).
  • Check the batch at this step for a data or tokenization bug (e.g. an unmasked pad token, a corrupted shard).
  • If it self-recovers within a few steps and doesn't recur, it may be benign -- confirm against the divergence detector's verdict.
ERR steps 321–322 spikes

Loss spike

train/loss jumped to inf at step 322, far beyond the local median-absolute-deviation scale (it's non-finite).

peak_step=322   peak_value=inf   robust_z=inf   window=21   threshold=6

  • Lower the learning rate or add/extend LR warmup.
  • Clip gradients (or lower the existing max-norm).
  • Check the batch at this step for a data or tokenization bug (e.g. an unmasked pad token, a corrupted shard).
  • If it self-recovers within a few steps and doesn't recur, it may be benign -- confirm against the divergence detector's verdict.
ERR steps 322 divergence

Diverged (non-finite values)

train/loss became Inf at step 322 (1 non-finite point(s) in this run).

first_bad_step=322   count=1

  • Lower the learning rate; this is the single most common cause.
  • Enable gradient clipping if it isn't already on.
  • Check for a divide-by-zero or log(0) in the loss (e.g. an unstable custom loss term, an empty batch after filtering).
  • If using mixed precision, check the loss-scaler didn't overflow, or try bf16 instead of fp16.
WARN steps 321 grad_norm

Gradient-norm explosion

grad_norm hit 23.02 at step 321, 183.4x the local MAD scale above its rolling median.

peak_step=321   peak_value=23.02   robust_z=183.4

  • Enable or lower gradient clipping (e.g. max-norm 1.0).
  • Lower the learning rate.
  • Check the batch at this step for an outlier example (e.g. a corrupted/extreme-length sample).
INFO steps 290–320 overfitting

Overfitting onset

Best eval/loss was 0.183 at step 290; it's since risen to 54.88 (300x its minimum, slope p=0.0833). (Downgraded: overlaps a spikes finding at step 320 -- this eval rise looks explained by that divergence, not by ordinary overfitting.)

best_step=290   best_value=0.183   current_value=54.88   relative_rise=298.9   slope_p_value=0.08326

  • Use the checkpoint at step 290 (best eval/loss), not the last one.
  • Add or increase regularization: weight decay, dropout, data augmentation.
  • Reduce model capacity or train for fewer steps / add early stopping.
  • Get more training data, or de-duplicate against the eval set if overlap is possible.
INFO steps 319–320 lr_schedule

LR discontinuity

lr jumped from 1.097e-05 to 30 between step 319 and 320 -- 100% of the run's whole LR range (30) in a single step, out of line with the logged intervals around it.

before=1.1e-05   after=30   jump_fraction_of_range=1

  • Check for a checkpoint resume where the scheduler's step counter didn't match the optimizer's.
  • Check for a manual LR override mid-run.