TRAINSPOTTER run diagnostic
source
examples/throughput_drop.trainer_state.json
format
hf
steps
0–299
generated
2026-10-01 17:38 UTC
0error
3warning
0info

train/loss / eval/loss

train/losseval/loss0120100200

lr

00.20.40100200

grad_norm

0120100200

eval/accuracy

00.510100200

step_time

00.010.020100200

Findings log

WARN steps 2 grad_norm

Gradient-norm explosion

grad_norm hit 1.466 at step 2, 14.0x the local MAD scale above its rolling median.

peak_step=2   peak_value=1.466   robust_z=13.97

  • Enable or lower gradient clipping (e.g. max-norm 1.0).
  • Lower the learning rate.
  • Check the batch at this step for an outlier example (e.g. a corrupted/extreme-length sample).
WARN steps 204 spikes

Loss spike

train/loss jumped to 0.4233 at step 204, 6.9x the local median-absolute-deviation scale.

peak_step=204   peak_value=0.4233   robust_z=6.911   window=21   threshold=6

  • Lower the learning rate or add/extend LR warmup.
  • Clip gradients (or lower the existing max-norm).
  • Check the batch at this step for a data or tokenization bug (e.g. an unmasked pad token, a corrupted shard).
  • If it self-recovers within a few steps and doesn't recur, it may be benign -- confirm against the divergence detector's verdict.
WARN steps 240–299 throughput

Throughput regression

Median step time over the last 60 points (8.841e-05s) is 1.40x the early-run baseline (6.31e-05s, first 60 points), derived from step_time.

baseline_median_s=6.31e-05   recent_median_s=8.84e-05   ratio=1.401   window=60   source=step_time

  • Check for a dataloader/IO bottleneck (e.g. a slow shard, network storage stall, growing prefetch queue).
  • Check for memory pressure/fragmentation causing more frequent allocator work or CPU-GPU sync.
  • Check other processes on the same GPU/host, or a thermal/power throttling event.
  • If a sequence-length or batch-size curriculum increases work per step by design, this may be expected -- not a regression.