WARN
steps 2
grad_norm
Gradient-norm explosion
grad_norm hit 1.466 at step 2, 14.0x the local MAD scale above its rolling median.
peak_step=2 peak_value=1.466 robust_z=13.97
- Enable or lower gradient clipping (e.g. max-norm 1.0).
- Lower the learning rate.
- Check the batch at this step for an outlier example (e.g. a corrupted/extreme-length sample).
WARN
steps 204
spikes
Loss spike
train/loss jumped to 0.4233 at step 204, 6.9x the local median-absolute-deviation scale.
peak_step=204 peak_value=0.4233 robust_z=6.911 window=21 threshold=6
- Lower the learning rate or add/extend LR warmup.
- Clip gradients (or lower the existing max-norm).
- Check the batch at this step for a data or tokenization bug (e.g. an unmasked pad token, a corrupted shard).
- If it self-recovers within a few steps and doesn't recur, it may be benign -- confirm against the divergence detector's verdict.
WARN
steps 240–299
throughput
Throughput regression
Median step time over the last 60 points (8.841e-05s) is 1.40x the early-run baseline (6.31e-05s, first 60 points), derived from step_time.
baseline_median_s=6.31e-05 recent_median_s=8.84e-05 ratio=1.401 window=60 source=step_time
- Check for a dataloader/IO bottleneck (e.g. a slow shard, network storage stall, growing prefetch queue).
- Check for memory pressure/fragmentation causing more frequent allocator work or CPU-GPU sync.
- Check other processes on the same GPU/host, or a thermal/power throttling event.
- If a sequence-length or batch-size curriculum increases work per step by design, this may be expected -- not a regression.