Findings log
ERR
steps 320
spikes
Loss spike
eval/loss jumped to 54.88 at step 320, 4273.9x the local median-absolute-deviation scale.
peak_step=320 peak_value=54.88 robust_z=4.27e+03 window=21 threshold=6
- Lower the learning rate or add/extend LR warmup.
- Clip gradients (or lower the existing max-norm).
- Check the batch at this step for a data or tokenization bug (e.g. an unmasked pad token, a corrupted shard).
- If it self-recovers within a few steps and doesn't recur, it may be benign -- confirm against the divergence detector's verdict.
ERR
steps 321–322
spikes
Loss spike
train/loss jumped to inf at step 322, far beyond the local median-absolute-deviation scale (it's non-finite).
peak_step=322 peak_value=inf robust_z=inf window=21 threshold=6
- Lower the learning rate or add/extend LR warmup.
- Clip gradients (or lower the existing max-norm).
- Check the batch at this step for a data or tokenization bug (e.g. an unmasked pad token, a corrupted shard).
- If it self-recovers within a few steps and doesn't recur, it may be benign -- confirm against the divergence detector's verdict.
ERR
steps 322
divergence
Diverged (non-finite values)
train/loss became Inf at step 322 (1 non-finite point(s) in this run).
first_bad_step=322 count=1
- Lower the learning rate; this is the single most common cause.
- Enable gradient clipping if it isn't already on.
- Check for a divide-by-zero or log(0) in the loss (e.g. an unstable custom loss term, an empty batch after filtering).
- If using mixed precision, check the loss-scaler didn't overflow, or try bf16 instead of fp16.
WARN
steps 321
grad_norm
Gradient-norm explosion
grad_norm hit 23.02 at step 321, 183.4x the local MAD scale above its rolling median.
peak_step=321 peak_value=23.02 robust_z=183.4
- Enable or lower gradient clipping (e.g. max-norm 1.0).
- Lower the learning rate.
- Check the batch at this step for an outlier example (e.g. a corrupted/extreme-length sample).
INFO
steps 290–320
overfitting
Overfitting onset
Best eval/loss was 0.183 at step 290; it's since risen to 54.88 (300x its minimum, slope p=0.0833). (Downgraded: overlaps a spikes finding at step 320 -- this eval rise looks explained by that divergence, not by ordinary overfitting.)
best_step=290 best_value=0.183 current_value=54.88 relative_rise=298.9 slope_p_value=0.08326
- Use the checkpoint at step 290 (best eval/loss), not the last one.
- Add or increase regularization: weight decay, dropout, data augmentation.
- Reduce model capacity or train for fewer steps / add early stopping.
- Get more training data, or de-duplicate against the eval set if overlap is possible.
INFO
steps 319–320
lr_schedule
LR discontinuity
lr jumped from 1.097e-05 to 30 between step 319 and 320 -- 100% of the run's whole LR range (30) in a single step, out of line with the logged intervals around it.
before=1.1e-05 after=30 jump_fraction_of_range=1
- Check for a checkpoint resume where the scheduler's step counter didn't match the optimizer's.
- Check for a manual LR override mid-run.