WARN
steps 0
lr_schedule
No LR warmup detected
lr starts at 0.5, already 100% of the run's peak (0.5) -- no ramp-up phase visible in the log.
initial_lr=0.5 peak_lr=0.5
- Add LR warmup (e.g. linear warmup over the first 1-10% of steps) -- it's a common fix for early spikes/instability.
- If warmup is configured but not logged until after it finishes, this may be a logging artifact -- check the logging interval.
WARN
steps 203
spikes
Loss spike
train/loss jumped to 0.5146 at step 203, 7.4x the local median-absolute-deviation scale.
peak_step=203 peak_value=0.5146 robust_z=7.419 window=21 threshold=6
- Lower the learning rate or add/extend LR warmup.
- Clip gradients (or lower the existing max-norm).
- Check the batch at this step for a data or tokenization bug (e.g. an unmasked pad token, a corrupted shard).
- If it self-recovers within a few steps and doesn't recur, it may be benign -- confirm against the divergence detector's verdict.
WARN
steps 230–249
overfitting
Overfitting onset
Best eval/loss was 0.146 at step 230; it's since risen to 0.1941 (32.9%, slope p=0.0245).
best_step=230 best_value=0.146 current_value=0.1941 relative_rise=0.3293 slope_p_value=0.02446
- Use the checkpoint at step 230 (best eval/loss), not the last one.
- Add or increase regularization: weight decay, dropout, data augmentation.
- Reduce model capacity or train for fewer steps / add early stopping.
- Get more training data, or de-duplicate against the eval set if overlap is possible.