ANLP Assignment 2, Task 2: hand-written optimizers

The dense 26.45M-parameter decoder of Task 1 (variant v1), pretrained from scratch on browndw/human-ai-parallel-corpus for 1x the training data (44.43M tokens) with AdamW, MARS, Lion and Muon (one per category of 'Fantastic Pretraining Optimizers'), plus Sophia-H from its Hessian-based category as a bonus.

optimizer val loss val ppl human LLM BLEU BLEU (sm.) ROUGE-L state MB peak MB tok/s min reaches AdamW-final at speedup vs AdamW
AdamW 3.5931 36.34 4.290 3.371 0.59 0.62 13.36 201.8 5,187 98,515 10.0 1.00x data 1.00x
MARS 3.4868 32.68 4.206 3.258 0.79 0.82 13.71 302.7 5,287 96,010 10.3 0.71x data 1.41x
Lion 3.6045 36.76 4.294 3.385 0.78 0.80 13.08 100.9 5,086 99,751 9.6 never 0.89x
Muon 3.3125 27.45 4.051 3.077 0.77 0.80 13.16 147.8 5,133 89,851 10.5 0.68x data 1.47x
Sophia 3.6592 38.83 4.341 3.442 0.70 0.73 13.78 201.8 5,187 92,881 10.5 never 0.75x

<optimizer>/final.pt holds {'model': state_dict, 'config': ...}; the tokenizer is Task 1's (tokenizer.json), and the code is in the submission repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support