ANLP Assignment 2, Task 2: hand-written optimizers
The dense 26.45M-parameter decoder of Task 1 (variant v1), pretrained from scratch on browndw/human-ai-parallel-corpus for 1x the training data (44.43M tokens) with AdamW, MARS, Lion and Muon (one per category of 'Fantastic Pretraining Optimizers'), plus Sophia-H from its Hessian-based category as a bonus.
| optimizer | val loss | val ppl | human | LLM | BLEU | BLEU (sm.) | ROUGE-L | state MB | peak MB | tok/s | min | reaches AdamW-final at | speedup vs AdamW |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AdamW | 3.5931 | 36.34 | 4.290 | 3.371 | 0.59 | 0.62 | 13.36 | 201.8 | 5,187 | 98,515 | 10.0 | 1.00x data | 1.00x |
| MARS | 3.4868 | 32.68 | 4.206 | 3.258 | 0.79 | 0.82 | 13.71 | 302.7 | 5,287 | 96,010 | 10.3 | 0.71x data | 1.41x |
| Lion | 3.6045 | 36.76 | 4.294 | 3.385 | 0.78 | 0.80 | 13.08 | 100.9 | 5,086 | 99,751 | 9.6 | never | 0.89x |
| Muon | 3.3125 | 27.45 | 4.051 | 3.077 | 0.77 | 0.80 | 13.16 | 147.8 | 5,133 | 89,851 | 10.5 | 0.68x data | 1.47x |
| Sophia | 3.6592 | 38.83 | 4.341 | 3.442 | 0.70 | 0.73 | 13.78 | 201.8 | 5,187 | 92,881 | 10.5 | never | 0.75x |
<optimizer>/final.pt holds {'model': state_dict, 'config': ...}; the tokenizer is
Task 1's (tokenizer.json), and the code is in the submission repository.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support