Qielbas Tiny 9M M2

A 9,557,716-parameter English causal language model.

Architecture

Same architecture and tokenizer as M1: four GatedDeltaNet2 blocks and one causal attention block, width 320, SwiGLU hidden size 1048, RMSNorm, RoPE, tied embeddings, context 256 and a 4096-token BPE vocabulary.

Changes from M1

Both releases start from the same Muon pretrained checkpoint average. M2 adds a 300M-token continuation with 15% WikiText-103 raw training text and 85% of the original data mixture. WikiText-2 and WikiText-103 validation/test articles were excluded by title and exact text before training.

The continuation uses fresh Muon + AdamW, peak LR 0.0001, and seed 7. ARC adaptation uses fresh AdamW, peak LR 0.00005, five epochs and seed 123. Its loss is 0.15 choice CE + 0.85 text replay CE; M1 used 1.0 + 0.25. Replay keeps the 15% WikiText share. Epoch 5 was selected by ARC-Easy held-out validation accuracy; M1 used mean Easy/Challenge accuracy with a language-loss guard. Benchmark test data was not used for training or epoch selection.

Data

The original 6.7B tokens comprise FineWeb-Edu (2.7B), Cosmopedia v2 (0.3B), FinePhrase FAQ and Tutorial (0.5B each), DCLM-Edu (1.0B), FinePDFs-Edu English (1.2B), and FineWiki English (0.5B). The additional 300M tokens contain 45M WikiText exposures and 255M exposures from that mixture, bringing pretraining to 7.0B tokens. The audited WikiText pool contains 176M unique tokens.

ARC adaptation uses 3345 Easy/Challenge training questions and 867 held-out validation questions, with general-text replay.

Benchmarks

Metric M1 M2
Tiny-ML Efficiency 80.1426 81.4170
Raw Overall 71.8696 73.0124
ARC-Easy accuracy 48.2744% 47.4327%
BLiMP accuracy 75.3224% 75.6836%
WikiText-2 byte perplexity 2.90791 2.33674

M2 gains 1.2744 Efficiency points (+1.59%).

Same frozen Glint-1.3 protocol as M1: 256-token contexts, no BOS, unnormalized answer likelihood, 2376 ARC-Easy questions and 67000 BLiMP pairs. Efficiency uses the retained Tiny-ML normalization. Results are self-reported; the separate official-harness M1 evaluation is not mixed into this table. Reports and predictions are included.

Usage

Native PyTorch loader; requires a BF16-capable CUDA GPU and the pinned dependencies.

pip install -r requirements.txt
import torch
from load_model import load_model

model, tokenizer = load_model(".")
inputs = torch.tensor([tokenizer.encode("The Earth revolves around").ids], device="cuda")
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
    next_token = model(inputs)[0, -1].argmax().item()
print(tokenizer.decode([next_token]))
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train dawidkielbasa/qielbas-tiny-9m-m2

Space using dawidkielbasa/qielbas-tiny-9m-m2 1