AutoResearch-tinystories-depth10
AutoResearch-tinystories-depth10 is a 122.6M parameter decoder-only Transformer trained from scratch on TinyStories (karpathy/tinystories-gpt4-clean).
This model is part of the AutoResearch project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows.
Overview
This is a 10-layer decoder-only Transformer trained on the TinyStories (karpathy/tinystories-gpt4-clean) dataset for 4.0 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 0.358686 (perplexity: 1.2823) on the held-out validation set.
References
Papers
- NanoGPT / NanoChat architecture patterns
Datasets
Related Projects
WANDB Run
Highlights
- Trained from scratch
- 122.6M parameters
- Trained on 331.4M tokens (632 steps)
- 10-layer decoder-only Transformer with sliding window attention
- RoPE positional encoding, RMSNorm, ReLUยฒ activation
- MuonAdamW optimizer (Muon for matrices, AdamW for embeddings)
- Hugging Face Transformers compatible
Model Architecture
| Property | Value |
|---|---|
| Architecture | Decoder-only Transformer |
| Parameters | 122,553,140 (122.6M) |
| Layers | 10 |
| Hidden Size | 640 |
| Attention Heads | 5 |
| KV Heads | 5 |
| Head Dimension | 128 |
| Feed Forward Size | 2560 |
| Context Length | 2048 |
| Vocabulary Size | 16,384 |
| Positional Encoding | RoPE |
| Activation | ReLUยฒ |
| Normalization | RMSNorm |
| Window Pattern | SSSL |
| Weight Tying | No |
Training
This model was trained from scratch for 4.0 hours (14416s) of wall-clock training time.
Training Configuration
| Setting | Value |
|---|---|
| Optimizer | MuonAdamW (Muon + AdamW) |
| Precision | torch.bfloat16 |
| Learning Rate | 0.04 (matrix) / 0.6 (embedding) |
| Weight Decay | 0.2 |
| Batch Size | 4 ร 2048 = 8,192 tokens/step |
| Gradient Accumulation | 64 steps |
| Total Batch Size | 524,288 tokens |
| Context Length | 2048 |
| Vocabulary | 16,384 tokens (BPE) |
| LR Scheduler | Linear warmdown (50%) |
| Activation Checkpointing | Enabled |
Hardware
- GPU: NVIDIA GeForce RTX 4060 Ti
- VRAM: 16.0 GB
- Peak VRAM Used: 3.5 GB
- MFU: 11.77%
- Framework: PyTorch 2.9.1+cu128
Dataset
- Name: TinyStories (karpathy/tinystories-gpt4-clean)
- Language: English
Preprocessing
Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.
Intended Use
This model is intended for:
- Educational purposes and research
- Text generation experiments
- Studying small language model training dynamics
Not recommended for:
- Production use or safety-critical applications
- Tasks requiring factual accuracy
Evaluation
Results
| Metric | Score |
|---|---|
| Validation BPB | 0.358686 |
| Perplexity | 1.2823 |
| Peak VRAM | 3.5 GB |
| MFU | 11.77% |
Example Generations
Example 1
Prompt
Once upon a time,
Generation
Once upon a time, 3 year old called Amy went to the beach. She ran and laughed in the sand. The sun was warm, and the water was cool. She even saw a dolphin! It was so funny and Amy smiled.
The beach was filled with all
Example 2
Prompt
A lonely dragon
Generation
A lonely dragon
Example 3
Prompt
The opposite of boy is
Generation
The opposite of boy is
Example 4
Prompt
The opposite of queen is
Generation
The opposite of queen is
Example 5
Prompt
My name is
Generation
My name is
Usage
import torch
import pickle
import json
from train import GPT, GPTConfig, Tokenizer
# Load config
with open('config.json', 'r') as f:
config_dict = json.load(f)
config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})
# Load model
model = GPT(config)
state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
model.load_state_dict(state_dict)
model.eval()
# Load tokenizer
with open('tokenizer.pkl', 'rb') as f:
tokenizer = pickle.load(f)
# Generate
prompt = 'Once upon a time, '
input_ids = tokenizer.encode(prompt)
x = torch.tensor([input_ids], dtype=torch.long)
with torch.no_grad():
for _ in range(50):
logits = model(x)
probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
next_token = torch.multinomial(probs, num_samples=1)
input_ids.append(next_token.item())
x = torch.tensor([input_ids], dtype=torch.long)
print(tokenizer.decode(input_ids))
Repository Structure
model.pt # Model weights
config.json # Model architecture config
dataset.txt # Dataset name used for training
token_bytes.pt # Token byte mappings
tokenizer.pkl # Trained BPE tokenizer
tokenizer_config.json # Tokenizer configuration
training_metrics.json # Training metrics
README.md # This file
Limitations
- Small model size limits language understanding and coherence
- Trained on a single dataset (TinyStories) โ limited domain
- Fixed time budget training โ not fully trained to convergence
- No RLHF or safety alignment
Ethical Considerations
- This is a research artifact, not a production model
- The training data consists of synthetic stories (GPT-4 generated)
- No harmful content filtering was applied
- Intended for research and educational use only
Citation
@misc{autoresearch_tinystories_depth10,
title={AutoResearch-tinystories-depth10},
author={Dustin Loring},
year={2026},
howpublished={\url{https://huggingface.co/quik-models/dry-blaze-34}}
}}
Version History
| Version | Date | Notes |
|---|---|---|
| v1.0 | 2026-07-26 | Initial release |
Acknowledgements
Built with the AutoResearch training framework.
Thanks to:
- Hugging Face
- PyTorch
- The creators of the TinyStories dataset
- The open-source AI research community
License
This model is released under the MIT License unless otherwise specified.
- Downloads last month
- 19
