--- library_name: transformers tags: - TinyLM - small - open-source - 50M license: mit datasets: - HuggingFaceFW/fineweb-edu - Salesforce/wikitext language: - en pipeline_tag: text-generation model-index: - name: TinyLM2-50M results: - task: type: multiple-choice name: HellaSwag dataset: type: hellaswag name: HellaSwag metrics: - type: accuracy_norm value: 0.2761 - task: type: multiple-choice name: CommonsenseQA dataset: type: commonsense_qa name: CommonsenseQA metrics: - type: accuracy value: 0.1957 - task: type: multiple-choice name: PIQA dataset: type: piqa name: PIQA metrics: - type: accuracy_norm value: 0.5930 - task: type: multiple-choice name: Winogrande dataset: type: winogrande name: Winogrande metrics: - type: accuracy value: 0.4972 - task: type: language-modeling name: WikiText dataset: type: wikitext name: WikiText-2 metrics: - type: perplexity value: 77.01 - task: type: linguistic-acceptability name: BLiMP dataset: type: blimp name: BLiMP metrics: - type: accuracy value: 0.7641 - task: type: multiple-choice name: ARC-Easy dataset: type: ai2_arc name: ARC-Easy config: ARC-Easy metrics: - type: accuracy_norm value: 0.4150 - task: type: multiple-choice name: ARC-Challenge dataset: type: ai2_arc name: ARC-Challenge config: ARC-Challenge metrics: - type: accuracy_norm value: 0.2449 - task: type: multiple-choice name: SciQ dataset: type: sciq name: SciQ metrics: - type: accuracy_norm value: 0.6010 - task: type: multiple-choice name: MMLU dataset: type: cais/mmlu name: MMLU metrics: - type: accuracy value: 0.2295 --- ## TinyLM2-50M **TinyLM2-50M** is a compact decoder-only Transformer language model designed for efficient instruction following and conversational AI. The model has approximately 50M parameters and has been pre-trained on 4 billion tokens using the ALiBi decoder-only architecture, providing a lightweight platform for researching and understanding language-model behavior while remaining suitable for local inference and resource-constrained environments. ## Evaluation All evaluations are zero-shot unless stated otherwise, and i used lm_eval to run them ## Model Architecture & Hyperparameters TinyLM2-50M is built on a custom ALiBi Decoder-Only Transformer architecture with pre-normalization and gated feedforward networks: | Hyperparameter | Value | Description | | :--- | :--- | :--- | | **Architecture** | ALiBi Decoder-Only Transformer | Autoregressive Decoder-Only Transformer | | **Total Parameters** | **~50.96M** (53,430,272) | Compact and ultra-fast for edge & local CPU/GPU inference | | **`vocab_size`** | **50,271** | Includes special chat tags (`<|SYSTEM|>`, `<|USER|>`, `<|ASSISTANT|>`) | | **`hidden_size` (`d_model`)** | **512** | Model hidden dimension | | **`intermediate_size` (`ff_hidden_d`)**| **819** | SwiGLU Gated Feedforward hidden dimension | | **`num_hidden_layers`** | **12** | Number of Transformer block layers | | **`num_attention_heads`** | **8** | Attention heads (Head dim = 64) | | **`max_position_embeddings`** | **2,048** | Maximum context sequence length | | **Normalization** | **RMSNorm** (`eps=1e-8`) | Scale normalization for accelerated throughput | | **Activation Function** | **SwiGLU (SiLU)** | Gated Feedforward activation | | **Positional Encoding** | **ALiBi** | Attention with Linear Biases | | **Tie Word Embeddings** | `True` | Tied input embedding and LM head projection | --- ## Tokenizer & Chat Template The model uses a custom Byte-Level BPE Tokenizer equipped with special tokens and a pre-configured Jinja2 `chat_template` for multi-turn conversations. | Property | Value | | :--- | :--- | | **Tokenizer Type** | `GPT2Tokenizer` (Byte-Level BPE) | | **Vocabulary Size** | 50,271 | | **Special Tokens** | `<|START|>` `<|END|>` `<|UNK|>` | | **Chat Control Tokens** | `<|SYSTEM|>` `<|USER|>` `<|ASSISTANT|>` | | **Extra Special Tokens** | `<|THINK|>` `<|/THINK|>` `<|AVAILABLE_TOOLS|>` `<|/AVAILABLE_TOOLS|>` `<|TOOL_CALLS|>` `<|/TOOL_CALLS|>` `<|TOOL_RESULTS|>` `<|/TOOL_RESULTS|>` | | **Chat Template** | Native Jinja2 support via `tokenizer.apply_chat_template()` | --- ## Training Configuration | Parameter | Value | | :--- | :--- | | **Pipeline Process** | Pre-Training (PT) | | **Dataset** | `HuggingFaceFW/fineweb-edu` (`sample-100BT`), `Salesforce/wikitext` (`wikitext-2-v1`) | | **Total Tokens** | 4,000,000,000 (4B) | | **Epochs** | 1 | | **Learning Rate** | `1e-4` | | **Learning Rate Schedule** | Cosine (`warmup_ratio=0.01`) | | **Micro-Batch Size** | 2 per device | | **Gradient Accumulation** | 16 steps | | **Effective Batch Size** | 32 × 2,048 tokens | | **Optimizer** | AdamW (`weight_decay=0.1`) | | **Max Sequence Length** | 2,048 tokens | | **Precision** | float16 | | **Hardware** | NVIDIA Tesla T4 x 2 GPU | --- ## Inference ```python # pip install torch transformers import torch from transformers import pipeline pipe = pipeline( "text-generation", model="Se00n00/TinyLM2-50M", trust_remote_code = True ) messages = [ {"role": "system", "content": "You are a helpful AI assistant."}, {"role": "user", "content": "Explain artificial intelligence in simple terms."} ] prompt = pipe.tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) result = pipe( prompt, max_new_tokens=120, do_sample=True, temperature=0.5, top_k=40, top_p=0.9 ) print(result[0]['generated_text']) ``` ────── ## Sample Outputs Raw next-token continuation (no chat template, temperature 0.8, top-p 0.9): **Prompt**: `Photosynthesis is the process by which` > Plants take up oxygen and use it to make energy. Plant leaves, flowers, and > even fruit can also convert carbon dioxide (CO2) into sugars. As you grow, > plants are able to store the carbon dioxide from their leaves and convert it > to sugars and carbohydrates. This is known as photosynthesis, and it is a > process that converts the carbon dioxide into a form of energy. …