b1ade
Collection
SLMs for RAG • 8 items • Updated • 1
How to use w601sxs/b1ade-1b-grpo with PEFT:
from peft import PeftModel
from transformers import AutoModelForCausalLM
base_model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.2-1B-Instruct")
model = PeftModel.from_pretrained(base_model, "w601sxs/b1ade-1b-grpo")A reasoning-optimized Llama-3.2-1B model trained with GRPO (Group Relative Policy Optimization) on chain-of-thought data.
B1ade-1B-GRPO is a 1 billion parameter language model fine-tuned using reinforcement learning (GRPO) to improve reasoning and chain-of-thought capabilities. The model is trained directly on the base Llama-3.2-1B-Instruct without an intermediate supervised fine-tuning step, using LoRA for parameter-efficient training.
This model is designed for:
LoRA Settings:
Training Hyperparameters:
GRPO Settings:
Dataset: w601sxs/simplecot_subset_50k
Model checkpoints were saved every 500 steps. The final checkpoint at step 5000 is the published model.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
# Load base model
base_model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.2-1B-Instruct",
torch_dtype="auto",
device_map="auto"
)
# Load LoRA adapters
model = PeftModel.from_pretrained(base_model, "w601sxs/b1ade-1b-grpo")
tokenizer = AutoTokenizer.from_pretrained("w601sxs/b1ade-1b-grpo")
# Generate
prompt = "What is 25 times 4? Think step by step."
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
@misc{b1ade-1b-grpo,
author = {w601sxs},
title = {B1ade-1B-GRPO: GRPO-trained Llama 1B for Reasoning},
year = {2026},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/w601sxs/b1ade-1b-grpo}}
}
w601sxs
Base model
meta-llama/Llama-3.2-1B-Instruct