Instructions to use robinzixuan/sink_token_llama3.2_1b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use robinzixuan/sink_token_llama3.2_1b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="robinzixuan/sink_token_llama3.2_1b")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("robinzixuan/sink_token_llama3.2_1b") model = AutoModelForCausalLM.from_pretrained("robinzixuan/sink_token_llama3.2_1b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use robinzixuan/sink_token_llama3.2_1b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "robinzixuan/sink_token_llama3.2_1b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "robinzixuan/sink_token_llama3.2_1b", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/robinzixuan/sink_token_llama3.2_1b
- SGLang
How to use robinzixuan/sink_token_llama3.2_1b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "robinzixuan/sink_token_llama3.2_1b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "robinzixuan/sink_token_llama3.2_1b", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "robinzixuan/sink_token_llama3.2_1b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "robinzixuan/sink_token_llama3.2_1b", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use robinzixuan/sink_token_llama3.2_1b with Docker Model Runner:
docker model run hf.co/robinzixuan/sink_token_llama3.2_1b
sink_token_llama3.2_1b
Sink baseline S2 (prefix sink) checkpoint built on Llama-3.2-1B: a learnable per-head sink K/V slot every query can attend to (carries a learnable value). Research checkpoint from the paper Attention Sinks and Outliers in Attention Residuals (OASIS), part of the OASIS collection.
Model summary
| Base model | meta-llama/Llama-3.2-1B |
| Variant | Sink baseline S2 (prefix sink) |
| Token-level attention | Softmax over learnable per-head sink K/V slot |
| Depth routing | AttentionResidual |
| Null → depth coupling | ✗ |
| Training | Continued training at block size 6,144 on BookCorpus + Wikipedia + LongAlpaca |
| Parameters | 1.30B (F32) |
| Config class | LlamaForCausalLM |
| License | llama3.2 (inherited from the base model) |
About OASIS
AttnResidual architectures add a depth-wise normalization channel that improves inter-layer routing flexibility, but it also amplifies attention sinks, activation outliers, and the resulting loss of inference stability and quantization robustness. OASIS adds a Softmax1-based null space and couples token-level null evidence to depth routing through an inter-layer null signal, reducing sink-dominated routing. This repo is one arm of that comparison (OASIS, its ablations, and sink / vanilla baselines); see the table at the bottom for all variants.
Usage
⚠️ Weights only. The checkpoint contains extra AttentionResidual (and, where applicable, Softmax1 / sink) parameters that are not part of the stock
LlamaForCausalLM. Loading it with plainAutoModelForCausalLMsilently drops those parameters and runs standard attention, so outputs will not reflect this variant. Build the model with the OASIS modeling code, then load the weights.
import glob, os, torch
from huggingface_hub import snapshot_download
from safetensors.torch import load_file
from transformers import AutoTokenizer
model_id = "robinzixuan/sink_token_llama3.2_1b"
path = snapshot_download(model_id)
tokenizer = AutoTokenizer.from_pretrained(path)
state_dict = {}
for f in sorted(glob.glob(os.path.join(path, "*.safetensors"))):
state_dict.update(load_file(f))
# TODO: replace with the model constructor from the OASIS codebase (variant="sink_token")
model = build_oasis_model(path, variant="sink_token")
missing, unexpected = model.load_state_dict(state_dict, strict=False)
assert not missing and not unexpected, (missing, unexpected)
model = model.to("cuda" if torch.cuda.is_available() else "cpu").eval()
inputs = tokenizer("Hello, ", return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
To see which parameters stock transformers would drop:
from transformers import AutoModelForCausalLM
_, info = AutoModelForCausalLM.from_pretrained("robinzixuan/sink_token_llama3.2_1b", output_loading_info=True)
print(info["unexpected_keys"][:10]) # AttentionResidual / sink parameters ignored by LlamaForCausalLM
Evaluation
See the paper for the attention-sink, outlier (‖·‖∞, kurtosis) and post-quantization (W8A8 / W4A4) comparisons across variants.
Base model
Architecture
config.jsondeclaresLlamaForCausalLMwith extra AttentionResidual + sink modules; requires the OASIS modeling code (see Usage).
OASIS collection
| Model | Base | Variant | Context |
|---|---|---|---|
| Phi4_attn_residual | Phi-4-mini-instruct | Vanilla AttnResidual baseline | standard |
| Phi4_attn_OASIS | Phi-4-mini-instruct | OASIS | standard |
| Phi4_attn_AoS | Phi-4-mini-instruct | AoS (attention-only Softmax1) | standard |
| qwen3-0.6b-oasis | Qwen3-0.6B | OASIS | standard |
| qwen3-0.6b-vanilla | Qwen3-0.6B | Vanilla AttnResidual baseline | standard |
| llama-3.2-1b-vanilla | Llama-3.2-1B | Vanilla AttnResidual baseline | standard |
| llama-3.2-1b-oasis | Llama-3.2-1B | OASIS | standard |
| oasis_llama3.2_1b_6k | Llama-3.2-1B | OASIS | 6k |
| vanilla_llama3.2_1b_6k | Llama-3.2-1B | Vanilla AttnResidual baseline | 6k |
| oasis_qwen3_0.6b_6k | Qwen3-0.6B | OASIS | 6k |
| vanilla_qwen3_0.6b_6k | Qwen3-0.6B | Vanilla AttnResidual baseline | 6k |
| sink_logit_llama3.2_1b | Llama-3.2-1B | Sink baseline S1 (sink logit) | 6k |
| sink_token_llama3.2_1b (this model) | Llama-3.2-1B | Sink baseline S2 (prefix sink) | 6k |
| s0_softmax1_noroute_llama3.2_1b | Llama-3.2-1B | Ablation S0 (Softmax1, no routing) | 6k |
Citation
@article{luo2026oasis,
title = {Attention Sinks and Outliers in Attention Residuals},
author = {Luo, Haozheng and Dai, Haoran and Zhang, Shaoyang and Chen, Xi and
Jiang, Eric Hanchen and Li, Yijiang and Huang, Jingyuan and Qiu, Chenghao and
Xu, Chenwei and Pan, Zhenyu and Zhang, Haotian and Wang, Binghui and Chen, Yan},
journal = {arXiv preprint arXiv:2605.17887},
year = {2026}
}
Built with Llama. Llama 3.2 is licensed under the Llama 3.2 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.
- Downloads last month
- 298
Model tree for robinzixuan/sink_token_llama3.2_1b
Base model
meta-llama/Llama-3.2-1B