sink_token_llama3.2_1b

Sink baseline S2 (prefix sink) checkpoint built on Llama-3.2-1B: a learnable per-head sink K/V slot every query can attend to (carries a learnable value). Research checkpoint from the paper Attention Sinks and Outliers in Attention Residuals (OASIS), part of the OASIS collection.

Model summary

Base model meta-llama/Llama-3.2-1B
Variant Sink baseline S2 (prefix sink)
Token-level attention Softmax over learnable per-head sink K/V slot
Depth routing AttentionResidual
Null → depth coupling ✗
Training Continued training at block size 6,144 on BookCorpus + Wikipedia + LongAlpaca
Parameters 1.30B (F32)
Config class LlamaForCausalLM
License llama3.2 (inherited from the base model)

About OASIS

AttnResidual architectures add a depth-wise normalization channel that improves inter-layer routing flexibility, but it also amplifies attention sinks, activation outliers, and the resulting loss of inference stability and quantization robustness. OASIS adds a Softmax1-based null space and couples token-level null evidence to depth routing through an inter-layer null signal, reducing sink-dominated routing. This repo is one arm of that comparison (OASIS, its ablations, and sink / vanilla baselines); see the table at the bottom for all variants.

Usage

⚠️ Weights only. The checkpoint contains extra AttentionResidual (and, where applicable, Softmax1 / sink) parameters that are not part of the stock LlamaForCausalLM. Loading it with plain AutoModelForCausalLM silently drops those parameters and runs standard attention, so outputs will not reflect this variant. Build the model with the OASIS modeling code, then load the weights.

import glob, os, torch
from huggingface_hub import snapshot_download
from safetensors.torch import load_file
from transformers import AutoTokenizer

model_id = "robinzixuan/sink_token_llama3.2_1b"
path = snapshot_download(model_id)
tokenizer = AutoTokenizer.from_pretrained(path)

state_dict = {}
for f in sorted(glob.glob(os.path.join(path, "*.safetensors"))):
    state_dict.update(load_file(f))

# TODO: replace with the model constructor from the OASIS codebase (variant="sink_token")
model = build_oasis_model(path, variant="sink_token")
missing, unexpected = model.load_state_dict(state_dict, strict=False)
assert not missing and not unexpected, (missing, unexpected)
model = model.to("cuda" if torch.cuda.is_available() else "cpu").eval()

inputs = tokenizer("Hello, ", return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

To see which parameters stock transformers would drop:

from transformers import AutoModelForCausalLM

_, info = AutoModelForCausalLM.from_pretrained("robinzixuan/sink_token_llama3.2_1b", output_loading_info=True)
print(info["unexpected_keys"][:10])  # AttentionResidual / sink parameters ignored by LlamaForCausalLM

Evaluation

See the paper for the attention-sink, outlier (‖·‖∞, kurtosis) and post-quantization (W8A8 / W4A4) comparisons across variants.

Base model

Architecture

  • config.json declares LlamaForCausalLM with extra AttentionResidual + sink modules; requires the OASIS modeling code (see Usage).

OASIS collection

Model Base Variant Context
Phi4_attn_residual Phi-4-mini-instruct Vanilla AttnResidual baseline standard
Phi4_attn_OASIS Phi-4-mini-instruct OASIS standard
Phi4_attn_AoS Phi-4-mini-instruct AoS (attention-only Softmax1) standard
qwen3-0.6b-oasis Qwen3-0.6B OASIS standard
qwen3-0.6b-vanilla Qwen3-0.6B Vanilla AttnResidual baseline standard
llama-3.2-1b-vanilla Llama-3.2-1B Vanilla AttnResidual baseline standard
llama-3.2-1b-oasis Llama-3.2-1B OASIS standard
oasis_llama3.2_1b_6k Llama-3.2-1B OASIS 6k
vanilla_llama3.2_1b_6k Llama-3.2-1B Vanilla AttnResidual baseline 6k
oasis_qwen3_0.6b_6k Qwen3-0.6B OASIS 6k
vanilla_qwen3_0.6b_6k Qwen3-0.6B Vanilla AttnResidual baseline 6k
sink_logit_llama3.2_1b Llama-3.2-1B Sink baseline S1 (sink logit) 6k
sink_token_llama3.2_1b (this model) Llama-3.2-1B Sink baseline S2 (prefix sink) 6k
s0_softmax1_noroute_llama3.2_1b Llama-3.2-1B Ablation S0 (Softmax1, no routing) 6k

Citation

@article{luo2026oasis,
  title   = {Attention Sinks and Outliers in Attention Residuals},
  author  = {Luo, Haozheng and Dai, Haoran and Zhang, Shaoyang and Chen, Xi and
             Jiang, Eric Hanchen and Li, Yijiang and Huang, Jingyuan and Qiu, Chenghao and
             Xu, Chenwei and Pan, Zhenyu and Zhang, Haotian and Wang, Binghui and Chen, Yan},
  journal = {arXiv preprint arXiv:2605.17887},
  year    = {2026}
}

Built with Llama. Llama 3.2 is licensed under the Llama 3.2 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.

Downloads last month
298
Safetensors
Model size
1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for robinzixuan/sink_token_llama3.2_1b

Finetuned
(941)
this model

Collection including robinzixuan/sink_token_llama3.2_1b

Paper for robinzixuan/sink_token_llama3.2_1b