Instructions to use SUSTech/Qwen3-8B-CalibSFT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SUSTech/Qwen3-8B-CalibSFT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="SUSTech/Qwen3-8B-CalibSFT") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("SUSTech/Qwen3-8B-CalibSFT") model = AutoModelForCausalLM.from_pretrained("SUSTech/Qwen3-8B-CalibSFT", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SUSTech/Qwen3-8B-CalibSFT with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SUSTech/Qwen3-8B-CalibSFT" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SUSTech/Qwen3-8B-CalibSFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/SUSTech/Qwen3-8B-CalibSFT
- SGLang
How to use SUSTech/Qwen3-8B-CalibSFT with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SUSTech/Qwen3-8B-CalibSFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SUSTech/Qwen3-8B-CalibSFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SUSTech/Qwen3-8B-CalibSFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SUSTech/Qwen3-8B-CalibSFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use SUSTech/Qwen3-8B-CalibSFT with Docker Model Runner:
docker model run hf.co/SUSTech/Qwen3-8B-CalibSFT
Qwen3-8B-CalibSFT
Introduction
Qwen3-8B-CalibSFT is Qwen3-8B fine-tuned with CalibSFT, the method proposed in our paper On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models. For each training question, CalibSFT labels every sampled response with a confidence that mixes the question's success rate with the response's correctness, balances the data across confidence levels from 0 to 1, and supervises reasoning and answers only on correct responses. Compared with Qwen3-8B, it substantially lowers the Brier score and ECE on both in-distribution and out-of-distribution benchmarks, and it serves as the initialization of Qwen3-8B-CalibSFT-RLCR. For training details, please refer to our GitHub repository.
Usage
Use the system prompt and user format from training:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "SUSTech/Qwen3-8B-CalibSFT"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="auto", device_map="auto")
SYSTEM_PROMPT = (
"A conversation between User and Assistant. The user asks a question, and the Assistant solves it. "
"The assistant first thinks about the reasoning process in the mind and analyzes its confidence about "
"the solution and then provides the user with the final answer as well as its confidence level. "
"The confidence level indicates how certain the Assistant is about its answer, expressed as a decimal "
"between 0 and 1 with exactly two decimal places, enclosed within <confidence> </confidence> tags. "
"The response must strictly follow this format: <think> reasoning process here </think> "
"<answer> final short answer only </answer> <confidence> 0.xx </confidence>. "
"The <answer> tag must contain only the final answer string needed for exact-match evaluation, "
"not a full sentence, explanation, or reasoning."
)
question = "What is the smallest positive integer n such that n^2 + n is divisible by 12?"
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"\n\nPROBLEM: {question}\n\n"},
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=16384, do_sample=True, temperature=0.6, top_p=0.95, top_k=20)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
# <think> ... </think> <answer> ... </answer> <confidence> 0.xx </confidence>
Evaluation
Results are averaged over 8 in-distribution math benchmarks (DeepScaleR-Eval, MATH-500, MinervaMath, OlympiadBench, GSM8K, AIME 2024–2026) and 8 out-of-distribution benchmarks (HotpotQA, TriviaQA, DROP, MuSiQue, LiveBenchReasoning, NQOpen, PopQA, WebQuestions), sampling at temperature 0.6 with 4 responses per question (32 on AIME). DeepScaleR-Eval consists of 2,000 DeepScaleR questions held out from training.
| Model | In-distribution | Out-of-distribution | ||||||
|---|---|---|---|---|---|---|---|---|
| Pass@1 ↑ | AUROC ↑ | Brier ↓ | ECE ↓ | Pass@1 ↑ | AUROC ↑ | Brier ↓ | ECE ↓ | |
| Qwen3-8B | 59.32 | 71.90 | 33.20 | 34.36 | 45.36 | 65.71 | 45.15 | 46.94 |
| Qwen3-8B-CalibSFT | 58.51 | 79.93 | 17.30 | 14.23 | 46.37 | 61.20 | 29.15 | 25.49 |
Citation
@article{wang2026pitfalls,
title={On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models},
author={Wang, Shuoyuan and Luo, Beier and Zeng, Hao and Yu, Chengyao and Zhang, Songxin and Xie, Zejian and Jing, Bingyi and Wei, Hongxin},
journal={arXiv preprint arXiv:2609.32470},
year={2026}
}
- Downloads last month
- 78