Qwen3-8B-CalibSFT

Paper Code Dataset

Introduction

Qwen3-8B-CalibSFT is Qwen3-8B fine-tuned with CalibSFT, the method proposed in our paper On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models. For each training question, CalibSFT labels every sampled response with a confidence that mixes the question's success rate with the response's correctness, balances the data across confidence levels from 0 to 1, and supervises reasoning and answers only on correct responses. Compared with Qwen3-8B, it substantially lowers the Brier score and ECE on both in-distribution and out-of-distribution benchmarks, and it serves as the initialization of Qwen3-8B-CalibSFT-RLCR. For training details, please refer to our GitHub repository.

Usage

Use the system prompt and user format from training:

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "SUSTech/Qwen3-8B-CalibSFT"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="auto", device_map="auto")

SYSTEM_PROMPT = (
    "A conversation between User and Assistant. The user asks a question, and the Assistant solves it. "
    "The assistant first thinks about the reasoning process in the mind and analyzes its confidence about "
    "the solution and then provides the user with the final answer as well as its confidence level. "
    "The confidence level indicates how certain the Assistant is about its answer, expressed as a decimal "
    "between 0 and 1 with exactly two decimal places, enclosed within <confidence> </confidence> tags. "
    "The response must strictly follow this format: <think> reasoning process here </think> "
    "<answer> final short answer only </answer> <confidence> 0.xx </confidence>. "
    "The <answer> tag must contain only the final answer string needed for exact-match evaluation, "
    "not a full sentence, explanation, or reasoning."
)
question = "What is the smallest positive integer n such that n^2 + n is divisible by 12?"
messages = [
    {"role": "system", "content": SYSTEM_PROMPT},
    {"role": "user", "content": f"\n\nPROBLEM: {question}\n\n"},
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=16384, do_sample=True, temperature=0.6, top_p=0.95, top_k=20)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
# <think> ... </think> <answer> ... </answer> <confidence> 0.xx </confidence>

Evaluation

Results are averaged over 8 in-distribution math benchmarks (DeepScaleR-Eval, MATH-500, MinervaMath, OlympiadBench, GSM8K, AIME 2024–2026) and 8 out-of-distribution benchmarks (HotpotQA, TriviaQA, DROP, MuSiQue, LiveBenchReasoning, NQOpen, PopQA, WebQuestions), sampling at temperature 0.6 with 4 responses per question (32 on AIME). DeepScaleR-Eval consists of 2,000 DeepScaleR questions held out from training.

Model In-distribution Out-of-distribution
Pass@1 ↑AUROC ↑Brier ↓ECE ↓ Pass@1 ↑AUROC ↑Brier ↓ECE ↓
Qwen3-8B59.3271.9033.2034.3645.3665.7145.1546.94
Qwen3-8B-CalibSFT58.5179.9317.3014.2346.3761.2029.1525.49

Citation

@article{wang2026pitfalls,
  title={On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models},
  author={Wang, Shuoyuan and Luo, Beier and Zeng, Hao and Yu, Chengyao and Zhang, Songxin and Xie, Zejian and Jing, Bingyi and Wei, Hongxin},
  journal={arXiv preprint arXiv:2609.32470},
  year={2026}
}
Downloads last month
78
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SUSTech/Qwen3-8B-CalibSFT

Finetuned
Qwen/Qwen3-8B
Finetuned
(2163)
this model
Finetunes
1 model

Dataset used to train SUSTech/Qwen3-8B-CalibSFT

Collection including SUSTech/Qwen3-8B-CalibSFT

Paper for SUSTech/Qwen3-8B-CalibSFT