Text Generation
Transformers
Safetensors
English
qwen3
agent
code
QASM
quantum
conversational
text-generation-inference
Instructions to use Benyucong/rl_quantum_4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Benyucong/rl_quantum_4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Benyucong/rl_quantum_4b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Benyucong/rl_quantum_4b") model = AutoModelForCausalLM.from_pretrained("Benyucong/rl_quantum_4b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Benyucong/rl_quantum_4b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Benyucong/rl_quantum_4b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Benyucong/rl_quantum_4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Benyucong/rl_quantum_4b
- SGLang
How to use Benyucong/rl_quantum_4b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Benyucong/rl_quantum_4b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Benyucong/rl_quantum_4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Benyucong/rl_quantum_4b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Benyucong/rl_quantum_4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Benyucong/rl_quantum_4b with Docker Model Runner:
docker model run hf.co/Benyucong/rl_quantum_4b
File size: 6,174 Bytes
07b4322 4758135 07b4322 4758135 07b4322 3753cce 07b4322 4b6ac2c 07b4322 4758135 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 3ee8f99 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 1498ea6 4b6ac2c 0d1de20 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 2f2b703 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 36040a5 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 07b4322 4b6ac2c 7f065c7 4b6ac2c af03636 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 | ---
base_model:
- Qwen/Qwen3-4B-Instruct-2507
datasets:
- Benyucong/graph-data-quantum-rl
language:
- en
library_name: transformers
license: apache-2.0
metrics:
- code_eval
pipeline_tag: text-generation
tags:
- agent
- code
- QASM
- quantum
---
# QUASAR: Quantum Assembly Code Generation with Tool-Augmented RL
[](https://huggingface.co/papers/2510.00967) [](https://github.com/benyucong/QUASAR) [](https://huggingface.co/datasets/Benyucong/graph-data-quantum-rl)
## Model Summary
**QUASAR** is a 4B-parameter model fine-tuned from **Qwen3-4B-Instruct-2507** using a two-stage process: supervised fine-tuning (SFT) followed by agentic reinforcement learning (RL) with tool-augmented feedback.
The model is designed to **generate OpenQASM 3.0 quantum circuits** for optimization problems such as **QAOA** and **VQE**, achieving **high syntactic validity and semantic fidelity**.
- **Framework:** Agentic RL with external quantum simulator verification
- **Reward:** Hierarchical 4-level reward (syntax, distribution alignment, expectation value, optimization progress)
- **Primary Domain:** Quantum circuit generation and quantum optimization algorithm design
---
## Model Details
- **Model type:** LLM fine-tuned with reinforcement learning
- **Languages:** English
- **License:** Apache-2.0
- **Base model:** Qwen/Qwen3-4B-Instruct-2507
---
## Uses
### Direct Use
- Generate OpenQASM 3.0 code from natural language descriptions
- Design ansatz circuits for quantum optimization tasks (QAOA, VQE)
### Downstream Use
- Integration into quantum compilers
- Research on LLM-guided quantum algorithm design
---
## Bias, Risks, and Limitations
- May produce a valid QASM that is semantically weak if prompts are ambiguous
- Tailored primarily to **graph-based quantum optimization problems**
- Evaluated mainly in simulation; hardware generalization remains untested
**Recommendation:** Always verify generated circuits with independent quantum simulators or compilers before deployment.
---
## How to Get Started
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Benyucong/rl_quantum_4b")
tokenizer = AutoTokenizer.from_pretrained("Benyucong/rl_quantum_4b")
prompt = """Design a QASM 3.0 quantum circuit with 3 qubits and 3 layers to solve the vertex_cover \
given the graph: {"directed": false, "multigraph": false, "graph": {}, "nodes": [{"id": 0}, {"id": 1}, {"id": 2}], \
"edges": [{"source": 0, "target": 1}, {"source": 0, "target": 2}, {"source": 1, "target": 2}]}. \
Provide valid QASM 3.0 code with optimal parameters."""
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=1024)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
---
## Training Details
### Training Data
- Dataset: [Benyucong/graph-data-quantum-rl](https://huggingface.co/datasets/Benyucong/graph-data-quantum-rl)
- Contains QASM 3.0 circuits, Hamiltonians, eigenvalues, and parameterized circuits for **12 quantum optimization problems**
### Training Setup
- **Stage 1:** Supervised fine-tuning (SFT), the model is available [here](https://huggingface.co/Benyucong/sft_quantum_circuit_gen_4B).
- **Stage 2:** Reinforcement learning with GRPO and hierarchical reward
### Hyperparameters
- **Batch size:** 128
- **Rollouts:** 16 per prompt (temperature = 0.7, top-p = 0.8)
- **Precision:** bf16 mixed precision
- **GPUs:** 16 × H100-64GB (FSDP enabled)
- **Training time:** ~48 hours
---
## Evaluation
### Metrics (Please check our paper for details)
- **SCR:** Syntactic Correctness Ratio
- **SREV:** Successful Rate of Expectation Value
- **RE:** Relative Entropy (distributional alignment)
- **HQCR:** High-Quality Circuit Ratio
### Results (QUASAR vs Baselines)
| Method | Pass@1 SCR ↑ | Pass@1 SREV ↑ | Pass@1 RE ↓ | Pass@1 HQCR ↑ | Pass@10 SCR ↑ | Pass@10 SREV ↑ | Pass@10 RE ↓ | Pass@10 HQCR ↑ |
|---------------------|--------------|---------------|-------------|---------------|---------------|----------------|--------------|----------------|
| DeepSeek-V3 | 94.83% | 12.24% | 19.20 | 10.00% | 98.97% | 26.38% | 16.39 | 16.38% |
| GPT-5 | 87.07% | 10.00% | 19.94 | 6.90% | 90.52% | 27.07% | 11.57 | 16.55% |
| GPT-4o | 87.93% | 9.83% | 19.42 | 6.38% | 88.79% | 18.62% | 14.08 | 12.07% |
| **Qwen3-4B SFT** | 97.41% | 18.97% | 12.74 | 15.17% | 99.65% | 31.55% | 10.81 | 23.62% |
| Cold Start GRPO | 84.48% | 19.84% | 14.32 | 12.41% | 95.17% | 27.59% | 11.38 | 18.96% |
| **QUASAR (ours)** | **99.31%** | **22.41%** | **11.61** | **17.24%** | **100%** | **33.10%** | **8.48** | **27.24%** |
---
## Environmental Impact
- **Hardware Type:** NVIDIA H100 (16×, 64GB)
- **Training Hours:** ~48
---
## Technical Specifications
- **Architecture:** Qwen3-4B-Instruct-2507
- **Fine-tuning:** SFT + RL (GRPO)
- **Reward Design:** Syntax validity, distributional alignment (JS distance), expectation-value matching, optimization-progress efficiency
- **Frameworks:** PyTorch, vLLM, Qiskit, OpenQASM
---
## Citation
```bibtex
@misc{yu2025quasarquantumassemblycode,
title={QUASAR: Quantum Assembly Code Generation Using Tool-Augmented LLMs via Agentic RL},
author={Cong Yu and Valter Uotila and Shilong Deng and Qingyuan Wu and Tuo Shi and Songlin Jiang and Lei You and Bo Zhao},
year={2025},
eprint={2510.00967},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2510.00967},
}
``` |