---
pipeline_tag: text-generation
base_model:
- Qwen/Qwen3.8-27B
license: apache-2.0
library_name: Model Optimizer
tags:
- nvidia
- ModelOpt
- Qwen3.8
- quantized
- FP4
- fp4
- FP8
- fp8
---
# Model Overview
## Description:
The NVIDIA Qwen3.8-27B NVFP4 model is a quantized version of Alibaba's Qwen3.8-27B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information on the model, please check [here](https://huggingface.co/Qwen/Qwen3.8-27B). The model is quantized with [Model Optimizer](https://github.com/NVIDIA/Model-Optimizer).
This model is ready for commercial or non-commercial use.
## Third-Party Community Consideration
This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party's requirements for this application and use case; see link to Non-NVIDIA [(Qwen3.8-27B) Model Card](https://huggingface.co/Qwen/Qwen3.8-27B) from Qwen.
### License/Terms of Use:
**GOVERNING DOWNLOAD TERMS:** Use of the model is governed by the [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
### Deployment Geography:
Global
### Use Case:
Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots, RAG systems, and other AI-powered applications.
### Release Date:
Hugging Face 09/08/2026 via https://huggingface.co/nvidia/Qwen3.8-27B-NVFP4
## References
NVIDIA Model Optimizer: https://github.com/NVIDIA/Model-Optimizer
## Model Architecture:
**Architecture Type:** Transformer
**Network Architecture:** Qwen3.8-27B (`Qwen3_5ForConditionalGeneration`)
**Number of Model Parameters:** 27B
## Input:
**Input Type(s):** Text, Image, Video
**Input Format(s):** String, Red, Green, Blue (RGB), Video (MP4/WebM)
**Input Parameters:** One-Dimensional (1D), Two-Dimensional (2D), Three-Dimensional (3D)
**Other Properties Related to Input:** Context length up to 262K
## Output:
**Output Type(s):** Text
**Output Format:** String
**Output Parameters:** One-Dimensional (1D): Sequences
**Other Properties Related to Output:** None
Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.
## Software Integration:
**Supported Runtime Engine(s):**
* **vLLM**
* **SGLang**
**Supported Hardware Microarchitecture Compatibility:**
* NVIDIA Blackwell
**Preferred Operating System(s):**
* Linux
The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.
## Model Version(s):
This checkpoint uses mixed NVFP4/FP8 quantization and was produced with nvidia-modelopt **v0.48.0**.
## Training and Evaluation Datasets:
## Calibration Dataset:
**Link:** [Nemotron-Post-Training-Dataset-v3](https://huggingface.co/collections/nvidia/nemotron-post-training-v3)
**Data Modality:** Text
**Data Collection Method by dataset:** Varies by dataset.
**Labeling Method by dataset:** Varies by dataset.
**Properties:** The Nemotron-Post-Training-Dataset-v3 is a multi-million-sample corpus developed by NVIDIA for Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) to power alignment, reasoning, and agentic capabilities in the Nemotron-3 model family.
## Training Dataset:
**Data Modality:** Undisclosed
**Data Collection Method by dataset:** Undisclosed
**Labeling Method by dataset:** Undisclosed
**Properties:** Undisclosed
**Image Training Data Size:** Undisclosed
**Text Training Data Size:** Undisclosed
**Video Training Data Size:** Undisclosed
## Evaluation Dataset:
**Datasets:** GPQA Diamond, Terminal-Bench, AA-LCR, MMMU-Pro, SciCode, IFBench
**Data Collection Method by dataset:** Hybrid: Automated, Manually-Collected
**Labeling Method by dataset:** Hybrid: Manually-Labeled, Automated
**Properties:** We evaluated the model on text-based reasoning, coding, agentic tasks, long-context recall, instruction following, and multimodal reasoning. GPQA Diamond contains graduate-level multiple-choice questions written by domain experts in biology, physics, and chemistry. Terminal-Bench evaluates agents on terminal-based tasks. AA-LCR (Artificial Analysis Long Context Recall) evaluates a model's ability to accurately retrieve and recall information from long input contexts. MMMU-Pro is a challenging multimodal understanding benchmark that measures college-level reasoning across diverse disciplines. SciCode evaluates scientific coding capabilities. IFBench evaluates instruction-following capabilities across diverse and structured task constraints.
## Inference:
**Acceleration Engine:** **vLLM**
**Test Hardware:** **NVIDIA Grace Blackwell GB300**
## Post-Training Quantization
Qwen3.8-27B was quantized using a mixed-precision recipe. NVFP4 quantization was applied to the MLP layers and language model head (`lm_head`), while FP8 quantization was applied to the self-attention and linear-attention layers. The NVFP4 layers were calibrated on 2,048 samples using the [Model Optimizer Local-Hessian calibration algorithm](https://nvidia.github.io/Model-Optimizer/reference/generated/modelopt.torch.quantization.config.html#modelopt.torch.quantization.config.LocalHessianCalibConfig).
To learn more about the algorithm, see [Local Hessian for NVFP4 Quantization](https://nvidia.github.io/Model-Optimizer/announcements/local-hessian.html). To reproduce this checkpoint, follow the command in the [Using Local Hessian](https://nvidia.github.io/Model-Optimizer/announcements/local-hessian.html#using-local-hessian) section.
## Usage
To serve this checkpoint with [vLLM](https://github.com/vllm-project/vllm), you can start the docker `vllm/vllm-openai:nightly` and run the sample command below:
```sh
vllm serve nvidia/Qwen3.8-27B-NVFP4 \
--port 8000 \
--kv-cache-dtype fp8_e4m3 \
--tensor-parallel-size 4 \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--mm-encoder-tp-mode data \
--seed 0 \
--gpu-memory-utilization 0.85 \
--max-num-seqs 32 \
--max-num-batched-tokens 32768 \
--enable-chunked-prefill
```
To serve this checkpoint with [SGLang](https://github.com/sgl-project/sglang), you can start the docker `lmsysorg/sglang:dev` and run the sample command below:
```sh
sglang serve \
--trust-remote-code \
--model-path nvidia/Qwen3.8-27B-NVFP4 \
--kv-cache-dtype fp8_e4m3 \
--mem-fraction-static 0.85 \
--chunked-prefill-size 2048 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--mamba-full-memory-ratio 4.59 \
--host 0.0.0.0 \
--port 30000 \
--mamba-radix-cache-strategy extra_buffer \
--mamba-ssm-dtype float32
```
For more details please refer to [SGLang cookbook](https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-27B).
## Evaluation
The accuracy benchmark results are presented in the table below:
| Benchmark | Qwen3.8-27B BF16 | Qwen3.8-27B NVFP4 |
|---|---|---|
| GPQA Diamond | 88.92 | 88.01 |
| Terminal-Bench | 75.56 | 74.02 |
| AA-LCR | 72.63 | 73.38 |
| MMMU-Pro | 75.14 | 74.86 |
| SciCode | 47.93 | 48.41 |
| IFBench | 80.07 | 78.93 |