myJEV-9B-RL
The larger confidence-aware myJEV experiment, with one-pass decisions and a learned confidence policy.
Give it a request and descriptions of the allowed choices. It returns the chosen ID, a score for each choice and an estimate of correctness, without generating an answer paragraph. The main training task is banking support routing. Possible workflows include choosing a support queue, selecting the next workflow branch and handing uncertain cases to a reviewer after setting a threshold on your own calibration data.
By Amit Bahree. Source and study | Training findings | Inference and hosting
Why choose this version?
Choose this version to compare reinforcement learning at larger capacity. Its mean BANKING77 accuracy gain over standard 9B was small, and the paired uncertainty interval included no improvement. Confidence Brier error was also higher. It is an experimental comparison, rather than an automatic upgrade over the smaller default.
If your labels are fixed, compare a smaller classifier too. A separate one-seed ModernBERT-base control reached 90.78% BANKING77 accuracy and 0.0555 correctness Brier after temperature calibration. It saw 23,997 training examples over three epochs, versus 8,000 example presentations in these myJEV runs, so this is a practical control with a different budget, not a matched architecture comparison. Its output head fixes the 77 labels; myJEV accepts candidate descriptions with each request. That flexibility does not establish accuracy on an unfamiliar taxonomy. See the decision guide and encoder control.
This is a Qwen/Qwen3.5-9B backbone plus a small adapter, confidence heads and calibration settings. The download here contains the learned additions; the loader also downloads the separately pinned backbone. A small adapter file does not eliminate the backbone’s memory or compute cost. Serving reads the input once and performs no autoregressive generation.
Results you can compare
The released checkpoint uses seed 11 by a fixed packaging convention. It was not selected for having the best test result. Accuracy and macro-F1 below use all 3,080 examples in BANKING77’s official test split. Brier measures squared error of reported correctness confidence, where lower is better.
| Evaluation | Accuracy | Macro-F1 | Correctness Brier |
|---|---|---|---|
| This released checkpoint, seed 11 | 89.06% | 0.8882 | 0.0872 |
| Mean of seeds 11, 22 and 33 | 89.34% | 0.8911 | 0.0942 |
Accuracy’s sample standard deviation across those three seeds is 0.28 percentage points. The paired analysis reports uncertainty for method contrasts. Three observed seeds do not establish performance across all future runs or user tasks.
For this seed, warm HTTP latency was 117.55 ms p50 / 123.42 ms p95, using 100 requests, concurrency one and a short three-candidate request on one NVIDIA A30. The maximum allocated GPU memory observed for direct scoring across the three benchmark workloads was 11.24 GiB. That excludes some driver/runtime allocations and is not a maximum-context memory guarantee. Startup to readiness was 18.42 seconds with cached weights, measured once. All three sizes were validated on 24 GB A30 hardware; precision here is NF4 four-bit QLoRA. Full serving conditions and evidence.
Run it
The reference environment is Linux, Python 3.12 and an NVIDIA GPU. The pinned requirements include PyTorch, Transformers, PEFT and the four-bit runtime. A working NVIDIA driver and a host C compiler are needed for the tested GPU path. On Debian/Ubuntu, install gcc and libc6-dev if missing.
git clone https://github.com/bahree/myJEV.git
cd myJEV
git checkout c0147b69d57fbe541cdc2e1bc85c472938e4dda0
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.lock
python -m pip install --no-deps .
Use the custom DecisionModel loader. Loading only the adapter, calling ordinary generate(), or using a generic classification widget does not execute this model’s complete decision interface.
from myjev import DecisionModel
model = DecisionModel.load(
"bahree/myJEV-9B-RL",
revision="d22fdbb4b33e76037178459d523010e9b045a40c",
device="cuda:0",
)
request = {
"context": "I was charged twice.",
"instructions": "Select the appropriate support route.",
"candidates": [
{
"id": "billing",
"description": "Charges, invoices, and refunds"
},
{
"id": "technical",
"description": "Errors and configuration"
},
{
"id": "other",
"description": "Neither listed route applies"
}
]
}
result = model.score(request)
print(result["selected_id"], result["confidence"])
The recorded pinned-runtime check selected billing with confidence approximately 0.8766 on that exact fixture. It demonstrates a working call, not a guarantee on a new support taxonomy. The immutable revision above contains the tested runtime files; subsequent card-only edits leave those files unchanged.
The same request is checked into the source repository as examples/request.json:
myjev score --artifact bahree/myJEV-9B-RL \
--revision d22fdbb4b33e76037178459d523010e9b045a40c \
--input examples/request.json
# Run the local HTTP service in a separate terminal.
myjev serve --artifact bahree/myJEV-9B-RL \
--revision d22fdbb4b33e76037178459d523010e9b045a40c
# Once /readyz succeeds:
curl -fsS http://127.0.0.1:8000/score \
-H 'Content-Type: application/json' --data-binary @examples/request.json
Python, CLI and HTTP return the same response fields. Keep your artifact revision pinned in deployment configuration. Publishing this package does not start a hosted endpoint. The source repository includes tested local Docker instructions and a separate, unexecuted managed-hosting recipe.
What the scores mean
Selection scores are normalized values used to rank the supplied candidates. The selected ID is the highest-scoring choice.
Reported confidence comes from a separate candidate-conditioned policy over 21 values, from 0.00 through 1.00 in steps of 0.05. Serving first chooses the highest-scoring answer, then reports the policy’s expected confidence for that answer. It does not sample an answer or confidence at inference. This objective did not make confidence better calibrated than temperature-scaled supervised training in the study.
The API also returns artifact_revision and calibration_revision so callers can identify the exact behavior they used. Set acceptance/deferral thresholds using representative calibration data. An explicit other candidate is a classification option; confidence-based deferral is a separate decision by your application.
How it was trained
After 4,000 supervised updates, this model received 4,000 updates using an exact expected-reward objective. The reward combines whether the answer is correct with squared error in its reported confidence: correct - (confidence - correct)^2. A penalty keeps the policy close to the frozen supervised reference. The finite action space lets training calculate the expectation directly rather than estimate it by sampling. This release is the exact-RL arm, not the separate sampled-REINFORCE arm.
The matched runs use one example per optimizer update, so the initial and continuation stages each present 4,000 training examples. Training uses the grouped training portion of BANKING77, with separate validation and calibration partitions and the official test split preserved. Candidate order is randomized. The backbone remains frozen while low-rank adapter updates and custom heads learn the task. That reduces training storage; it does not turn the large pretrained backbone into a tiny inference model.
BANKING77 supplies 77 closely related banking intents. It is one task family, not a generalist instruction mixture. CLINC150 was kept outside training and tuning for transfer and unsupported-request evaluation. These released weights have no archive adaptation. The separate blog-archive and scratch-model experiments are documented in the project rather than folded into these results.
The model family
Quality columns are three-seed BANKING77 means. Runtime columns are measured on the released seed-11 artifacts under the same short-request conditions described above.
| Model | Accuracy | Confidence Brier | HTTP p50 | Allocated VRAM |
|---|---|---|---|---|
| myJEV-0.8B | 82.93% | 0.1073 | 57.48 ms | 1.55 GiB |
| myJEV-0.8B-RL | 81.36% | 0.1321 | 60.88 ms | 1.55 GiB |
| myJEV-4B | 89.23% | 0.0740 | 82.18 ms | 8.18 GiB |
| myJEV-4B-RL | 90.27% | 0.0818 | 81.85 ms | 8.18 GiB |
| myJEV-9B | 89.15% | 0.0742 | 115.09 ms | 11.24 GiB |
| myJEV-9B-RL | 89.34% | 0.0942 | 117.55 ms | 11.24 GiB |
The standard releases use continued supervised training and temperature scaling. -RL releases use exact confidence-aware reward training. The 9B models also change backbone precision to four-bit NF4, so differences cannot be attributed solely to parameter count. The study includes a separate 4B precision control.
Tested scope and limits
The manifest accepts 2 to 160 candidates and up to 4,096 tokenizer tokens for the entire rendered prompt. Duplicate IDs and oversized requests are rejected instead of silently truncated. Short 160-candidate smoke checks passed for every release; that is not evidence of equally good accuracy or latency at 160 choices. Candidate wording, order, missing correct options and quoted instructions can change decisions. Check the transfer and robustness results before choosing this model for a new workflow.
The reference backend has been checked for artifact reloads and Python/CLI/HTTP/Docker consistency, including pinned Hub download checks. New quantization, merged weights or optimized backends need their own equivalence and calibration measurements. No matched speed comparison with the proprietary Jev service has been performed.
Files, licenses and provenance
The package includes adapter/, heads.safetensors, manifest.json, LICENSE, BACKBONE_LICENSE and NOTICE.md. It does not include backbone weights, training text or optimizer state.
- Original myJEV code, adapters and heads: MIT, copyright Amit Bahree.
- Qwen backbone: Apache-2.0, with its original terms retained in
BACKBONE_LICENSE. - BANKING77: CC BY 4.0; Casanueva et al., Efficient Intent Detection with Dual Sentence Encoders (2020).
- Backbone revision:
c202236235762e1c871ad0ccb60c8ee5ba337b9a. - Artifact revision:
93499e5c987b1c4bc5278c3fa28df87dc15a56e1f92452d21db4ff143ef4fe0c. - Calibration revision:
uncalibrated.
Source repository | Dataset rationale | All releases and deployment notes
Model tree for bahree/myJEV-9B-RL
Dataset used to train bahree/myJEV-9B-RL
Evaluation results
- Accuracy (%) for released seed-11 checkpoint on BANKING77 official testtest set self-reported89.058