JEV-9B with vision (Core ML)
AutoTrust's JEV-9B System 1, with vision, converted to Core ML for Apple
silicon. It makes typed decisions over text and images: yes/no (noul), a 0–5 score, or a choice among 2–256
options. Each decision returns a calibrated probability for every option from a single forward pass; nothing is
generated.
JEV-9B with vision is the stock Qwen3.5-9B, including its unchanged vision
tower, plus JEV-9B's System 1 LoRA and decision head (vl/adapter_vllm). In this repository:
- The LoRA is merged into the decoder weights.
- The decision head is reduced to the adapted
lm_headrows it reads. - The prompt is exactly JEV's
serve_decide.pySystem 1 template.
Runtime: JevManager in FluidUse (Swift, macOS 15+), with a SwiftUI
demo where the model plays Super Mario Bros. 1-1 (JevMarioDemo).
Files
| File | Role | Size | Precision |
|---|---|---|---|
part00 … part07.mlpackage |
decoder with the System 1 LoRA merged, 4 layers per part; functions L256 and L512 (sequence buckets), weights shared |
6.5 GB total | 8-bit weights (per channel), fp16 compute |
VisionTower_P256.mlpackage, VisionTower_P576.mlpackage |
Qwen3.5-9B vision tower, one image per call, up to 256 / 576 patches (16 px) | 1.7 GB each | fp32 (fp16 is not accurate enough for this ViT) |
pos_embed.f32 |
the vision tower's learned position grid; the host resamples it per image | 10.6 MB | fp32 |
embeddings.f16 |
input token embeddings (gathered on the host) | 2.0 GB | fp16 |
head_rows.f32, head.json |
adapted lm_head rows at the 24 verbalizer ids (false/true, 0–5, A–P) plus the other single-token option labels; per-slot bias and per-kind temperature |
4.3 MB | fp32 |
tokenizer.json, config.json |
tokenizer; the manifest with shapes, buckets and token ids for the host |
Decision:
- Run the vision tower once per image.
- Splice its tokens into the prompt at the
<|image_pad|>positions. - Chain the 8 decoder parts. Their I/O is
hidden [1, L, 4096]andcos/sin [L, 64], all fp16, with interleaved M-RoPE tables built on the host. - Take the last prompt row times the head rows, add the bias, divide by the temperature, and softmax over the question's options.
Quality
The fp32 reference is Hugging Face's Qwen3_5VisionModel, plus the same decoder in fp32 with the LoRA merged
(streamed 4 layers at a time), plus the head rows. It was run on 12 Super Mario Bros. frames with a choice, a yes/no
and a score question on each:
| 8-bit Core ML vs fp32 | |
|---|---|
| answers that differ | 0 / 36 |
| max probability difference (choice / noul / score) | 0.013 / 0.023 / 0.006 |
| vision tokens, max relative error | 1.4e-4 |
Speed (M5 Pro, 24 GB, GPU; Swift JevManager)
| Median | |
|---|---|
| vision tower, one image (256 / 576-patch bucket) | ~60 / ~120 ms |
| decoder, one question (L256 bucket) | ~270 ms |
| one image + 3 questions | 0.95 s |
| load (compiled): 8 decoder parts + 2 vision towers | ~40 s |
The decoder needs ~7 GB of memory to itself. The Qwen3.5 decoder (Gated DeltaNet layers) does not run on the Neural Engine, so this is a GPU model.
Usage (Swift)
import FluidUse
let jev = try await JevManager.load(from: bundleDirectory, bucket: 256)
let result = try await jev.decide(
state: [.text("Close-up of the area just ahead of Mario in a Super Mario Bros. level:"), .image(view)],
questions: [JevQuestion(.noul, "Is there an enemy (a brown Goomba or a Koopa) in this picture?")])
print(result.decisions[0].yes ?? 0, result.totalMilliseconds)
License and credits
Apache-2.0, the same as autotrust/JEV-9B and
Qwen/Qwen3.5-9B (LICENSE is Qwen's). All weights belong to AutoTrust
and Qwen; this repository only changes their format and precision. Conversion by
Fluid Inference.
- Downloads last month
- -