Instructions to use ScooterLacroix/qwen3-embed-0.6b-int4-code with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ScooterLacroix/qwen3-embed-0.6b-int4-code with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="ScooterLacroix/qwen3-embed-0.6b-int4-code")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ScooterLacroix/qwen3-embed-0.6b-int4-code", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Qwen3-Embedding-0.6B โ int4-blockwise ONNX (code-optimized)
Quantized ONNX build of Qwen/Qwen3-Embedding-0.6B targeting local code-search embeddings in agentic tools: 331 MB total (graph 0.8 MB + external weights 331 MB), down from 1.19 GiB bf16 โ a ~3.9ร reduction with retrieval behavior preserved against the previous dynamic-uint8 build.
Files
| File | Size | What |
|---|---|---|
onnx/qwen3-embed-0.6b-dynamic-uint8.onnx |
0.8 MB | graph (external-data format) |
onnx/qwen3-embed-0.6b-dynamic-uint8.onnx_data |
331 MB | weights โ required next to the .onnx |
onnx/tokenizer.json |
10.9 MB | Qwen3 tokenizer |
variants/dynamic-uint8-legacy.onnx |
585 MB | previous dynamic-uint8 single-file build |
PERFORMANCE_DATA.json |
โ | raw diff + retrieval-battery results |
The .onnx file is useless without .onnx_data in the same directory.
Usage (onnxruntime, CPU or CUDA EP)
import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer
tok = Tokenizer.from_file("onnx/tokenizer.json")
sess = ort.InferenceSession("onnx/qwen3-embed-0.6b-dynamic-uint8.onnx",
providers=["CPUExecutionProvider"]) # or CUDAExecutionProvider
def embed(texts: list[str]) -> np.ndarray:
encs = [tok.encode(t) for t in texts]
S = max(len(e.ids) for e in encs)
ids = np.zeros((len(encs), S), dtype=np.int64)
am = np.zeros((len(encs), S), dtype=np.int64)
for i, e in enumerate(encs):
ids[i, :len(e.ids)] = e.ids
am[i, :len(e.ids)] = e.attention_mask
feed = {"input_ids": ids, "attention_mask": am,
"position_ids": np.tile(np.arange(S, dtype=np.int64), (len(encs), 1))}
for inp in sess.get_inputs(): # zero-initialize KV-cache states
if "past_key_values" in inp.name:
dt = np.float16 if "float16" in inp.type else np.float32
shape = [int(d) if str(d).isdigit() else 0 for d in inp.shape]
feed[inp.name] = np.zeros(shape, dtype=dt)
out = sess.run(None, feed)[0]
seq = am.sum(axis=1) - 1 # last-token pooling (Qwen3-Embedding)
emb = out[np.arange(out.shape[0]), seq]
return emb / np.clip(np.linalg.norm(emb, axis=-1, keepdims=True), 1e-12, None)
for query in ["handle async retries with exponential backoff",
"debounce a function call"]:
print(embed([query]).shape) # (1, 1024) L2-normalized
Task instruction prefix (recommended for queries, per upstream Qwen3-Embedding):
"Instruct: {task_description}\nQuery: {query}".
Quantization
Derived from the bf16 upstream via dynamic-uint8 (MatMulInteger) ONNX export, then restructured to int4-blockwise weights with external-data storage. ~4.4 bits/param effective. Blockwise group quantization methodology + quality-gate discipline documented in the accompanying research notes; gates were paired renders/queries scored against the unquantized baseline with a bf16-vs-bf16 noise-floor control.
Measured: int4-blockwise vs dynamic-uint8 build
Code-retrieval battery: 10 queries ร 10 TypeScript/Node docs (retry/backoff, sqlite WAL transactions, minhash dedup, NDJSON streaming, async locks, env validation, path safetyโฆ), last-token pooled, L2-normalized, same inputs to both builds:
| Metric | Result |
|---|---|
| Mean doc-embedding cosine (current vs uint8) | 0.638 (min 0.583) |
| Top-1 retrieval agreement | 10/10 queries |
| Retrieval correctness (both builds) | 10/10 queries |
| Mean Kendall ฯ (rank order agreement) | 0.60 |
Interpretation: absolute embedding directions drift under int4 (expected โ see limitation), but the functionally relevant signal survives: both builds retrieve the identical correct document for every query. Re-embed any existing index after switching.
Limitations
- Not index-compatible with other builds: mean cosine to the dynamic-uint8 build is 0.64 โ re-embed your corpus.
- Absolute embedding drift is substantial even though ranking is preserved; do not mix embeddings from different builds in one vector space.
- Verified on 10-query/10-doc synthetic code battery โ directional evidence, not a full benchmark (MTEB/CoIR run pending).
- Weight-only quantization: activations remain fp32; MatMulNBits-style int4 requires onnxruntime โฅ 1.16.
- Multilingual capability inherited from upstream; only EN verified in this evaluation.
Provenance & license
Base: Qwen/Qwen3-Embedding-0.6B (Apache-2.0). Quantization tooling: onnxruntime blockwise quantization + custom verification harness. This derivative inherits Apache-2.0.