Qwen3-Embedding-0.6B โ€” int4-blockwise ONNX (code-optimized)

Quantized ONNX build of Qwen/Qwen3-Embedding-0.6B targeting local code-search embeddings in agentic tools: 331 MB total (graph 0.8 MB + external weights 331 MB), down from 1.19 GiB bf16 โ€” a ~3.9ร— reduction with retrieval behavior preserved against the previous dynamic-uint8 build.

Files

File Size What
onnx/qwen3-embed-0.6b-dynamic-uint8.onnx 0.8 MB graph (external-data format)
onnx/qwen3-embed-0.6b-dynamic-uint8.onnx_data 331 MB weights โ€” required next to the .onnx
onnx/tokenizer.json 10.9 MB Qwen3 tokenizer
variants/dynamic-uint8-legacy.onnx 585 MB previous dynamic-uint8 single-file build
PERFORMANCE_DATA.json โ€” raw diff + retrieval-battery results

The .onnx file is useless without .onnx_data in the same directory.

Usage (onnxruntime, CPU or CUDA EP)

import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer

tok = Tokenizer.from_file("onnx/tokenizer.json")
sess = ort.InferenceSession("onnx/qwen3-embed-0.6b-dynamic-uint8.onnx",
                            providers=["CPUExecutionProvider"])  # or CUDAExecutionProvider

def embed(texts: list[str]) -> np.ndarray:
    encs = [tok.encode(t) for t in texts]
    S = max(len(e.ids) for e in encs)
    ids = np.zeros((len(encs), S), dtype=np.int64)
    am  = np.zeros((len(encs), S), dtype=np.int64)
    for i, e in enumerate(encs):
        ids[i, :len(e.ids)] = e.ids
        am[i, :len(e.ids)] = e.attention_mask
    feed = {"input_ids": ids, "attention_mask": am,
            "position_ids": np.tile(np.arange(S, dtype=np.int64), (len(encs), 1))}
    for inp in sess.get_inputs():            # zero-initialize KV-cache states
        if "past_key_values" in inp.name:
            dt = np.float16 if "float16" in inp.type else np.float32
            shape = [int(d) if str(d).isdigit() else 0 for d in inp.shape]
            feed[inp.name] = np.zeros(shape, dtype=dt)
    out = sess.run(None, feed)[0]
    seq = am.sum(axis=1) - 1                 # last-token pooling (Qwen3-Embedding)
    emb = out[np.arange(out.shape[0]), seq]
    return emb / np.clip(np.linalg.norm(emb, axis=-1, keepdims=True), 1e-12, None)

for query in ["handle async retries with exponential backoff",
              "debounce a function call"]:
    print(embed([query]).shape)  # (1, 1024) L2-normalized

Task instruction prefix (recommended for queries, per upstream Qwen3-Embedding): "Instruct: {task_description}\nQuery: {query}".

Quantization

Derived from the bf16 upstream via dynamic-uint8 (MatMulInteger) ONNX export, then restructured to int4-blockwise weights with external-data storage. ~4.4 bits/param effective. Blockwise group quantization methodology + quality-gate discipline documented in the accompanying research notes; gates were paired renders/queries scored against the unquantized baseline with a bf16-vs-bf16 noise-floor control.

Measured: int4-blockwise vs dynamic-uint8 build

Code-retrieval battery: 10 queries ร— 10 TypeScript/Node docs (retry/backoff, sqlite WAL transactions, minhash dedup, NDJSON streaming, async locks, env validation, path safetyโ€ฆ), last-token pooled, L2-normalized, same inputs to both builds:

Metric Result
Mean doc-embedding cosine (current vs uint8) 0.638 (min 0.583)
Top-1 retrieval agreement 10/10 queries
Retrieval correctness (both builds) 10/10 queries
Mean Kendall ฯ„ (rank order agreement) 0.60

Interpretation: absolute embedding directions drift under int4 (expected โ€” see limitation), but the functionally relevant signal survives: both builds retrieve the identical correct document for every query. Re-embed any existing index after switching.

Limitations

  • Not index-compatible with other builds: mean cosine to the dynamic-uint8 build is 0.64 โ€” re-embed your corpus.
  • Absolute embedding drift is substantial even though ranking is preserved; do not mix embeddings from different builds in one vector space.
  • Verified on 10-query/10-doc synthetic code battery โ€” directional evidence, not a full benchmark (MTEB/CoIR run pending).
  • Weight-only quantization: activations remain fp32; MatMulNBits-style int4 requires onnxruntime โ‰ฅ 1.16.
  • Multilingual capability inherited from upstream; only EN verified in this evaluation.

Provenance & license

Base: Qwen/Qwen3-Embedding-0.6B (Apache-2.0). Quantization tooling: onnxruntime blockwise quantization + custom verification harness. This derivative inherits Apache-2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ScooterLacroix/qwen3-embed-0.6b-int4-code

Quantized
(261)
this model