qwen2.5-coder-0.5b-coreml
Qwen2.5-Coder-0.5B-Instruct converted to Core ML: it writes Python from a plain-English request, fully on device. One 944 MB fp16 package with two functions that share the weights:
| Function | Shape | Runs on | p50 (M5 Pro, macOS 27, from Swift) |
|---|---|---|---|
prefill |
448 tokens, fixed | Neural Engine (1,451 / 1,456 ops) | 52 ms |
decode |
1 token, 1,024-slot KV cache in MLState |
GPU | 25β28 ms / token (36β40 tok/s) |
A typical function (100β330 tokens) is written in 3β9 s. The prompt runs on the Neural Engine; the token-by-token writing runs on the GPU, because the Neural Engine rejects the decoder's in-graph KV-cache write. By time, the Neural Engine carries about 1% of a task.
Quality
Greedy decoding, pass@1, same prompt and test harness for both rows:
| Benchmark | PyTorch fp32 | This Core ML package |
|---|---|---|
| HumanEval (164) | 90 / 164 = 54.9% | 89 / 164 = 54.3% |
| MBPP sanitized, test split (257) | β | 119 / 257 = 46.3% |
On HumanEval, 142 of 164 completions are token-for-token identical to PyTorch; one task flips from pass to fail.
Qwen reports 61.6% HumanEval for this model with its own evaluation harness; the numbers above use plain greedy
decoding (repetition_penalty 1.0) and a simple code extractor, applied identically to both rows.
Files
qwen2_5_coder_0_5b.mlpackageβ multifunction Core ML package (prefill,decode), token-id inputs (embedding lookup and RoPE tables are inside the graph).tokenizer.jsonβ Qwen2.5 byte-level BPE.config.jsonβ host settings: prefill length 448, cache 1,024, pad / stop ids, max new tokens 512, system prompt.
Demo
CodeWriterDemo in FluidInference/FluidUse: a list of ten short
Python tasks (from MBPP, one per Python feature: list comprehension, slicing, dict, string methods, sorted, β¦).
The model writes each one live into an editor, then python3 runs the task's asserts. The demo list is
hand-picked for the recording; the MBPP row above is the unbiased number.
swift run -c release CodeWriterDemo # downloads this repository on first launch
Use
Swift host: CodeWriterManager in the same repository.
let writer = try await CodeWriterManager.load(from: modelDirectory)
let result = try await writer.write(task: "Write a function to reverse each string in a list.") { text in
print(text) // the text so far, after every token
}
print(result.code, result.timing.tokensPerSecond)
Prompt (Qwen2.5 chat template):
<|im_start|>system
{config.systemPrompt}<|im_end|>
<|im_start|>user
{task}<|im_end|>
<|im_start|>assistant
Left-pad the token ids to 448 with padID, mask padded columns, run prefill, copy its keys / values
([24, 2, 448, 64] fp16) into the decoder state (k_cache_i / v_cache_i, [1, 2, 1024, 64]), then call
decode one token at a time (input_ids, attention_mask [1, 1, 1, 1024], cache_position) until
<|im_end|>. Greedy decoding reproduces the benchmarked completions.
Provenance
Base: Qwen/Qwen2.5-Coder-0.5B-Instruct (Apache-2.0),
weights unchanged apart from fp16 storage. Conversion: coremltools 9.0, torch 2.7.0, minimum_deployment_target
macOS 15 / iOS 18. Benchmarks: HumanEval (MIT),
MBPP (CC-BY-4.0).
- Downloads last month
- 2
Model tree for FluidInference/qwen2.5-coder-0.5b-coreml
Base model
Qwen/Qwen2.5-0.5B