qwen2.5-coder-0.5b-coreml

Qwen2.5-Coder-0.5B-Instruct converted to Core ML: it writes Python from a plain-English request, fully on device. One 944 MB fp16 package with two functions that share the weights:

Function Shape Runs on p50 (M5 Pro, macOS 27, from Swift)
prefill 448 tokens, fixed Neural Engine (1,451 / 1,456 ops) 52 ms
decode 1 token, 1,024-slot KV cache in MLState GPU 25–28 ms / token (36–40 tok/s)

A typical function (100–330 tokens) is written in 3–9 s. The prompt runs on the Neural Engine; the token-by-token writing runs on the GPU, because the Neural Engine rejects the decoder's in-graph KV-cache write. By time, the Neural Engine carries about 1% of a task.

Quality

Greedy decoding, pass@1, same prompt and test harness for both rows:

Benchmark PyTorch fp32 This Core ML package
HumanEval (164) 90 / 164 = 54.9% 89 / 164 = 54.3%
MBPP sanitized, test split (257) β€” 119 / 257 = 46.3%

On HumanEval, 142 of 164 completions are token-for-token identical to PyTorch; one task flips from pass to fail. Qwen reports 61.6% HumanEval for this model with its own evaluation harness; the numbers above use plain greedy decoding (repetition_penalty 1.0) and a simple code extractor, applied identically to both rows.

Files

  • qwen2_5_coder_0_5b.mlpackage β€” multifunction Core ML package (prefill, decode), token-id inputs (embedding lookup and RoPE tables are inside the graph).
  • tokenizer.json β€” Qwen2.5 byte-level BPE.
  • config.json β€” host settings: prefill length 448, cache 1,024, pad / stop ids, max new tokens 512, system prompt.

Demo

CodeWriterDemo in FluidInference/FluidUse: a list of ten short Python tasks (from MBPP, one per Python feature: list comprehension, slicing, dict, string methods, sorted, …). The model writes each one live into an editor, then python3 runs the task's asserts. The demo list is hand-picked for the recording; the MBPP row above is the unbiased number.

swift run -c release CodeWriterDemo   # downloads this repository on first launch

Use

Swift host: CodeWriterManager in the same repository.

let writer = try await CodeWriterManager.load(from: modelDirectory)
let result = try await writer.write(task: "Write a function to reverse each string in a list.") { text in
    print(text)  // the text so far, after every token
}
print(result.code, result.timing.tokensPerSecond)

Prompt (Qwen2.5 chat template):

<|im_start|>system
{config.systemPrompt}<|im_end|>
<|im_start|>user
{task}<|im_end|>
<|im_start|>assistant

Left-pad the token ids to 448 with padID, mask padded columns, run prefill, copy its keys / values ([24, 2, 448, 64] fp16) into the decoder state (k_cache_i / v_cache_i, [1, 2, 1024, 64]), then call decode one token at a time (input_ids, attention_mask [1, 1, 1, 1024], cache_position) until <|im_end|>. Greedy decoding reproduces the benchmarked completions.

Provenance

Base: Qwen/Qwen2.5-Coder-0.5B-Instruct (Apache-2.0), weights unchanged apart from fp16 storage. Conversion: coremltools 9.0, torch 2.7.0, minimum_deployment_target macOS 15 / iOS 18. Benchmarks: HumanEval (MIT), MBPP (CC-BY-4.0).

Downloads last month
2
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for FluidInference/qwen2.5-coder-0.5b-coreml

Quantized
(93)
this model