Text Generation
GGUF
llama.cpp
conversational
on-device
mobile
iphone
android
cpu
local-llm
edge
mixture-of-experts
Mixture of Experts
quantized
pocket
vidraft
qwen3_5_moe
darwin
imatrix
Instructions to use FINAL-Bench/POCKET-35B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use FINAL-Bench/POCKET-35B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf FINAL-Bench/POCKET-35B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf FINAL-Bench/POCKET-35B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf FINAL-Bench/POCKET-35B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf FINAL-Bench/POCKET-35B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use FINAL-Bench/POCKET-35B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FINAL-Bench/POCKET-35B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FINAL-Bench/POCKET-35B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
- Ollama
How to use FINAL-Bench/POCKET-35B-GGUF with Ollama:
ollama run hf.co/FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use FINAL-Bench/POCKET-35B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "FINAL-Bench/POCKET-35B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use FINAL-Bench/POCKET-35B-GGUF with Docker Model Runner:
docker model run hf.co/FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
- Lemonade
How to use FINAL-Bench/POCKET-35B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.POCKET-35B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use FINAL-Bench/POCKET-35B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use FINAL-Bench/POCKET-35B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FINAL-Bench/POCKET-35B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "FINAL-Bench/POCKET-35B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| license: apache-2.0 | |
| library_name: llama.cpp | |
| pipeline_tag: text-generation | |
| base_model: | |
| - FINAL-Bench/Darwin-36B-Opus | |
| tags: | |
| - gguf | |
| - llama.cpp | |
| - conversational | |
| - on-device | |
| - mobile | |
| - iphone | |
| - android | |
| - cpu | |
| - local-llm | |
| - edge | |
| - mixture-of-experts | |
| - moe | |
| - quantized | |
| - vidraft | |
| - qwen3_5_moe | |
| - darwin | |
| > ### π Collections | |
| > **βΆ [POCKET Models](https://huggingface.co/collections/FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6)** β this family (on-device, no GPU) | |
| > [Darwin Family](https://huggingface.co/collections/FINAL-Bench/darwin-family-699987b1f652864af0122193) Β· [Aether Foundation](https://huggingface.co/collections/FINAL-Bench/aether-foundation-model-6a5c7f2fa1a4165c0414e53a) Β· [VKAE Accelerated](https://huggingface.co/collections/FINAL-Bench/vkae-accelerated-6a47231d7e7999dd8227675a) | |
|  | |
| # POCKET-35B-GGUF | |
| ### A **35B** model that runs on your **PC with no GPU** β and on your phone. Just stock `llama.cpp`. **No fork**, no CUDA, no cloud. | |
| > π **Try it live, no install β** [](https://huggingface.co/spaces/FINAL-Bench/POCKET-35B-CPU) [](https://huggingface.co/spaces/FINAL-Bench/POCKET-26B-CPU) β both answering on a **CPU-only** box (no GPU). POCKET-26B is Gemma4-based. | |
| [](https://www.apache.org/licenses/LICENSE-2.0) [](https://github.com/ggml-org/llama.cpp) []() []() | |
| **Pick your build β** [](https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF) [](https://huggingface.co/FINAL-Bench/POCKET-KR-GGUF) [-0f6e56)](https://huggingface.co/FINAL-Bench/POCKET-KR-MLX) [](https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF) [](https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF) | |
| ## The POCKET lineup β pick by your device | |
| | Repo | File | Size | Runs on | Best for | Korean PPL* | | |
| |---|---|---|---|---|---| | |
| | **POCKET-35B-GGUF** | `Q4_K_M` | 21 GB | PC / server (32 GB RAM) | top quality | **5.79** | | |
| | **POCKET-35B-GGUF** | `Q2_K` β | 13 GB | mini-PC, no GPU | **daily driver** | 6.49 | | |
| | **POCKET-35B-GGUF** | `IQ1_M` | 8.2 GB | 16 GB RAM box | smallest full model | 9.69 | | |
| | **POCKET-KR-GGUF** | `IQ2_M` | 5.1 GB | Android 8 GB+ | π°π· Korean phone | 7.95 | | |
| | **POCKET-KR-MLX** | 2-bit | 5.1 GB | π **iPhone / iPad / Mac** | π°π· Korean, Apple-native | 7.95 | | |
| | **POCKET-EN-GGUF** | `iPhone-mix` | 5.3 GB | π iPhone (PocketPal) | π English phone | β | | |
| | **POCKET-EN-GGUF** | `PC-mix` | 6.8 GB | PC / Android | π English, best quality | β | | |
| *Wikipedia-Korean perplexity, lower is better. `Q4_K_M` = 5.79 baseline. English builds are tuned on English; see each repo. | |
| > π **Why MLX for Korean but GGUF for English on iPhone?** Apple-native MLX only does uniform quantization. Korean survives it (96 experts hold up); English needs our proprietary quantization, which only GGUF supports β so the English iPhone build ships as a GGUF you run with [PocketPal](https://github.com/a-ghorbani/pocketpal-ai). Honest, not lazy. | |
| > π **POCKET-26B** β a **Gemma4-26B-A4B**-based sibling that loads in **any app today** (Ollama Β· LM Studio Β· PocketPal Β· MLX), no bleeding-edge runtime needed: **[GGUF](https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF)** (`Q2_K` 11 GB Β· `Q4_K_M` 17 GB Β· GPQA-Diamond **67%**). Universal compatibility for 12 GB phones, PC, and browser. | |
|  | |
| ## Benchmarks β what is measured, what is not | |
| **We measure Bonsai on the same machine with the same stock `llama.cpp`, and we tell you where we lose.** | |
| `[measured]` Generation speed β POCKET wins on both CPU and GPU: | |
| | | POCKET-35B IQ1_M | Bonsai-27B Q1_0 | | | |
| |---|---|---|---| | |
| | CPU generate (Xeon, 16t) | **27.0 tok/s** | 10.1 | π’ 2.69Γ | | |
| | GPU generate (H100) | **197 tok/s** | 89 | π’ 2.22Γ | | |
| | GPU prompt (H100) | 753 | **1816** | π΄ 0.41Γ | | |
| | Quality (HellaSwag, 400q) | 61.0% | 60.0% | βͺ tie (CI overlaps) | | |
| `[measured on a MacBook M3 Pro, 18 GB]` β and on a laptop, POCKET wins **every** axis, including prompt processing: | |
| | | POCKET-35B IQ1_M | Bonsai-27B Q1_0 | | | |
| |---|---|---|---| | |
| | Metal generate (tg64) | **25.4 tok/s** | 12.8 | π’ 1.99Γ | | |
| | CPU generate (8 threads) | **13.8 tok/s** | 4.4 | π’ 3.13Γ | | |
| | Metal prompt (pp128) | **240.7 tok/s** | 73.4 | π’ 3.28Γ | | |
| | CPU prompt (pp128) | **45.5 tok/s** | 9.6 | π’ 4.75Γ | | |
| On a laptop GPU the arithmetic headroom that let Bonsai win prefill on an H100 is gone, so MoE sparsity wins across the board. `POCKET-35B-Q2_K` runs on the M3 Pro's CPU at **19.5 tok/s** β on an 18 GB Mac, run Q2_K on CPU (`-ngl 0`); its 13 GB exceeds the recommended Metal budget. | |
| `[measured β GPQA Diamond, 198q, greedy]` reasoning quality vs quantization: | |
| | Model | GPQA-Diamond (greedy) | | |
| |---|---| | |
| | Qwen3.6-35B-A3B | 73.2% | | |
| | POCKET-35B Q4_K_M | 68.7% | | |
| | POCKET-35B Q2_K | 60.1% | | |
| `[pending β community reports welcome]` on-device **iPhone** and **Strix Halo** throughput. We publish only what we ran ourselves; help us fill the rest. | |
| > The same-size rival `Ternary-Bonsai-27B-Q2_0` (7.2 GB) **fails to load in upstream llama.cpp** β it needs the PrismML fork. POCKET runs on the tools you already have. | |
| ## Files in this repo | |
| | File | Size | bpw | Runs on | Korean PPL | | |
| |---|---|---|---|---| | |
| | `POCKET-35B-Q4_K_M.gguf` | 21 GB | 4.5 | PC 32 GB RAM | 5.79 (top) | | |
| | `POCKET-35B-Q3_K_M.gguf` | 16 GB | 3.4 | PC 24 GB | 6.06 | | |
| | **`POCKET-35B-Q2_K.gguf`** β | 13 GB | 2.6 | mini-PC 16β24 GB | 6.49 (**best value**) | | |
| | `POCKET-35B-IQ1_M.gguf` | 8.2 GB | 1.9 | 16 GB RAM | 9.69 (smallest) | | |
| ## Quickstart β no fork needed | |
| ```bash | |
| # any recent llama.cpp β brew / winget / apt, or LM Studio / Ollama | |
| llama-cli -m POCKET-35B-Q2_K.gguf -p "μλ νμΈμ" -ngl 0 -t 8 | |
| # reproduce our CPU numbers: | |
| llama-bench -m POCKET-35B-IQ1_M.gguf -p 128 -n 64 -ngl 0 -t 16 | |
| ``` | |
| Use physical-core count for `-t` (max ~32). Do **not** pass all threads β it can slow down sharply. | |
| ## Lineage β where POCKET comes from | |
| POCKET is quantized from **[Darwin-36B-Opus](https://huggingface.co/FINAL-Bench/Darwin-36B-Opus)**, VIDRAFT's flagship β a model bred and evolved over several generations on the **Darwin platform** (crossbreeding, healing, expert surgery). Darwin-36B-Opus itself traces back to a Qwen3.5-family MoE architecture. | |
| | Component | Origin | | |
| |---|---| | |
| | **Starting checkpoint** | **[Darwin-36B-Opus](https://huggingface.co/FINAL-Bench/Darwin-36B-Opus)** β VIDRAFT, multi-generation Darwin evolution | | |
| | Base architecture | Qwen3.5-family MoE (256 experts, top-8), unchanged | | |
| | Quantization (`Q4_K_M`β¦`IQ1_M`) | stock llama.cpp β no custom format | | |
| | Runtime | upstream llama.cpp / Apple MLX β unmodified | | |
| | Proprietary language-specific tuning (KR/EN builds) | **ours** (VIDRAFT) | | |
| The CPU/GPU speed comes from the sparse-MoE architecture plus ordinary quantization β reproducible with the same base and the same tools. What we add is the Darwin-evolved weights, the honest measurement, the Korean tuning, and the pruning that makes the 5 GB phone builds. | |
| ## Limitations | |
| - The iPhone/Mac speed is **not yet measured by us** β community reports welcome. | |
| - Extreme quants (`IQ1_M`) hurt Korean ~2.8Γ more than English; use `Q2_K` or larger for quality. | |
| - English phone builds trade quality for size; the PC build (`PC-mix`) is much closer to full quality. | |
| ## License | |
| Apache-2.0. | |
| --- | |
| *POCKET is a VIDRAFT model family. 35B, in your pocket. No GPU.* | |
| ## Learn more | |
| - On-device LLMs without a GPU β and how POCKET measures up: [Can you run a large LLM without a GPU?](https://vidraft.net/insights/on-device-llm-without-gpu.html) | |
| - What model quantization is, and why a 4-bit model stays smart: [What is model quantization?](https://vidraft.net/insights/what-is-quantization-llm.html) | |
| <!-- POCKET-FAMILY --> | |
| --- | |
| ## π§© The POCKET Family β On-device AI by VIDRAFT | |
| *Big models, small hardware. No GPU, no cloud.* | |
| **Models** | |
| - π¦ [POCKET-35B-GGUF](https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF) β flagship, PC / server, no GPU | |
| - π¦ [POCKET-26B-GGUF](https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF) β compact 26B | |
| - π°π· [POCKET-KR-GGUF](https://huggingface.co/FINAL-Bench/POCKET-KR-GGUF) β Korean, Android | |
| - π [POCKET-KR-MLX](https://huggingface.co/FINAL-Bench/POCKET-KR-MLX) β Korean, iPhone / Mac | |
| - π [POCKET-EN-GGUF](https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF) β English, phone / PC | |
| - πΌοΈ [POCKET-Image-Zimage](https://huggingface.co/FINAL-Bench/POCKET-Image-Zimage) β character-perfect text in any image | |
| **Demos & tools (Spaces)** | |
| - π¨ [POCKET-Image Studio](https://huggingface.co/spaces/FINAL-Bench/POCKET-Image-Studio) β text-in-image, generate in-page | |
| - π₯οΈ [POCKET-35B-CPU](https://huggingface.co/spaces/FINAL-Bench/POCKET-35B-CPU) β 35B answering on a CPU | |
| - π₯οΈ [POCKET-26B-CPU](https://huggingface.co/spaces/FINAL-Bench/POCKET-26B-CPU) β 26B on a CPU | |
| π [Full POCKET collection](https://huggingface.co/collections/FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6) | |
| <!-- /POCKET-FAMILY --> | |