Presets

Configuration and automation for local LLM and image-generation inference on a dual-NVIDIA GPU host. The repository includes llama.cpp model presets, a llama-swap gateway configuration, and Docker-based build/run scripts for llama.cpp, stable-diffusion.cpp, and Unsloth Studio.

Hardware

Linux host has:

  • RTX 2080 SUPER: CUDA device 0, compute capability sm_75, ~8 GB VRAM
  • GTX 980: CUDA device 1, compute capability sm_52, ~4 GB VRAM
  • Proprietary NVIDIA driver 580
  • CUDA 12.9.1

llama.cpp CUDA builds target both GPUs with:

CMAKE_CUDA_ARCHITECTURES=52;75

The host keeps CUDA 13.4. The llama.cpp build uses CUDA 12.9 inside an Ubuntu 24.04 container because CUDA 12.9 is incompatible with host Ubuntu 26 headers.

Preset files

The repository has two preset layers:

  • preset.ini is the multi-model llama.cpp router preset used by the optional standalone service in scripts/llama-server.sh.
  • llama-swap-presets/*.ini contains one model configuration per file. These are the presets launched by the llama-swap gateway defined in llama-swap.yaml.

The root server.sh controls the gateway and selects models by their llama-swap model ID; it does not take a preset filename. The standalone service can select a preset explicitly, for example:

./scripts/llama-server.sh start

preset.ini

preset.ini defines a llama.cpp router with these active model sections:

  • Qwen3.6-35B
  • Ornith-1.5-35B
  • Gemma-4-12B
  • Gemma-4-26B
  • Qwopus3.6-35B

The [*] section supplies defaults, while a named model section overrides those defaults.

The main defaults in preset.ini include:

c = 131072
image-min-tokens = 1024
device = CUDA0,CUDA1
split-mode = layer
main-gpu = 0
device-draft = CUDA1
mmproj-device = CUDA1
models-max = 1

CUDA0 is the RTX 2080 SUPER and CUDA1 is the GTX 980. Layer splitting keeps the model distributed across both GPUs, while draft-model and multimodal projector work is assigned to CUDA1. The GTX 980 is connected over PCIe x1.

llama-swap-presets/

Each single-model preset has a [*] section for shared llama.cpp settings and one named section containing the Hugging Face model, sampling parameters, context size, and speculative-decoding configuration. The active gateway presets are:

File Model section Context Gateway status
qwen3.6-35b-mtp.ini Qwen3.6-35B 131,072 Active
qwen3.8-27b-pi.ini Qwen3.8-27B-Pi 65,536 Active
ornith-1.5-35b-a3b-mtp.ini Ornith-1.5-35B 131,072 Active
qwopus3.6-35b-mtp.ini Qwopus3.6-35B 131,072 Active
gemma-4-12b-it-qat.ini Gemma-4-12B 65,536 Active
gemma-4-26b-a4b-it-qat-mtp.ini Gemma-4-26B 98,304 Active
qwen3.8-flash-next.ini Qwen3.8-Flash-Next 32,768 Not currently exposed; YAML entry is commented out
laya-bf16.ini Laya Default Not currently listed in llama-swap.yaml

The hf setting identifies the Hugging Face repository and model variant, for example unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL. The model files are resolved through the container's /models Hugging Face cache. Set HF_TOKEN in .env when a model requires authentication.

Common preset settings include:

  • c: context length; model-specific values override the global default.
  • device, split-mode, and main-gpu: GPU placement and layer splitting.
  • device-draft and spec-type: speculative decoding device and method.
  • spec-draft-*: speculative draft limits and acceptance thresholds.
  • load-mode and lazy-mode: model loading and memory-mapping behavior.
  • ctx-checkpoints and checkpoint-min-step: long-context checkpointing for the MTP presets that enable it.
  • reasoning, reasoning-format, and chat-template-kwargs: reasoning and chat-template behavior.
  • ctk and ctv: key/value cache quantization for selected low-memory presets.

Changes to a preset affect the next process start. Gateway model IDs, preset paths, context values, and capabilities must remain consistent between the selected file and its entry in llama-swap.yaml.

Configuration

Scripts load .env from the repository root. Copy the template before using any build or service script:

cp .env.example .env

Keep .env local; it may contain access tokens, API keys, and machine-specific paths. Commit .env.example, not .env. Values such as PRESETS_DIR="${MODELS_ROOT}/presets" are shell expansions, so define the base variables before the derived variables.

.env.example settings

Paths

Variable Purpose
MODELS_ROOT Base directory containing the source checkouts, presets, model cache, and Unsloth work directory.
PRESETS_DIR Presets repository directory mounted at /presets; normally the checkout containing this README.
LLAMA_DIR Existing llama.cpp Git checkout used for builds and mounted at /src. Set this before using scripts/build-from-source.sh llama.cpp.
STABLE_DIFFUSION_DIR Existing stable-diffusion.cpp Git checkout used for image-generation builds and mounted at /stable-diffusion or /src. Set this before using scripts/build-from-source.sh stable-diffusion.cpp.
MODELS_DIR Host directory mounted at /models, including the Hugging Face model cache and diffusion assets.

LLAMA_DIR and STABLE_DIFFUSION_DIR must point to existing Git checkouts; build-from-source.sh updates them with Git but does not clone them. The llama-swap gateway requires both completed builds because it launches the custom llama-server and sd-server binaries.

CUDA and Docker images

Variable Purpose
CUDA_IMAGE Base CUDA development image used when building the local builder image.
BUILDER_IMAGE Local Docker image tag produced by scripts/create-build-image.sh and used by the llama.cpp and stable-diffusion server scripts.
CUDA_BUILD Build-directory suffix, such as 12 for build-12; it must match the binaries built in both source checkouts.
CUDA_ARCHITECTURES Semicolon-separated CMake CUDA targets; the template uses 52;75 for the GTX 980 and RTX 2080 SUPER.
CUDA_FLAGS Additional flags passed to CMake during CUDA builds.

Container names

Variable Purpose
BUILD_CONTAINER_NAME Name of the temporary Compose build container.
SERVER_CONTAINER_NAME Name used by scripts/llama-server.sh for the standalone llama.cpp server.
BENCH_CONTAINER_PREFIX Prefix reserved for benchmark containers.
LLAMA_SWAP_CONTAINER_NAME Name used by the root server.sh llama-swap gateway.

Standalone server and RPC settings

Variable Purpose
SERVER_HOST Address used by the standalone llama.cpp and stable-diffusion servers inside their containers.
SERVER_PORT Host port and server listen port used by the standalone server scripts.
SERVER_INTERNAL_PORT Container-side published port. The current server scripts pass SERVER_PORT to the process as well, so keep these two values equal unless the scripts are changed together.
RPC_HOSTS Host value passed to ggml-rpc-server, normally HOST when using host networking.
RPC_PORT RPC server listening port.

Authentication and Hugging Face access

Variable Purpose
HF_TOKEN Optional Hugging Face access token passed to model-serving and Unsloth containers. Set it for gated or private repositories.
LLAMA_API_KEY API key passed to the standalone llama.cpp server and used as the default source for LLAMA_SWAP_API_KEY.

llama-swap gateway

Variable Purpose
LLAMA_SWAP_VERSION Tag selected for the llama-swap image, such as unified-cuda.
LLAMA_SWAP_IMAGE Full gateway image reference; normally derived from LLAMA_SWAP_VERSION.
LLAMA_SWAP_PORT Host port published to the gateway's container port 8080.
LLAMA_SWAP_API_KEY Gateway API key. The template defaults it to LLAMA_API_KEY; it must be non-empty when starting the gateway.

Unsloth Studio

Variable Purpose
UNSLOTH_IMAGE Unsloth Studio/JupyterLab image.
UNSLOTH_CONTAINER_NAME Unsloth container name.
UNSLOTH_WORK_DIR Host directory mounted as the Unsloth workspace.
HF_CACHE_DIR Host Hugging Face cache mounted into the Unsloth container.
UNSLOTH_BIND_ADDRESS Host bind address for Studio and JupyterLab; use 127.0.0.1 with SSH forwarding.
UNSLOTH_STUDIO_HOST_PORT Host port mapped to Studio's container port 8000.
UNSLOTH_JUPYTER_HOST_PORT Host port mapped to JupyterLab's container port 8888.
UNSLOTH_STUDIO_PASSWORD Unsloth Studio password.
JUPYTER_PASSWORD JupyterLab password.
UNSLOTH_STUDIO_BOOTSTRAP_TIMEOUT Maximum Studio bootstrap time in seconds.

Service lifecycle

Variable Purpose
STOP_TIMEOUT Graceful Docker stop timeout in seconds for the service scripts.

Key files:

File Purpose
.env Local machine configuration; ignored by Git
.env.example Portable configuration template

Docker setup

Host requirements:

  • Docker
  • NVIDIA Container Toolkit
  • Proprietary NVIDIA driver 580

Before building, clone llama.cpp into LLAMA_DIR; for gateway image generation, also clone stable-diffusion.cpp into STABLE_DIFFUSION_DIR. scripts/build-from-source.sh updates existing Git checkouts; it does not clone them.

Build image once after changing docker/Dockerfile:

./scripts/create-build-image.sh

Files:

File Purpose
docker/Dockerfile CUDA 12.9.1 build image with CMake, OpenSSL, GCC, and ccache
docker/docker-compose.llamacpp.yml GPU-enabled llama.cpp build service
docker/docker-compose.stabledifusion.yml GPU-enabled stable-diffusion.cpp build service
scripts/build-from-source.sh Pull/update the selected source tree and run its containerized build

Unsloth Studio

docker/docker-compose.unsloth.yml runs the official Unsloth image with Unsloth Studio and JupyterLab. Configure the UNSLOTH_* values in .env, set UNSLOTH_STUDIO_PASSWORD and JUPYTER_PASSWORD, then start it with:

./scripts/unsloth.sh start

The script defaults to start; use ./scripts/unsloth.sh stop, ./scripts/unsloth.sh restart, ./scripts/unsloth.sh status, or ./scripts/unsloth.sh logs for container lifecycle management.

By default, Studio is available at http://<server-host>:8000 and JupyterLab at http://<server-host>:8888; change UNSLOTH_STUDIO_HOST_PORT or UNSLOTH_JUPYTER_HOST_PORT in .env to use different ports. The service persists the Hugging Face cache, Studio data, Triton kernels, and files under UNSLOTH_WORK_DIR. The published ports default to 0.0.0.0; set UNSLOTH_BIND_ADDRESS="127.0.0.1" when using SSH port forwarding instead of exposing the services to the LAN.

The CUDA builder image includes libssl-dev, so llama.cpp can download Hugging Face models over HTTPS.

Scripts

The llama-swap gateway is managed by the root server.sh script. Supporting scripts are under scripts/: manage-models.py, build-from-source.sh, check-hf-updates.py, create-build-image.sh, llama-server.sh, rpc-server.sh, stable-server.sh, and unsloth-studio.sh.

Build (Linux)

./scripts/create-build-image.sh                    # First time or after docker/Dockerfile changes
./scripts/build-from-source.sh llama.cpp           # Build standard llama.cpp
./scripts/build-from-source.sh stable-diffusion.cpp  # Required for gateway image generation

Manage gateway models

For a llama.cpp text model, use scripts/manage-models.py add to create a single-model preset and register it in llama-swap.yaml in one step. Pass a Hugging Face repository reference with an optional model variant:

./scripts/manage-models.py add "unsloth/GLM-5.3-GGUF:UD-Q4_K_XL"

This derives the llama-swap model ID GLM-5.3, creates llama-swap-presets/glm-5.3-ud-q4-k-xl.ini, adds the model to the exclusive single-model group, and appends its gateway configuration and log path to llama-swap.yaml. The generated preset uses a 65,536-token context, text input/output, tools enabled, dual-GPU layer splitting, and no speculative decoding by default.

Adjust the defaults when needed:

./scripts/manage-models.py add \\
  "unsloth/GLM-5.3-GGUF:UD-Q4_K_XL" \\
  --model-id "GLM-5.3" \\
  --context 131072 \\
  --no-tools

For a Stable Diffusion model, pass --type stable-diffusion. This creates an SDAPI gateway entry instead of an INI preset. The diffusion-model and LLM paths must be valid inside the gateway container, usually under /models:

./scripts/manage-models.py add \\
  --type stable-diffusion \\
  "unsloth/Z-Image-Turbo-GGUF:Q4_K_M" \\
  --model-id "Z-Image-Turbo" \\
  --diffusion-model /models/hub/models--unsloth--Z-Image-Turbo-GGUF/snapshots/6c80814333b7b6a70a2e5b469a7c6437ce65de0f/z-image-turbo-Q4_K_M.gguf \\
  --vae /models/ae.safetensors \\
  --llm /models/hub/models--unsloth--Qwen3-4B-Instruct-2507-GGUF/snapshots/a06e946bb6b655725eafa393f4a9745d460374c9/Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf \\
  --backend "diffusion=cuda0&cuda1" \\
  --max-vram "cuda0=7,cuda1=3.5"

--vae defaults to /models/ae.safetensors; the default backend is diffusion=cuda0&cuda1, and the default VRAM limit is cuda0=7,cuda1=3.5. Stable Diffusion entries use ${sd_server}, the /sdapi/v1/samplers health check, and image input/output capabilities. The Hugging Face reference is used to derive the model ID; the script does not download the model or infer the VAE and LLM paths.

Remove a model setup by its exact llama-swap model ID:

./scripts/manage-models.py remove "GLM-5.3"

Removal deletes the model entry from llama-swap.yaml, removes it from the exclusive single-model group, and deletes the managed preset file. Existing model cache files and llama-swap logs are intentionally left untouched. The script can be run from any working directory, refuses duplicate additions, and refuses to remove an unknown model ID. Review generated presets and capabilities before starting a new model; model-specific sampling, reasoning, multimodal, and speculative-decoding settings cannot be inferred from the Hugging Face repository reference alone.

Hugging Face cache updates

scripts/check-hf-updates.py checks cached Hugging Face model repositories for new revisions on main. It runs in check-only mode by default and reports the cached files affected by an update without downloading anything:

./scripts/check-hf-updates.py

Use --update to update only files that are already cached locally and still exist in the remote revision:

./scripts/check-hf-updates.py --update

Newly added repository files are not downloaded. Deleted upstream files are reported and omitted. If no cached files remain in a repository, the script still advances the cached main revision without downloading new model files. The command exits nonzero only when one or more repositories fail to process.

The script automatically creates a virtual environment and installs huggingface_hub when needed. Set HF_VENV_DIR to choose its location; otherwise it uses $XDG_CACHE_HOME/hf-cache-updates-venv or ~/.cache/hf-cache-updates-venv.

The Hugging Face cache is selected in this order:

  1. HF_HUB_CACHE
  2. ${HF_HOME}/hub
  3. ~/.cache/huggingface/hub

Hugging Face authentication is read from the normal environment and local configuration used by huggingface_hub.

llama-swap gateway

server.sh exposes the configured models from llama-swap.yaml. The following six text models and the image model are members of the exclusive single-model group:

  • Qwen3.6-35B
  • Qwen3.8-27B-Pi
  • Ornith-1.5-35B
  • Gemma-4-12B
  • Gemma-4-26B
  • Qwopus3.6-35B
  • Z-Image-Turbo

Z-Image-Turbo uses the locally built sd-server from STABLE_DIFFUSION_DIR, mounted read-only into the gateway container. It uses the same model paths and multi-GPU settings as scripts/stable-server.sh, and is placed in the exclusive single-model group so image generation unloads any active model in that group before using the GPUs. Build stable-diffusion.cpp first with ./scripts/build-from-source.sh stable-diffusion.cpp.

On the target server, from the presets directory configured in .env:

./server.sh start
./server.sh status
./server.sh restart
./server.sh stop

Quick Start

cd <presets-directory>
./scripts/create-build-image.sh
./scripts/build-from-source.sh llama.cpp
./scripts/build-from-source.sh stable-diffusion.cpp
./server.sh start
./server.sh status

For image generation, select Z-Image-Turbo as the model and use the OpenAI-compatible images endpoint:

{
  "model": "Z-Image-Turbo",
  "prompt": "a watercolor painting of a mountain cabin",
  "size": "1024x1024"
}

The Stable Diffusion WebUI-compatible endpoints are also available through the same gateway, including /sdapi/v1/txt2img and /sdapi/v1/img2img.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support