Presets
Configuration and automation for local LLM and image-generation inference on a dual-NVIDIA GPU host. The repository includes llama.cpp model presets, a llama-swap gateway configuration, and Docker-based build/run scripts for llama.cpp, stable-diffusion.cpp, and Unsloth Studio.
Hardware
Linux host has:
- RTX 2080 SUPER: CUDA device 0, compute capability
sm_75, ~8 GB VRAM - GTX 980: CUDA device 1, compute capability
sm_52, ~4 GB VRAM - Proprietary NVIDIA driver 580
- CUDA 12.9.1
llama.cpp CUDA builds target both GPUs with:
CMAKE_CUDA_ARCHITECTURES=52;75
The host keeps CUDA 13.4. The llama.cpp build uses CUDA 12.9 inside an Ubuntu 24.04 container because CUDA 12.9 is incompatible with host Ubuntu 26 headers.
Preset files
The repository has two preset layers:
preset.iniis the multi-model llama.cpp router preset used by the optional standalone service inscripts/llama-server.sh.llama-swap-presets/*.inicontains one model configuration per file. These are the presets launched by the llama-swap gateway defined inllama-swap.yaml.
The root server.sh controls the gateway and selects models by their
llama-swap model ID; it does not take a preset filename. The standalone
service can select a preset explicitly, for example:
./scripts/llama-server.sh start
preset.ini
preset.ini defines a llama.cpp router with these active model sections:
Qwen3.6-35BOrnith-1.5-35BGemma-4-12BGemma-4-26BQwopus3.6-35B
The [*] section supplies defaults, while a named model section overrides
those defaults.
The main defaults in preset.ini include:
c = 131072
image-min-tokens = 1024
device = CUDA0,CUDA1
split-mode = layer
main-gpu = 0
device-draft = CUDA1
mmproj-device = CUDA1
models-max = 1
CUDA0 is the RTX 2080 SUPER and CUDA1 is the GTX 980. Layer splitting
keeps the model distributed across both GPUs, while draft-model and multimodal
projector work is assigned to CUDA1. The GTX 980 is connected over PCIe x1.
llama-swap-presets/
Each single-model preset has a [*] section for shared llama.cpp settings
and one named section containing the Hugging Face model, sampling parameters,
context size, and speculative-decoding configuration. The active gateway
presets are:
| File | Model section | Context | Gateway status |
|---|---|---|---|
qwen3.6-35b-mtp.ini |
Qwen3.6-35B |
131,072 | Active |
qwen3.8-27b-pi.ini |
Qwen3.8-27B-Pi |
65,536 | Active |
ornith-1.5-35b-a3b-mtp.ini |
Ornith-1.5-35B |
131,072 | Active |
qwopus3.6-35b-mtp.ini |
Qwopus3.6-35B |
131,072 | Active |
gemma-4-12b-it-qat.ini |
Gemma-4-12B |
65,536 | Active |
gemma-4-26b-a4b-it-qat-mtp.ini |
Gemma-4-26B |
98,304 | Active |
qwen3.8-flash-next.ini |
Qwen3.8-Flash-Next |
32,768 | Not currently exposed; YAML entry is commented out |
laya-bf16.ini |
Laya |
Default | Not currently listed in llama-swap.yaml |
The hf setting identifies the Hugging Face repository and model variant,
for example unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL. The model files are
resolved through the container's /models Hugging Face cache. Set HF_TOKEN
in .env when a model requires authentication.
Common preset settings include:
c: context length; model-specific values override the global default.device,split-mode, andmain-gpu: GPU placement and layer splitting.device-draftandspec-type: speculative decoding device and method.spec-draft-*: speculative draft limits and acceptance thresholds.load-modeandlazy-mode: model loading and memory-mapping behavior.ctx-checkpointsandcheckpoint-min-step: long-context checkpointing for the MTP presets that enable it.reasoning,reasoning-format, andchat-template-kwargs: reasoning and chat-template behavior.ctkandctv: key/value cache quantization for selected low-memory presets.
Changes to a preset affect the next process start. Gateway model IDs, preset
paths, context values, and capabilities must remain consistent between the
selected file and its entry in llama-swap.yaml.
Configuration
Scripts load .env from the repository root. Copy the template before using
any build or service script:
cp .env.example .env
Keep .env local; it may contain access tokens, API keys, and machine-specific
paths. Commit .env.example, not .env. Values such as
PRESETS_DIR="${MODELS_ROOT}/presets" are shell expansions, so define the
base variables before the derived variables.
.env.example settings
Paths
| Variable | Purpose |
|---|---|
MODELS_ROOT |
Base directory containing the source checkouts, presets, model cache, and Unsloth work directory. |
PRESETS_DIR |
Presets repository directory mounted at /presets; normally the checkout containing this README. |
LLAMA_DIR |
Existing llama.cpp Git checkout used for builds and mounted at /src. Set this before using scripts/build-from-source.sh llama.cpp. |
STABLE_DIFFUSION_DIR |
Existing stable-diffusion.cpp Git checkout used for image-generation builds and mounted at /stable-diffusion or /src. Set this before using scripts/build-from-source.sh stable-diffusion.cpp. |
MODELS_DIR |
Host directory mounted at /models, including the Hugging Face model cache and diffusion assets. |
LLAMA_DIR and STABLE_DIFFUSION_DIR must point to existing Git checkouts;
build-from-source.sh updates them with Git but does not clone them. The
llama-swap gateway requires both completed builds because it launches the
custom llama-server and sd-server binaries.
CUDA and Docker images
| Variable | Purpose |
|---|---|
CUDA_IMAGE |
Base CUDA development image used when building the local builder image. |
BUILDER_IMAGE |
Local Docker image tag produced by scripts/create-build-image.sh and used by the llama.cpp and stable-diffusion server scripts. |
CUDA_BUILD |
Build-directory suffix, such as 12 for build-12; it must match the binaries built in both source checkouts. |
CUDA_ARCHITECTURES |
Semicolon-separated CMake CUDA targets; the template uses 52;75 for the GTX 980 and RTX 2080 SUPER. |
CUDA_FLAGS |
Additional flags passed to CMake during CUDA builds. |
Container names
| Variable | Purpose |
|---|---|
BUILD_CONTAINER_NAME |
Name of the temporary Compose build container. |
SERVER_CONTAINER_NAME |
Name used by scripts/llama-server.sh for the standalone llama.cpp server. |
BENCH_CONTAINER_PREFIX |
Prefix reserved for benchmark containers. |
LLAMA_SWAP_CONTAINER_NAME |
Name used by the root server.sh llama-swap gateway. |
Standalone server and RPC settings
| Variable | Purpose |
|---|---|
SERVER_HOST |
Address used by the standalone llama.cpp and stable-diffusion servers inside their containers. |
SERVER_PORT |
Host port and server listen port used by the standalone server scripts. |
SERVER_INTERNAL_PORT |
Container-side published port. The current server scripts pass SERVER_PORT to the process as well, so keep these two values equal unless the scripts are changed together. |
RPC_HOSTS |
Host value passed to ggml-rpc-server, normally HOST when using host networking. |
RPC_PORT |
RPC server listening port. |
Authentication and Hugging Face access
| Variable | Purpose |
|---|---|
HF_TOKEN |
Optional Hugging Face access token passed to model-serving and Unsloth containers. Set it for gated or private repositories. |
LLAMA_API_KEY |
API key passed to the standalone llama.cpp server and used as the default source for LLAMA_SWAP_API_KEY. |
llama-swap gateway
| Variable | Purpose |
|---|---|
LLAMA_SWAP_VERSION |
Tag selected for the llama-swap image, such as unified-cuda. |
LLAMA_SWAP_IMAGE |
Full gateway image reference; normally derived from LLAMA_SWAP_VERSION. |
LLAMA_SWAP_PORT |
Host port published to the gateway's container port 8080. |
LLAMA_SWAP_API_KEY |
Gateway API key. The template defaults it to LLAMA_API_KEY; it must be non-empty when starting the gateway. |
Unsloth Studio
| Variable | Purpose |
|---|---|
UNSLOTH_IMAGE |
Unsloth Studio/JupyterLab image. |
UNSLOTH_CONTAINER_NAME |
Unsloth container name. |
UNSLOTH_WORK_DIR |
Host directory mounted as the Unsloth workspace. |
HF_CACHE_DIR |
Host Hugging Face cache mounted into the Unsloth container. |
UNSLOTH_BIND_ADDRESS |
Host bind address for Studio and JupyterLab; use 127.0.0.1 with SSH forwarding. |
UNSLOTH_STUDIO_HOST_PORT |
Host port mapped to Studio's container port 8000. |
UNSLOTH_JUPYTER_HOST_PORT |
Host port mapped to JupyterLab's container port 8888. |
UNSLOTH_STUDIO_PASSWORD |
Unsloth Studio password. |
JUPYTER_PASSWORD |
JupyterLab password. |
UNSLOTH_STUDIO_BOOTSTRAP_TIMEOUT |
Maximum Studio bootstrap time in seconds. |
Service lifecycle
| Variable | Purpose |
|---|---|
STOP_TIMEOUT |
Graceful Docker stop timeout in seconds for the service scripts. |
Key files:
| File | Purpose |
|---|---|
.env |
Local machine configuration; ignored by Git |
.env.example |
Portable configuration template |
Docker setup
Host requirements:
- Docker
- NVIDIA Container Toolkit
- Proprietary NVIDIA driver 580
Before building, clone llama.cpp into LLAMA_DIR; for gateway image
generation, also clone stable-diffusion.cpp into STABLE_DIFFUSION_DIR.
scripts/build-from-source.sh updates existing Git checkouts; it does not clone them.
Build image once after changing docker/Dockerfile:
./scripts/create-build-image.sh
Files:
| File | Purpose |
|---|---|
docker/Dockerfile |
CUDA 12.9.1 build image with CMake, OpenSSL, GCC, and ccache |
docker/docker-compose.llamacpp.yml |
GPU-enabled llama.cpp build service |
docker/docker-compose.stabledifusion.yml |
GPU-enabled stable-diffusion.cpp build service |
scripts/build-from-source.sh |
Pull/update the selected source tree and run its containerized build |
Unsloth Studio
docker/docker-compose.unsloth.yml runs the official Unsloth image with Unsloth Studio
and JupyterLab. Configure the UNSLOTH_* values in .env, set
UNSLOTH_STUDIO_PASSWORD and JUPYTER_PASSWORD, then start it with:
./scripts/unsloth.sh start
The script defaults to start; use ./scripts/unsloth.sh stop,
./scripts/unsloth.sh restart, ./scripts/unsloth.sh status, or
./scripts/unsloth.sh logs for container lifecycle management.
By default, Studio is available at http://<server-host>:8000 and JupyterLab
at http://<server-host>:8888; change UNSLOTH_STUDIO_HOST_PORT or
UNSLOTH_JUPYTER_HOST_PORT in .env to use different ports. The service
persists the Hugging Face cache, Studio data, Triton kernels, and files under
UNSLOTH_WORK_DIR. The published
ports default to 0.0.0.0; set UNSLOTH_BIND_ADDRESS="127.0.0.1" when using
SSH port forwarding instead of exposing the services to the LAN.
The CUDA builder image includes libssl-dev, so llama.cpp can download
Hugging Face models over HTTPS.
Scripts
The llama-swap gateway is managed by the root server.sh script. Supporting
scripts are under scripts/: manage-models.py, build-from-source.sh,
check-hf-updates.py, create-build-image.sh, llama-server.sh,
rpc-server.sh, stable-server.sh, and unsloth-studio.sh.
Build (Linux)
./scripts/create-build-image.sh # First time or after docker/Dockerfile changes
./scripts/build-from-source.sh llama.cpp # Build standard llama.cpp
./scripts/build-from-source.sh stable-diffusion.cpp # Required for gateway image generation
Manage gateway models
For a llama.cpp text model, use scripts/manage-models.py add to create a
single-model preset and register it in llama-swap.yaml in one step. Pass a
Hugging Face repository reference with an optional model variant:
./scripts/manage-models.py add "unsloth/GLM-5.3-GGUF:UD-Q4_K_XL"
This derives the llama-swap model ID GLM-5.3, creates
llama-swap-presets/glm-5.3-ud-q4-k-xl.ini, adds the model to the exclusive
single-model group, and appends its gateway configuration and log path to
llama-swap.yaml. The generated preset uses a 65,536-token context, text
input/output, tools enabled, dual-GPU layer splitting, and no speculative
decoding by default.
Adjust the defaults when needed:
./scripts/manage-models.py add \\
"unsloth/GLM-5.3-GGUF:UD-Q4_K_XL" \\
--model-id "GLM-5.3" \\
--context 131072 \\
--no-tools
For a Stable Diffusion model, pass --type stable-diffusion. This creates an
SDAPI gateway entry instead of an INI preset. The diffusion-model and LLM paths
must be valid inside the gateway container, usually under /models:
./scripts/manage-models.py add \\
--type stable-diffusion \\
"unsloth/Z-Image-Turbo-GGUF:Q4_K_M" \\
--model-id "Z-Image-Turbo" \\
--diffusion-model /models/hub/models--unsloth--Z-Image-Turbo-GGUF/snapshots/6c80814333b7b6a70a2e5b469a7c6437ce65de0f/z-image-turbo-Q4_K_M.gguf \\
--vae /models/ae.safetensors \\
--llm /models/hub/models--unsloth--Qwen3-4B-Instruct-2507-GGUF/snapshots/a06e946bb6b655725eafa393f4a9745d460374c9/Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf \\
--backend "diffusion=cuda0&cuda1" \\
--max-vram "cuda0=7,cuda1=3.5"
--vae defaults to /models/ae.safetensors; the default backend is
diffusion=cuda0&cuda1, and the default VRAM limit is cuda0=7,cuda1=3.5.
Stable Diffusion entries use ${sd_server}, the /sdapi/v1/samplers health
check, and image input/output capabilities. The Hugging Face reference is used
to derive the model ID; the script does not download the model or infer the
VAE and LLM paths.
Remove a model setup by its exact llama-swap model ID:
./scripts/manage-models.py remove "GLM-5.3"
Removal deletes the model entry from llama-swap.yaml, removes it from the
exclusive single-model group, and deletes the managed preset file. Existing
model cache files and llama-swap logs are intentionally left untouched. The
script can be run from any working directory, refuses duplicate additions, and
refuses to remove an unknown model ID. Review generated presets and capabilities
before starting a new model; model-specific sampling, reasoning, multimodal,
and speculative-decoding settings cannot be inferred from the Hugging Face
repository reference alone.
Hugging Face cache updates
scripts/check-hf-updates.py checks cached Hugging Face model repositories for
new revisions on main. It runs in check-only mode by default and reports the
cached files affected by an update without downloading anything:
./scripts/check-hf-updates.py
Use --update to update only files that are already cached locally and still
exist in the remote revision:
./scripts/check-hf-updates.py --update
Newly added repository files are not downloaded. Deleted upstream files are
reported and omitted. If no cached files remain in a repository, the script
still advances the cached main revision without downloading new model files.
The command exits nonzero only when one or more repositories fail to process.
The script automatically creates a virtual environment and installs
huggingface_hub when needed. Set HF_VENV_DIR to choose its location;
otherwise it uses $XDG_CACHE_HOME/hf-cache-updates-venv or
~/.cache/hf-cache-updates-venv.
The Hugging Face cache is selected in this order:
HF_HUB_CACHE${HF_HOME}/hub~/.cache/huggingface/hub
Hugging Face authentication is read from the normal environment and local
configuration used by huggingface_hub.
llama-swap gateway
server.sh exposes the configured models from llama-swap.yaml. The
following six text models and the image model are members of the exclusive
single-model group:
Qwen3.6-35BQwen3.8-27B-PiOrnith-1.5-35BGemma-4-12BGemma-4-26BQwopus3.6-35BZ-Image-Turbo
Z-Image-Turbo uses the locally built sd-server from
STABLE_DIFFUSION_DIR, mounted read-only into the gateway container. It uses
the same model paths and multi-GPU settings as scripts/stable-server.sh, and is
placed in the exclusive single-model group so image generation unloads any active model
in that group before using the GPUs. Build stable-diffusion.cpp first with
./scripts/build-from-source.sh stable-diffusion.cpp.
On the target server, from the presets directory configured in .env:
./server.sh start
./server.sh status
./server.sh restart
./server.sh stop
Quick Start
cd <presets-directory>
./scripts/create-build-image.sh
./scripts/build-from-source.sh llama.cpp
./scripts/build-from-source.sh stable-diffusion.cpp
./server.sh start
./server.sh status
For image generation, select Z-Image-Turbo as the model and use the
OpenAI-compatible images endpoint:
{
"model": "Z-Image-Turbo",
"prompt": "a watercolor painting of a mountain cabin",
"size": "1024x1024"
}
The Stable Diffusion WebUI-compatible endpoints are also available through the
same gateway, including /sdapi/v1/txt2img and /sdapi/v1/img2img.