Laya RA2 v33 — state-to-choice decision model

Published by jayzou3773. This is the exact fine-tuned Laya checkpoint used in our September 25, 2026 Red Alert 2 development matches against the adapted Supalosa bot. It scores context + currently offered choices; it does not generate chat, free-form commands, coordinates, or chain-of-thought.

The release includes full weights, tokenizer, encoder configuration, a portable Python inference interface, real decision examples, and the frozen JavaScript context/choice/game adapter. The weight file is approximately 842 MB (804 MiB).

Verified results: 11 wins / 0 losses / 1 unresolved on mp03t4; 2 wins / 0 losses / 10 unresolved on mp06t2; 13 wins / 0 losses / 11 unresolved across 24 development games. “12 wins out of 13” is not supported by these records. These are development results for the whole model + context + candidate + inference system, not held-out performance or a bare-checkpoint win rate. See per-game evidence and the limitations below.

The trained contexts and choice labels are English. A Chinese integration guide is provided for documentation: 中文接入说明.

What is included, and what you still need

Included here Supplied separately for a live game
Full model.safetensors and local tokenizer/config Linux and Node.js 22+
Python state/choice inference; CPU or CUDA ra2-game-env SDK and engine binary
Frozen RA2 normal-view context and staged choices Complete, legally obtained RA2 game resources
Supalosa compatibility adapter and match runner Pinned Supalosa source/dependencies (setup script downloads them)
Training provenance, results and policy-only examples A suitable CUDA PyTorch environment for fast matches

The tested engine was version 0.1.0, commit abad0cc238bd74b9e64a3d7e92c4631418d8dfae. The SDK directory must contain dist/index.js and its associated engine executable. Neither that engine nor RA2 assets are redistributed here. Model-only inference requires neither. This is not an OpenRA adapter or a plugin for the original Windows RA2 executable. A different engine needs its own observation/action translation.

Quickstart: score a real recorded decision

Python 3.12 was used for validation. Download the public repository without a token:

python -m venv .venv
source .venv/bin/activate
pip install huggingface_hub
python -c 'from huggingface_hub import snapshot_download; snapshot_download("jayzou3773/laya-ra2-v33", local_dir="laya-ra2-v33")'
cd laya-ra2-v33
pip install -r requirements.txt
USE_TF=0 python inference.py --device cpu --request examples/opening-request.json

The opening sample is a real normal-view request at tick 0. Its recorded decision was C: deploy the construction vehicle. CPU is useful for integration checks; CUDA is preferable for thousands of decisions. CUDA inference should use a suitable PyTorch build for the host (historically torch 2.7.0+cu128, bf16).

import json
from pathlib import Path
from inference import load_agent, decide_stage

# Load ONCE, keep resident throughout the match.
agent = load_agent(Path.cwd(), device="cpu")  # or device="cuda"
request = json.loads(Path("examples/opening-request.json").read_text())
result = decide_stage(agent, request)
print(result["answers"]["next_action"]["choice"])

The underlying API is laya.load(local_checkpoint, device=...).predict(state, questions, max_len=2048, head_max_len=512). Use the supplied decide_stage wrapper when reproducing this experiment: it enforces lossless context/option budgeting and the conditional target permutation described below. Standard Transformers AutoModelForCausalLM, chat templates, and generate() are not the loading or inference interface for this custom encoder-and-decision-head checkpoint.

How context is constructed

Context is built deterministically from the acting player's normal observation, public rules/map metadata, and that player's prior observations and receipts. No LLM writes the context. See context.mjs and intent-policy.mjs.

The input has four ordered sections:

  1. essential: current stage/tick, prior selections within this decision, game objective, own credits and power, harvester and army counts, production queue states/progress, currently visible enemies, operating advice, and the age of the last accepted combat command.
  2. details: compact own-unit groups, idle/order information and representative positions; explanations of search/cleanup behavior.
  3. memory: previously visible enemies with last-seen location and age, explicitly marked as unknown current survival/location. They are not live hidden-state reports and cannot become direct entity attack targets while invisible.
  4. logs: a short recent history of own observed count/queue changes and submitted commands/receipts. A receipt is not proof of arrival, damage, mining income, or successful execution.

The operating advice describes a recurring loop: establish power/refinery/miners, keep mining and production working, replace combat losses, scout and attack, place completed buildings, resume useful paused queues, and preserve the base. The code also reports net credit change (including spending/refunds, not gross mining income), miner Gather/Dock orders, production stalls and observed army count decline. These are input features and recommendations; the model still chooses the command. Historical enemy memory expires after 900 ticks.

Token budget: 2048 total tokens, including a 512-token question head. Each option is limited to 48 tokens. If needed, optional logs, then memory, then details are removed. Essential state and all offered options must fit intact or inference fails. See token_budget.py. There is no silent truncation of legal choices or essential commitments.

How choices are constructed and executed

The adapter consults the engine's action catalog, availability masks, own units, public unit/building rules, production state and normal-view targets to build concrete candidate actions. The set changes each tick. Examples include production, placement, queue controls, harvesting/support, scouting/movement, attacking, deployment, repairs, selling, diplomacy and resignation where supported. An option's letter has meaning only inside that request, not globally.

This is an engineered candidate policy, not exhaustive enumeration of every parameter combination. Grouping, finite parameter shortlists, public search waypoints and policy restrictions influence performance. For example, harvesters are excluded from distant military commands, experimentally ineffective construction-yard deployment bindings are excluded, and camera actions are merged with waiting. The source and omission records make these restrictions inspectable; passing an availability check does not guarantee pathfinding or a useful effect.

A decision is resolved in stages:

normal observation + player memory
    -> generate concrete candidate leaves
    -> select intent
    -> select function, if ambiguous
    -> select acting unit(s), if ambiguous
    -> select target/location/remaining parameters, if ambiguous
    -> confirm destructive actions when required
    -> validate and submit the selected concrete engine action
    -> advance simulation and update memory from observations/receipts

All stages use the same paused simulation tick. The next request's essential context contains previous selections and bound parameters. They are commitments, not observations that an action has executed. A single surviving leaf needs no extra call. Root A waits; later A cancels the pending action. Existing game orders continue when the simulation advances. Selling/resignation need explicit model confirmation. There are at most 8 stages per decision, up to 12 non-cancel branches on ordinary pages and 8 on target pages; larger sets are reached through additional page-selection calls.

questions contains task instructions and labels, while parts contains state. They enter the same model input during encoding; context is not missing merely because it is not duplicated into instructions:

request = {
    "parts": {
        "essential": ["Tick 0. Own construction vehicle available; no base yet."],
        "details": [], "memory": [], "logs": []
    },
    "questions": {
        "next_action": {
            "type": "choice",
            "instructions": "Choose the next action that advances your objective.",
            "criteria": {"A": "Wait", "B": "Deploy construction vehicle"}
        }
    },
    "dimensions": {"next_action": "intent"}
}
# This tiny example illustrates the schema only. Use the adapter and full
# recorded examples for the actual trained context and action bindings.

See staged.mjs, real multi-stage trace, and real target trace. Never submit the option letter as an engine action: resolve it through that request's candidate mapping.

Inference settings used for the reported games

  • Both sides advance/decide at 3-tick cadence (15 ticks/game second).
  • Laya selects at most one complete command per decision. The baseline/engine command cap is 16 per seat. Equal cadence does not mean equal command budgets.
  • All model calls pause simulation for both sides; inference wall time does not grant the opponent extra game time.
  • Deterministic probability argmax, no generative sampling; confidence is not calibrated as a probability of tactical success.
  • For a target stage with exactly Cancel + target B + target C, score once in original order and once with B/C swapped. Map probabilities back, normalize rounding, average and take argmax. Other stages use one forward call. This is binary-target-permutation-mean-v1, not a general ensemble over every decision.
  • Original inference used bf16 on CUDA. CPU fp32 inference may differ near ties; the CPU smoke test is not a new win-rate measurement.

Run against the baseline

cd ra2
node scripts/setup-dependencies.mjs
cd ..
export RA2_GAME_DIR=/absolute/path/to/RA2-assets
export RA2_SDK_DIR=/absolute/path/to/ra2-sdk-package
export LAYA_PYTHON=/absolute/path/to/your/venv/bin/python
# Defaults to CUDA GPU 0; choose an idle GPU with --gpu=N.
bash ra2/run-match.sh --map=mp03t4.map --seed=1004 --seat=red \
  --gpu=0 --ticks=18000 --out=runs/public-demo-1004-red

The dependency setup pins Laya to 970dc8c5f63d7b886a68409493f37d569424f933 and Supalosa to 165b77a71d0cf5ebd27c65b19d0486bcbae78d0f. It does not download RA2 assets or the separately distributed game engine. No parent-checkpoint download is needed: the published checkpoint contains all model weights. For a short CPU integration check, set LAYA_DEVICE=cpu and --ticks=90. Use a fresh --out directory for each run.

The match runner saves decisions, receipts, result, source hashes, observer-only frames and .rplx replay. A tick-cap termination has no winner; it is not a loss or a score-based win. The replay needs the compatible engine and is not an original RA2/Chrono Divide/OpenRA replay.

Training provenance

Architecture: Laya's ModernBERT-large encoder plus typed decision heads. Upstream base is convaiinnovations/laya-typed-decisions at 1a793eb568e6718f15941d08f85432581df534e3. This checkpoint continued from the RA2 v29 fine-tune (1aff9a70…), not directly from unmodified ModernBERT.

The RA2 fine-tuning objective was supervised cross-entropy over staged choice labels, using rule-based teacher demonstrations and corrections on learner-visited states (DAgger-style data aggregation). There was no runtime teacher takeover. The private TimS-ml/ra2-replay dataset was not used for this checkpoint.

The v33 data contained 9,462 deduplicated training decisions and 763 novel validation decisions, sourced from ten teacher/correction episodes. Validation used seed 105; overlapping encoded inputs were removed from validation. The lineage includes seeds 105, 110, 113, 114, 115. Development seeds 1000–1005 were not SFT source episodes, but the development maps/results were used in iteration.

Four continuation epochs, batch size 8, seed 7, AdamW with encoder LR 1e-5, head LR 1e-4, weight decay 0.01. The released weights are epoch 1, selected by minimum validation loss 0.1508458, accuracy 0.9501966 on the 763 validation choices. Choice accuracy is not game win rate; many validation labels are waiting. The later four-epoch history is in training-history.json. The original rl_agent_config.json includes inherited upstream training metadata; training.json is authoritative for this RA2 continuation.

Evaluation limitations and normal-view boundary

  • Six seeds, two maps, two fixed starts with swapped Laya seat: 24 development games, up to 18,000 ticks. No held-out generalization claim or causal attribution of gains to a single change is supported.
  • Both players: Russians, 10,000 initial credits, no starting army, normal view, short game on, superweapons/crates off. Offscreen camera masking was disabled for both; fog of war remained active.
  • Policy inputs exclude live hidden opponents, opponent assigned spawn/credits, privileged observer frames and replay hashes. Public map start candidates and aged own observations are allowed. The historical boundary is in-process, not an OS-level security boundary.
  • On mp03, the historical audit reviewed 60,860 decision records and 1,567 model actions. Five winning matches contained baseline resignation commands; six wins did not. Do not describe all wins as total annihilation. Disabling baseline resignation was not evaluated as a counterfactual.
  • The Supalosa bot runs through a compatibility adapter on another engine. Its native strength and every translated API behavior are not proven equivalent.
  • Equal observations and cadence do not imply equal computation/action budgets. The hand-designed context, candidate policy, and target averaging are part of the result. Audit/replay consistency does not prove absence of every possible bug, hidden information leak or evaluation bias.

Files, attribution and integrity

  • model.safetensors, encoder/, tokenizer/, rl_agent_config.json: original checkpoint.
  • inference.py: portable checked loading, token budget and staged choice scoring.
  • ra2/integration/: historical policy/adapter/runner snapshot. Only worker.py is replaced for portable release loading and optional CPU use; its original is retained as worker.original.py. The policy and target-ensemble code are unchanged.
  • examples/: extracted policy-side inputs, historical answers and selection traces.
  • provenance/, evaluation/: sanitized training/evaluation records; no private credentials, game assets or full training datasets.
  • SHA256SUMS: published-file checksums; release.json pins lineage.

Weight SHA-256: 2ed46ac94c366db2dc68751d0eabfe945ceaeaf430e9fea1ec9e6c60c0207925.

Apache-2.0 for the released weights/code; upstream attribution is in NOTICE, LICENSE and LICENSE-LAYA. Game assets and the separately acquired engine are not covered by this model repository's license. This is an independent research release, not an official RA2, OpenRA, Supalosa or Convai product.

Release validation

The original checkpoint hash matches the recorded matches. CPU checks loaded all weights, confirmed every parameter finite, reproduced recorded context packing, and compared the Python target-averaging wrapper to the frozen JavaScript code. See model checks. A real-engine 90-tick CPU smoke completed 30 decisions with two accepted Laya commands and no rejected Laya commands; see engine smoke. This short test establishes packaging/connectivity, not match performance.

To repeat the model-only checks, with Node.js 22+ installed:

USE_TF=0 python validate_release.py
Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jayzou3773/laya-ra2-v33

Finetuned
(9)
this model