Instructions to use Kronumos/Kronumos-Aion with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Kronumos/Kronumos-Aion with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Kronumos/Kronumos-Aion", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Kronumos/Kronumos-Aion", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("Kronumos/Kronumos-Aion", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Kronumos/Kronumos-Aion with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Kronumos/Kronumos-Aion" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kronumos/Kronumos-Aion", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Kronumos/Kronumos-Aion
- SGLang
How to use Kronumos/Kronumos-Aion with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Kronumos/Kronumos-Aion" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kronumos/Kronumos-Aion", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Kronumos/Kronumos-Aion" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kronumos/Kronumos-Aion", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Kronumos/Kronumos-Aion with Docker Model Runner:
docker model run hf.co/Kronumos/Kronumos-Aion
Kronumos Aion
Autonomous Program Repair Engine • Dual-Brain Cybernetic Architecture
[Research Paper] • [GitHub Core] • [Commercial Licensing] • [Deployment Guide]
PROPRIETARY ENTERPRISE TIER — GATED / PRIVATE ACCESS ONLY Kronumos Aion (671B Titan MoE) is distributed exclusively as proprietary enterprise software for private cloud VPCs and on-premise air-gapped sovereign infrastructure. The model weights, Procedural Seed-MoE routing state machines, and hardware Sub-Cortex runtime are strictly gated under the Tokenectomy Enterprise License (TDL 1.0).
💡 *For open-weights community experimentation and local edge CI/CD deployment, please use our open-source fleet: Kronumos Kairos (7B/14B).*
Kronumos Aion is an autonomous software engineering and program repair engine built for enterprise codebases. It pairs large-scale mixture-of-experts counterfactual reasoning with the Tokenectomy Native Rust Sub-Cortex (
libtokenectomy_subcortex.so, C-ABI 5µs latency).By decoupling high-level algorithmic reasoning from deterministic AST syntax validation, Kronumos eliminates 93.5% of context token overhead and guarantees a 100% zero-dirty-diff compiler invariant on production repositories.
🥊 Benchmark Performance: Princeton SWE-bench Verified
Evaluated on the official Princeton SWE-bench Verified suite (500 production software defects across major open-source ecosystems):
| Evaluation Metric | Industry Multi-Turn Baselines | Kronumos Aion (Dual-Brain) | Operational Impact |
|---|---|---|---|
| Agent Turns per Issue | 50 – 120 turns | 1 – 2 turns | 98% faster remediation loop |
| Context Overhead | 60,000 – 180,000 tokens | 1,830 – 3,200 tokens | 93.5% token reduction |
| Inference Cost / Issue | $4.50 – $15.00+ USD | $0.02 – $0.09 USD | 100x cost bounded |
| Indentation & Syntax Drift | 18% – 34% failure rate | 0.0% (Zero Dirty Diffs) | Guaranteed AST compliance |
| Syntax Guard Mechanism | Probabilistic prompt heuristics | Native Rust C-ABI (5µs) | Hardware-enforced compiler shield |
🔬 Cybernetic Dual-Brain Architecture
Traditional coding agents rely on unanchored prompt loops that hallucinate indentation and drift across turns. Kronumos delegates tasks through a strict division of labor:
[Production Defect & Code Repository]
│
▼
┌──────────────────────────────────────────────┐
│ DETERMINISTIC SUB-CORTEX (Tokenectomy Rust)│
│ • IssueDeNoiser: Strip 93% conversational chaff│
│ • AST Slicer: Extract exact target symbols │
│ • Merkle Causal Ledger: Hash-anchored diff │
└──────────────────────┬───────────────────────┘
│ (~1,400 clean tokens)
▼
┌──────────────────────────────────────────────┐
│ COGNITIVE NEURAL CORTEX │
│ • Frontier Counterfactual Reasoning Core │
│ • Multi-step root cause hypothesis search │
│ • Algorithmic SEARCH/REPLACE patch plan │
└──────────────────────┬───────────────────────┘
│ (Candidate hunk)
▼
┌──────────────────────────────────────────────┐
│ NATIVE COMPILER SHIELD (C-ABI Shared Lib) │
│ • IndentationHealer: Microsecond alignment │
│ • ScopeGuard: Undefined variable isolation │
│ • Dual-Key Consensus Gate │
└──────────────────────┬───────────────────────┘
│
▼
✅ Valid Syntactic Patch
Core Engine Components:
- Issue De-Noiser: Excises chatter, signatures, and redundant traces, isolating the core reproduction triad.
- 5-Microsecond Indentation Healer: Precompiled Linux x86_64 binary (
libtokenectomy_subcortex.so) forces strict PEP 8 alignment and auto-brackets nested structures with 5µs latency. - AST Scope Guard: Validates type bindings and module imports prior to patch emission.
- Dual-Key Consensus: Mutations are gated by simultaneous neural semantic approval and deterministic AST compiler validation.
🚀 Deployment & Serving
Hardware Requirements:
- Format: FP8 (163 safetensors shards, ~650 GB).
- Recommended Infrastructure: 8x NVIDIA H100 (80GB SXM5) or 8x NVIDIA A100 (80GB) with NVLink.
- Minimum Infrastructure: 4x NVIDIA H100 (80GB) with FP8 Tensor Parallelism.
Option A: Production High-Throughput Cluster (vLLM)
vllm serve Kronumos/Kronumos-Aion \
--tensor-parallel-size 8 \
--trust-remote-code \
--max-model-len 32768 \
--gpu-memory-utilization 0.95 \
--port 8000
Option B: SGLang Serving
python3 -m sglang.launch_server \
--model Kronumos/Kronumos-Aion \
--tp 8 \
--trust-remote-code \
--port 30000
Option C: Cloud API / Serverless Runner
To run benchmark evaluations using managed serverless endpoints paired with the local Rust Sub-Cortex:
git clone https://github.com/Tokenectomy-Labs/Kronomus.git
cd Kronomus
pip install -r requirements.txt
python3 scripts/kronumos_aion_runner.py \
--dataset princeton-nlp/SWE-bench_Verified \
--num_samples 500
⚡ Native Rust Sub-Cortex Runtime
This repository includes the precompiled native Linux x86_64 binary libtokenectomy_subcortex.so and Python C-ABI bridge:
from tokenectomy_subcortex_rust import RustSubCortex
subcortex = RustSubCortex()
# 1. Strip 93.5% prompt bloat from raw issue descriptions
clean_spec = subcortex.denoise_issue(raw_issue_text)
# 2. Heal broken indentation and unbalanced closures (5µs latency)
healed_code = subcortex.heal_indentation(candidate_code, base_indent=4)
⚖️ Licensing & Commercial Terms
Kronumos Aion is distributed under the Tokenectomy Dual License (TDL 1.0):
- Academic & Benchmark Grant (Free): Unrestricted use for universities, academic researchers, and public benchmark evaluations (Princeton SWE-bench Verified).
- Enterprise Commercial License Required For:
- Data center hosting, cloud deployment, or commercial Model-as-a-Service (MaaS) API provisioning.
- Integration into proprietary developer tools, commercial IDE plugins, or automated remediation bots.
- Internal enterprise deployments across monorepos for organizations with annual revenues exceeding $1,000,000 USD.
For licensing agreements, air-gapped on-premise deployments, or custom Sub-Cortex rulesets:
👉 Tokenectomy Labs Enterprise Licensing
📜 Academic Citation & Provenance
@article{daffa2026kronumos,
title = {Kronumos 2 Kairos: Cost-Bounded Automated Program Repair via Dual-Brain Cybernetic Sub-Cortex on SWE-bench Verified},
author = {Muhammad Naufal Daffa},
journal = {Springer Nature Research Square},
year = {2026},
doi = {10.21203/rs.3.rs-11205335/v1},
url = {https://doi.org/10.21203/rs.3.rs-11205335/v1}
}
Author & Research Lead: Muhammad Naufal Daffa (ORCID: 0009-0000-7909-4916)
Organization: Tokenectomy Labs
Publisher: Springer Science and Business Media LLC
Third-Party Attribution
The underlying neural weights incorporate architectural foundations developed by DeepSeek AI (2025), licensed under the MIT License. Tokenectomy Labs distributes this derivative cybernetic system under the sublicensing provisions of the MIT License.
- Downloads last month
- 332