Text Generation
PEFT
lora
trl
naming
brand-generation
controllable-generation
nomen-ai / FINAL_REPORT.md
krystv's picture
Add consolidated final report
39399b9 verified
|
Raw History Blame Contribute Delete
2.85 kB

Nomen-AI Final Report

Mission

Build an end-to-end, production-ready, T4-compatible pipeline for controllable cross-lingual morpho-phonetic brand/channel name synthesis.

Completed public assets

Architecture

  • Base model: Qwen/Qwen2.5-1.5B-Instruct
  • Fine-tuning: LoRA + TRL SFTTrainer
  • Preference tuning: TRL DPOTrainer
  • Controls: ROOT, THEME, SYL, LEN, CREATIVE
  • Anti-duplication: fuzzy similarity + character n-gram overlap
  • Creativity decoding: contrastive search for low creativity, min-p sampling for high creativity

Completed engineering

  • Modular Python package under nomen_ai/
  • Synthetic dataset builder
  • SFT training script
  • DPO training script
  • Smoke test script
  • Evaluation script
  • Artifact checker
  • Adapter card updater
  • CPU validation tests
  • Colab notebooks
  • Dockerfile and docker-compose GPU training path
  • Makefile command map
  • Gradio demo Space
  • Research/citation/license docs

Datasets

SFT

DPO

Execution status

The following were attempted but could not be completed from the agent environment because GPU/HF Jobs execution was repeatedly rejected:

  1. GPU sandbox creation
  2. HF Jobs T4 smoke test
  3. Retried HF Jobs smoke test
  4. Additional HF Jobs validation/training attempts

Current artifact state:

  • krystv/nomen-ai-sft-lora: repo exists, no adapter weights yet
  • krystv/nomen-ai-dpo-lora: repo exists, no adapter weights yet

How to complete training

Colab T4

git clone https://huggingface.co/krystv/nomen-ai
cd nomen-ai
pip install -q -r requirements.txt
huggingface-cli login
bash scripts/train_all_colab.sh

Docker GPU

git clone https://huggingface.co/krystv/nomen-ai
cd nomen-ai
export HF_TOKEN=hf_...
docker compose up --build

Expected trained outputs

  • krystv/nomen-ai-sft-lora/adapter_model.safetensors
  • krystv/nomen-ai-dpo-lora/adapter_model.safetensors

Live demo

The current demo is CPU-safe and uses the morpheme synthesizer fallback:

https://huggingface.co/spaces/krystv/nomen-ai-demo

It displays live artifact status and can be upgraded to model-backed inference after DPO weights are present.