VoTSpeech Model Weights
VoTSpeech is an instruction-following text-to-speech model that separates voice design from speech generation. Given synthesis text and a natural-language voice description, it designs a continuous voice representation and generates 48 kHz speech in the requested voice.
The model extends the continuous autoregressive
dots.tts-soar backbone
with an instruction-conditioned voice-design stage. This repository provides
the released VoTSpeech checkpoint and the configuration and tokenizer assets
needed to load it.
Code and usage
For inference code, installation, and usage instructions, see
ywbn/VoTSpeech.
Weight files
model.safetensors: VoTSpeech model weightsvocoder.safetensors: 48 kHz vocoder weightsspeaker_encoder.safetensors: speaker encoder weightslatent_stats.pt: voice-latent normalization statisticsconfig.json,llm_config.json: model configurationtokenizer.json,tokenizer_config.json,chat_template.jinja: tokenizer assetsSHA256SUMS: release integrity checksums
License
The VoTSpeech model weights are licensed under CC BY-NC 4.0. The license permits non-commercial use, sharing, and adaptation with appropriate credit and an indication of any changes. The inference code is distributed separately in the GitHub repository.
VoTSpeech is intended for non-commercial research and evaluation. Publicly shared outputs should be identified as synthetic speech, and permission should be obtained before imitating an identifiable person. This repository does not distribute training data; rights in third-party source materials remain with their respective holders. Users should review generated audio and avoid deceptive, harmful, or privacy-invasive uses.
- Downloads last month
- 31