VoTSpeech Model Weights

VoTSpeech is an instruction-following text-to-speech model that separates voice design from speech generation. Given synthesis text and a natural-language voice description, it designs a continuous voice representation and generates 48 kHz speech in the requested voice.

The model extends the continuous autoregressive dots.tts-soar backbone with an instruction-conditioned voice-design stage. This repository provides the released VoTSpeech checkpoint and the configuration and tokenizer assets needed to load it.

Code and usage

For inference code, installation, and usage instructions, see ywbn/VoTSpeech.

Weight files

  • model.safetensors: VoTSpeech model weights
  • vocoder.safetensors: 48 kHz vocoder weights
  • speaker_encoder.safetensors: speaker encoder weights
  • latent_stats.pt: voice-latent normalization statistics
  • config.json, llm_config.json: model configuration
  • tokenizer.json, tokenizer_config.json, chat_template.jinja: tokenizer assets
  • SHA256SUMS: release integrity checksums

License

The VoTSpeech model weights are licensed under CC BY-NC 4.0. The license permits non-commercial use, sharing, and adaptation with appropriate credit and an indication of any changes. The inference code is distributed separately in the GitHub repository.

VoTSpeech is intended for non-commercial research and evaluation. Publicly shared outputs should be identified as synthetic speech, and permission should be obtained before imitating an identifiable person. This repository does not distribute training data; rights in third-party source materials remain with their respective holders. Users should review generated audio and avoid deceptive, harmful, or privacy-invasive uses.

Downloads last month
31
Safetensors
Model size
2B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support