Text-to-Speech
Safetensors
fish_qwen3_omni
instruction-following
multilingual

Fine tune possibility

#18
by Darn9027 - opened

Hi there, if I have a few hours of audio data available, should I use the fine-tuning feature instead of the voice cloning feature? The fine-tuning tutorial on your website still uses the fishaudio/openaudio-s1-mini model, do the same method apply to the s2-pro model as well?

I haven't been able to (locally) fine-tune the s2-pro model. The structure of the data you download for s1-mini and s2-pro are different. I haven't figured out how to fine-tune s2-pro. If anyone else has I would be happy to hear about it.

Sign up or log in to comment