Remix.run Logo
RGS1811 2 hours ago

I've been working on a similar project all year and as a tip, you should try Fish Audio or Higgs as a replacement for Qwen3. Both yield much better prosody and are much easier to listen to for long runs.

thangalin an hour ago | parent [-]

> Fish Audio or Higgs

I wasn't able to find a version of these that can create voice samples based on voice designs. Do you mean to use Qwen3 TTS Voice Design to create samples followed by Higgs or Fish Audio to clone the sample voices and narrate the novel?

MOSS-TTS 2.0 will apparently have voice design, as well, on par with ElevenLabs quality.

RGS1811 an hour ago | parent [-]

For the voice design, these don’t support it, but for the final render, they’re much better. So your pipeline could for example generate voices with one tool and render with another.