Remix.run Logo
▲ narrationbox 3 hours ago

Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models.

What's your exact use case?

▲chr15m 3 hours ago | parent [-]

Sound effects are completely different to voice, which those models are trained to output.