| ▲ | Ask HN: Are there AI models for generating sounds based on a text and reference? | |||||||||||||||||||
| 22 points by onemiketwelve a day ago | 12 comments | ||||||||||||||||||||
I've been having a hard time finding a solution. Is there really no commercialized model that I can feed in a reference sound and text instruction and get another sound out? Right now having a multimodal inputs to image or text output is a commodotized, solved problem. IE you can put a prompt for some image and use a reference image to guide the model on what you want. After all, a picture is worth a thousand words right? Ive also used a text and audio input in order to get a text description or classification out. I cannot for the life of me find a solution for Audio + text -> Audio My usecase is that I'm trying to generate new sound effects based on a source that doesn't have that many clean examples and thought this would be just another solved workflow but I'm not finding much. Tried elevenlabs SFX but it's only text->audio and without the reference, it's hard to guide with just text and recreate a sound with any accuracy. I tried the Stable Audio 3 model which is supposed to do what I want but the results were terrible. This modality seems to be stuck where image generation was in 2016. Is there something else or is this just not a common need? | ||||||||||||||||||||
| ▲ | soundworlds 3 minutes ago | parent | next [-] | |||||||||||||||||||
You should look into using AI to generate code that synthesizes sounds. Not exactly what you asked for, but I do think it is an approach worth considering: https://m.youtube.com/watch?v=1-i45X5aj94 | ||||||||||||||||||||
| ▲ | moonu 2 hours ago | parent | prev | next [-] | |||||||||||||||||||
Even though they're technically trained for music, it might be worth testing Suno/Lyria to see if they're able to do this. You might have to isolate it afterwards, but seems like it could be viable | ||||||||||||||||||||
| ▲ | thangalin an hour ago | parent | prev | next [-] | |||||||||||||||||||
https://github.com/OpenMOSS/MOSS-TTS KeenLore is my locally hosted full-cast emotive audiobook generator I'm developing for my novel: https://www.youtube.com/watch?v=WAeHgE94rVo Another HN user suggested embedding sound effects. The lead author of MOSS-TTS emailed me that the next version will have a SOTA voice designer. Poke around the voice design space, you'll find text-to-effects here and there. I certainly have my eyes on this space. | ||||||||||||||||||||
| ||||||||||||||||||||
| ▲ | jallmann an hour ago | parent | prev | next [-] | |||||||||||||||||||
Daydream Music - https://daydream.live The DEMON realtime engine is open source. The hosted service has a number of integrations with popular audio tools. (I'm on the team) | ||||||||||||||||||||
| ▲ | xg15 2 hours ago | parent | prev | next [-] | |||||||||||||||||||
Maybe a dumb idea, but how good are image models with manipulating spectrogram images? Then a workaround could be: convert input audio to spectrogram -> pass spectrogram + prompt to an image model -> convert modified spectrogram back to audio waveform. | ||||||||||||||||||||
| ||||||||||||||||||||
| ▲ | chr15m 2 hours ago | parent | prev | next [-] | |||||||||||||||||||
Apparently the AudioX and AudioLDM(2) models do this but I think you've found a genuine gap. | ||||||||||||||||||||
| ▲ | narrationbox 3 hours ago | parent | prev [-] | |||||||||||||||||||
Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models. What's your exact use case? | ||||||||||||||||||||
| ||||||||||||||||||||