| ▲ | xg15 3 hours ago | |
Maybe a dumb idea, but how good are image models with manipulating spectrogram images? Then a workaround could be: convert input audio to spectrogram -> pass spectrogram + prompt to an image model -> convert modified spectrogram back to audio waveform. | ||
| ▲ | 0x20cowboy 2 hours ago | parent | next [-] | |
https://huggingface.co/docs/transformers/model_doc/audio-spe... | ||
| ▲ | Buttons840 3 hours ago | parent | prev | next [-] | |
I wouldn't call it a dumb idea, but there's soooo much subtlety to sound that wont be visible in any reasonably sized image. | ||
| ▲ | narrationbox 2 hours ago | parent | prev [-] | |
Could be interesting, though spectrogram voice gen is a step back in terms of advancement. Vocoder models that take in Mel Specs suffer from a lot of issues like hissing due to Griffin Lim issues. Would be curious to see if diffusion models can work on neural codecs directly. I think Google had one called riffusion (the first version was designed for specs) | ||