MIDREAL

Ask HN: Are there AI models for generating sounds based on a text and reference?

34 points | 2d ago | Discuss on Hacker News | Back to Radar

Comments

narrationbox 44h ago on HN
Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models.

What's your exact use case?

chr15m 43h ago on HN
Sound effects are completely different to voice, which those models are trained to output.
xg15 44h ago on HN
Maybe a dumb idea, but how good are image models with manipulating spectrogram images? Then a workaround could be: convert input audio to spectrogram -> pass spectrogram + prompt to an image model -> convert modified spectrogram back to audio waveform.
Buttons840 44h ago on HN
I wouldn't call it a dumb idea, but there's soooo much subtlety to sound that wont be visible in any reasonably sized image.
0x20cowboy 43h ago on HN
narrationbox 43h ago on HN
Could be interesting, though spectrogram voice gen is a step back in terms of advancement. Vocoder models that take in Mel Specs suffer from a lot of issues like hissing due to Griffin Lim issues. Would be curious to see if diffusion models can work on neural codecs directly.

I think Google had one called riffusion (the first version was designed for specs)

chr15m 43h ago on HN
Apparently the AudioX and AudioLDM(2) models do this but I think you've found a genuine gap.
moonu 43h ago on HN
Even though they're technically trained for music, it might be worth testing Suno/Lyria to see if they're able to do this. You might have to isolate it afterwards, but seems like it could be viable
thangalin 43h ago on HN
https://github.com/OpenMOSS/MOSS-TTS

KeenLore is my locally hosted full-cast emotive audiobook generator I'm developing for my novel:

https://www.youtube.com/watch?v=WAeHgE94rVo

Another HN user suggested embedding sound effects. The lead author of MOSS-TTS emailed me that the next version will have a SOTA voice designer. Poke around the voice design space, you'll find text-to-effects here and there. I certainly have my eyes on this space.

ninininino 42h ago on HN
KeenLore is really cool.
jallmann 43h ago on HN
Daydream Music - https://daydream.live

The DEMON realtime engine is open source. The hosted service has a number of integrations with popular audio tools. (I'm on the team)

soundworlds 42h ago on HN
You should look into using AI to generate code that synthesizes sounds. Not exactly what you asked for, but I do think it is an approach worth considering: https://m.youtube.com/watch?v=1-i45X5aj94
SyneRyder 36h ago on HN
I've had some luck with that approach - giving a frontier LLM a reference sound, asking it to replicate the sound via code (basically asking it to create a physical model, I guess) and being able to steer the output by talking with the LLM.

I was working with very simple elements, but I was surprised by some of the outputs.

I would have also suggested Waves Illugen, but it turns out that is text-only, you can't give it an audio reference.

dmos62 24h ago on HN
Great video!
bobosha 32h ago on HN
Not exactly what you asked for, but are you aware of CLAP? https://github.com/LAION-AI/CLAP

This generates audio embeddings - much like CLIP does for visual inputs.

jajazheng 31h ago on HN
Hi! I’m unsure if this is what you were looking for, QWEN3 TTS has a “clone” feature. You can use the local version, give it a voice reference and text and it will read the text with your instructions with the sample voice you selected. I used this combo to make a customised “audiobook” for my mother. It worked quite well.
ruiqingcn 19h ago on HN
INDEX TTS2
brudgers 2h ago on HN
Sometimes doing the work is easier than shopping for a tool that does the work for us.

Particularly when the application is a corner case or requires aesthetic judgment.

Of course in the case of sound effects shopping is a well established practice. Sound libraries and foley artists are widely available.

Finally, sound is much harder than images because psychoacoustics are more complex than the mechanical descriptions while human visual experience is a massively filtered set of stimuli.

We chunk the visible into symbols/archetypes like chairs and trees to a much much greater degree than we chunk sounds into symbols.

What is the sound of a chair? Of a tree?

LLM music is heavily dependent on genre. There is no genre for SFX. Good luck.

Comments are loaded live from Hacker News and are not stored by Mid or Real.