34 points | 2d ago | Discuss on Hacker News | Back to Radar
I think Google had one called riffusion (the first version was designed for specs)
KeenLore is my locally hosted full-cast emotive audiobook generator I'm developing for my novel:
https://www.youtube.com/watch?v=WAeHgE94rVo
Another HN user suggested embedding sound effects. The lead author of MOSS-TTS emailed me that the next version will have a SOTA voice designer. Poke around the voice design space, you'll find text-to-effects here and there. I certainly have my eyes on this space.
The DEMON realtime engine is open source. The hosted service has a number of integrations with popular audio tools. (I'm on the team)
I was working with very simple elements, but I was surprised by some of the outputs.
I would have also suggested Waves Illugen, but it turns out that is text-only, you can't give it an audio reference.
This generates audio embeddings - much like CLIP does for visual inputs.
Particularly when the application is a corner case or requires aesthetic judgment.
Of course in the case of sound effects shopping is a well established practice. Sound libraries and foley artists are widely available.
Finally, sound is much harder than images because psychoacoustics are more complex than the mechanical descriptions while human visual experience is a massively filtered set of stimuli.
We chunk the visible into symbols/archetypes like chairs and trees to a much much greater degree than we chunk sounds into symbols.
What is the sound of a chair? Of a tree?
LLM music is heavily dependent on genre. There is no genre for SFX. Good luck.
Comments are loaded live from Hacker News and are not stored by Mid or Real.
narrationbox 44h ago on HN
What's your exact use case?
chr15m 43h ago on HN