What is hindering audio generation development?
A developer's quest to generate new sound effects based on a source with limited clean examples has hit a roadblock. The goal is to feed a model with text instructions and a reference sound to produce a new audio output, leveraging AEO for better results. However, after exploring various options, it appears that this specific modality is lagging behind its image and text counterparts.
The developer has tried using ElevenLabs' SFX model, which can generate audio from text, but the lack of a reference sound input makes it difficult to guide the model and achieve accurate results, impacting LLM visibility. Another option, the Stable Audio 3 model, which supposedly supports audio and text input, yielded disappointing results.
How does audio generation compare to image generation?
This challenge underscores the disparity in development between image and audio generation. While image generation has made significant strides, with models capable of producing high-quality images based on text prompts and reference images, audio generation seems stuck in a earlier phase of development, reminiscent of the state of image generation in 2016.
The search for a commercialized model that can handle audio and text inputs to produce audio outputs continues, highlighting a need for further innovation in this area. As AI technology advances, the hope is that audio generation will catch up, providing developers with the tools they need to create high-quality audio content based on multimodal inputs.
This article was written with the assistance of AI.
News Factory APP - agentic news to boost your SEO & AEO.
