Skip to content
Tech News
← Back to articles

No Mic Needed: You Can Create Music and Speech With Adobe’s AI Audio Tools

read original more articles
Why This Matters

Adobe’s new AI audio tools empower creators with professional-grade capabilities to generate speech, music, and sound effects directly within their workflow, streamlining content creation while maintaining high quality. This development is significant as it offers more accessible, customizable, and copyright-safe audio options, catering to filmmakers, musicians, and digital creators. By integrating AI-driven audio into Adobe’s ecosystem, the company enhances its position as a leader in creative technology and simplifies complex audio production tasks for users.

Key Takeaways

Adobe announced Thursday that it’s adding to its vast collection of artificial intelligence tools, this time with a focus on audio. The new Firefly AI tools use AI to generate speech, music and sound effects, forming “a very stable foundation” of audio options for creators, said Jay LeBoeuf, Adobe’s head of AI audio.

Unlike Suno or other AI music generators that can create entire songs in seconds, Adobe is offering more targeted, professional-grade tools for filmmakers, musicians and creators. Think about AI converting creators’ scripts into audio to be overlaid on a TikTok video or creating custom soundtracks without worrying about copyright infringement, thanks to Adobe’s universal license.

“We’re not trying to be somebody’s wedding music here,” LeBoeuf said in an interview. The goal is to build AI “tools that are useful” and address pain points in the audio creation and editing process.

To use the new tools, you’ll need access to Firefly, Adobe’s AI hub, which may be included in your current Creative Cloud subscription, depending on your specific plan or your company’s AI permissions. AI audio generations will count as generative credits, so keep an eye on how quickly you use those up. You can nab a Firefly-only subscription starting at $10 per month.

AI audio that isn’t robotic

To create speech, upload a script you’ve written, and it will transform it into an audio file. You’ll be able to choose from several artificial voices in a variety of genders and ages, and you can translate the audio into over 20 languages.

If you have names or products that are hard for the AI to read, you can add pronunciation guidance. My last name, Chedraoui, for example, could be phonetically spelled out as “Shed-rao-wee,” instead of whatever hideous sound the AI produces when pronouncing four vowels in a row.

One of the biggest challenges with AI audio is making voices sound less robotic. Monotonous audio is boring to listen to, and it’s a clear sign of AI. Adobe built its AI audio model to recognize and apply various emotions to its outputs. When you use Adobe’s generate speech tool you can use “emotion tags” to direct the AI to apply different expressions.

To help the models understand emotions, Adobe collected more emotive training material, LeBoeuf said. “Our design team knew that we were going to control them with these adjectives and these verbs. So because it’s been part of the training since the get-go, we have this nice vertically integrated stack that allows for the highest quality expressiveness.”

Adobe’s AI policy says that it only uses licensed and publicly available content and data to train its AI models. The company says it never trains on customers’ work to improve its services.

... continue reading