Google's New AI Turns Text into Music

January 29, 2023

Google researchers have made an AI that can generate minutes-long musical pieces from text prompts, and can even transform a whistled or hummed melody into other instruments, similar to how systems like DALL-E generate images from written prompts (via TechCrunch). The model is called MusicLM, and while you can't play around with it for yourself, the company has uploaded a bunch of samples that it produced using the model.

The examples are impressive. There are 30-second snippets of what sound like actual songs created from paragraph-long descriptions that prescribe a genre, vibe, and even specific instruments, as well as five-minute-long pieces generated from one or two words like "melodic techno." Perhaps my favorite is a demo of "story mode," where the model is basically given a script to morph between prompts.

It may not be for everyone, but I could totally see this being composed by a human (I also listened to it on loop dozens of times while writing this article). Also featured on the demo site are examples of what the model produces when asked to generate 10-second clips of instruments like the cello or maracas (the later example is one where the system does a relatively poor job), eight-second clips of a certain genre, music that would fit a prison escape, and even what a beginner piano player would sound like versus an advanced one. It also includes interpretations of phrases like "futuristic club" and "accordion death metal."

MusicLM can even simulate human vocals, and while it seems to get the tone and overall sound of voices right, there's a quality to them that's definitely off. The best way I can describe it is that they sound grainy or staticky. That quality isn't as clear in the example above, but I think this one illustrates it pretty well.

That, by the way, is the result of asking it to make music that would play at a gym. You may also have noticed that the lyrics are nonsense, but in a way that you may not necessarily catch if you're not paying attention, kind of like if you were listening to someone singing in Simlish or that one song that's meant to sound like English but isn't.

I won't pretend to know how Google achieved these results, but it's released a research paper explaining it in detail if you're the type of person who would understand.

AI-generated music has a long history dating back decades; there are systems that have been credited with composing pop songs, copying Bach better than a human could in the 90s, and accompanying live performances. One recent version uses AI image generation engine StableDiffusion to turn text prompts into spectrograms that are then turned into music. The paper says that MusicLM can outperform other systems in terms of its "quality and adherence to the caption," as well as the fact that it can take in audio and copy the melody.

That last part is perhaps one of the coolest demos the researchers put out. The site lets you play the input audio, where someone hums or whistles a tune, then lets you hear how the model reproduces it as an electronic synth lead, string quartet, guitar solo, etc. From the examples I listened to, it manages the task very well.

Like with other forays into this type of AI, Google is being significantly more cautious with MusicLM than some of its peers may be with similar tech. "We have no plans to release models at this point," concludes the paper, citing risks of "potential misappropriation of creative content" (read: plagiarism) and potential cultural appropriation or misrepresentation.

It's always possible the tech could show up in one of Google's fun musical experiments at some point, but for now, the only people who will be able to make use of the research are other people building musical AI systems. Google says it's publicly releasing a dataset with around 5,500 music-text pairs, which could help when training and evaluating other musical AIs.

Source: Re-posted and Summarized from Mitchell Clark at theverge.

[BACK]

TESTIMONIALS

What Our Clients

Are Saying About Us

We engaged The Computer Geeks in mid-2023 as they have a reputation for API integration within the T . . . [MORE].

James Hesse

November 4, 2023

We all have been VERY pleased with Adrian's vigilance in monitoring the website and his quick and su . . . [MORE].

Kenneth Bruscia PhD

June 6, 2023

FIVE STARS + It's true, this is the place to go for your web site needs. In my case, Justin fixed my . . . [MORE].

Paul Adler

March 28, 2023

We reached out to Rich and his team at Computer Geek in July 2021. We were in desperate need of help . . . [MORE].

Leigh Hutchens

March 2, 2023

Just to say thank you for all the hard work. I can't express enough how great it's been to send proj . . . [MORE].

Curtis Williams

February 28, 2023

I would certainly like to recommend that anyone pursing maintenance for a website to contact The Com . . . [MORE].

David Pappas

February 17, 2023

Google's New AI Turns Text into Music

Google's New AI Turns Text into Music

905-426-1784

To Make a Request For Further Information

5K

Happy Clients

12,800+

Cups Of Coffee

5K

Finished Projects

72+

Awards

TESTIMONIALS

What Our Clients