← All articles

Neural network for online music generation: how to create a track based on text

Neural network for online music generation: how to create a track based on text - NeuralSpace

Briefly about the main thing (BLUF)

A track based on a text description is already a reality: you write the genre, mood and whether words are needed, the neural network returns the finished music in a couple of minutes. We show what can be obtained, how it works and how to connect generation with other tools.

A week ago I needed a jingle for a podcast. Previously, I would have gone to a freelance exchange, explained the terms of reference, waited three days, asked for corrections. And here I described the mood and genre, and a minute later I listened to four versions with vocals and arrangement. A neural network for generating music is crazy, but it works. IN module “Music” NeuralSpace you can try it right now.

What you can get

  • Full songs - with vocals, verses, chorus
  • Instrumentals and background music for videos
  • Jingles for podcasts and YouTube
  • Demo versions: you hum the melody with the lyrics and you get an arrangement. Well, not exactly “humming”, but describing the rhythm and mood
The neural network generates a music track based on the text

How does this work

Describe the track: genre, tempo, mood, reference. “Lo-fi, calm, 85 BPM, like in a playlist for studying” - and the model understands. Separately, you can give the lyrics of the song in Russian or English. The output is 2-4 options up to several minutes long.

Linking with other tools

The track can be immediately thrown as a soundtrack to video. Cover - draw in image generator. Description for sites - write to GPT chat. Everything is within one subscription, payment for real results.

Tips from practice

  • Specify BPM and mood - without this the result is unpredictable
  • 1-2 references maximum. “Like Radiohead + Russian rock of the 90s + lo-fi” - the model will get confused
  • The lyrics are better in short lines - the model gets into the rhythm more accurately
  • Generate 3-4 options and take the best one. The first option is almost never ideal

An honest minus: the neural network is still rather weak in complex arrangements and non-standard sizes. But for background music, jingles and demos - more than enough.

Register — the first tracks on welcome tokens.

Frequently asked questions (FAQ)

Question: How to get the best result from a neural network?
Answer: Use detailed prompts (descriptions) in English, set the style and details of the scene.

Question: Can these materials be used for commercial purposes?
Answer: Yes, the generated content is entirely yours.