Neural network for online music generation: how to create a track based on text
Briefly about the main thing (BLUF)
A track based on a text description is already a reality: you write the genre, mood and whether words are needed, the neural network returns the finished music in a couple of minutes. We show what can be obtained, how it works and how to connect generation with other tools.
A week ago I needed a jingle for a podcast. Previously, I would have gone to a freelance exchange, explained the terms of reference, waited three days, asked for corrections. And here I described the mood and genre, and a minute later I listened to four versions with vocals and arrangement. A neural network for generating music is crazy, but it works. IN module “Music” NeuralSpace you can try it right now.
What you can get
- Full songs - with vocals, verses, chorus
- Instrumentals and background music for videos
- Jingles for podcasts and YouTube
- Demo versions: you hum the melody with the lyrics and you get an arrangement. Well, not exactly “humming”, but describing the rhythm and mood

How does this work
Describe the track: genre, tempo, mood, reference. “Lo-fi, calm, 85 BPM, like in a playlist for studying” - and the model understands. Separately, you can give the lyrics of the song in Russian or English. The output is 2-4 options up to several minutes long.
Linking with other tools
The track can be immediately thrown as a soundtrack to video. Cover - draw in image generator. Description for sites - write to GPT chat. Everything is within one subscription, payment for real results.
Tips from practice
- Specify BPM and mood - without this the result is unpredictable
- 1-2 references maximum. “Like Radiohead + Russian rock of the 90s + lo-fi” - the model will get confused
- The lyrics are better in short lines - the model gets into the rhythm more accurately
- Generate 3-4 options and take the best one. The first option is almost never ideal
An honest minus: the neural network is still rather weak in complex arrangements and non-standard sizes. But for background music, jingles and demos - more than enough.
Register — the first tracks on welcome tokens.
Frequently asked questions (FAQ)
Question: How to get the best result from a neural network?
Answer: Use detailed prompts (descriptions) in English, set the style and details of the scene.
Question: Can these materials be used for commercial purposes?
Answer: Yes, the generated content is entirely yours.