How to describe a sound an AI synthesizer can actually make
A practical guide to prompting instrument sounds: name the family, describe timbre with material and motion words, say what the sound is for, and iterate with takes instead of rewrites.

Most first prompts to a generative instrument look like film reviews: epic, cinematic, emotional, huge. They are honest descriptions of a feeling, and they are almost useless to a model that has to produce a waveform. The result is the audio equivalent of beige - a sound that is a bit of everything and therefore nothing.
The fix is not a longer prompt. It is a prompt that talks about sound the way sound designers do. This guide walks through the four moves that reliably turn a vague wish into a playable instrument.
Name the instrument family first
The single most useful word in any prompt is the family: bass, lead, pad, pluck, keys, organ, texture, impact. It sets the register, the envelope, and the role in a track before anything else is decided.
Compare warm and dreamy with warm dreamy pad. The first could be a piano, a choir, or a wind chime. The second is a sustained, chordal, slow-attack sound - and the model can now spend its effort on making it warm and dreamy instead of guessing what “it” is.
If you are not sure which family you want, pick the role the sound will play: what carries the low end (bass), what sits on top (lead), what fills the space between (pad), what marks the rhythm (pluck, impact).
Describe timbre with material and motion words
Once the family is fixed, describe how it feels, using words that map to physics:
- Material: analog, digital, glassy, metallic, wooden, felt, tube, saturated
- Motion: filtered, wobbling, breathing, evolving, modulated, reversed, granular
- Weight: sub, deep, thin, bright, dark, airy, fat
Fat moog bass, glassy fm bell, breathing organic pad, filtered analog arp - each of these gives the model a texture to aim at and a way to move. Emotional words are fine as a third word (“haunting choir pad”), but they should never be the only word.
Say what the sound is for
Context resolves ambiguity. Bass alone could be a smooth sub for a ballad or a distorted growl for a drop. Sub bass for a powerful trap drop and cello-like bass that sounds sad and lonely are both bass prompts, and they should sound nothing alike.
Genre and function are shortcuts the model understands well: berlin techno lead, lo-fi pad with a cinematic quality, progressive house arp, neurofunk bass design. One of these phrases carries more information than a paragraph of adjectives.
Use the vocabulary the model already knows
Every generative model has a home vocabulary - the phrases that appear often enough in its training that it has a confident answer for them. In AI Synthesizer, that vocabulary is exposed directly: the suggestions that appear as you type, the idea chips under the prompt bar, and the dice button all draw from it.
Starting from a known phrase and changing one word is the fastest route to a specific sound. Warm analog lead becomes warm analog bass, then dark analog bass, then dark bass sequence. Each step moves one variable, so you always know what changed.
Iterate with takes, not rewrites
When a result has the right character but the wrong detail, do not rewrite the prompt. Ask for another take. AI Synthesizer generates four takes per prompt precisely because the same words produce a family of related sounds, and the second or third take is often the one.
Rewrite only when the character is wrong: the wrong family, the wrong weight, the wrong motion. Then change one word at a time, so a better result can be traced back to the word that caused it.
A prompt checklist
Before pressing Create, read the prompt back and check:
- Is there an instrument family (bass, lead, pad, pluck, keys)?
- Is there at least one material or motion word?
- Does it say what the sound is for, or which genre it lives in?
- Is it under eight words?
- Is it in English, using terms a sound designer would use?
A prompt that passes all five will rarely disappoint, and when it does, the fix will be obvious.
Main claim: A good instrument prompt names the family, describes the timbre in material and motion words, and states the use - then relies on takes for variation and on single-word edits for direction.
Try AI Synthesizer
Describe a sound and play it seconds later - free in your browser.
Frequently asked questions
How long should a prompt for an instrument sound be?
Three to eight words is the sweet spot. One instrument family, one or two timbre words, and optionally a use. Longer prompts tend to add contradictions rather than detail, and the model has to guess which part matters.
Why does 'epic cinematic sound' give a muddy result?
Because it describes a feeling, not a sound. The model has no instrument family, no register, and no texture to work from, so it averages everything it associates with 'epic'. Replace it with a concrete family and a material word: 'wide brass pad' or 'dark bowed bass'.
Should I write prompts in English?
Yes, for now. The vocabulary the model was trained on is English sound design language. The interface and this blog are available in more languages, and the prompt tips translate directly - the words themselves work best in English.
What is the difference between a take and a new prompt?
A take is another generation of the same prompt - same description, different result. A new prompt changes the description. When the character is right but the detail is off, ask for takes; when the character is wrong, change the words.

