Today, you have a:

60-second video.

The footage is already edited.

You then open Lyria 3.5,

and enter:

“Make me a 60-second warm, tasteful instrumental.”

AI generates it.

The style:

Perfect.

The instruments:

Sound great too.

But when you put it inside the video,

you realize:

It doesn’t fit at all.

The product only just appears,

and the music is already full.

The key scenes show up,

but the music doesn’t change.

The video is about to end,

but the track feels like:

it still has two more minutes to play.

The issue isn’t that:

AI can’t compose music.

But rather:

you only told it:

what the music should sound like.

But you didn’t tell it:

how the music should move with the video.

Today, learn just one method.

Before making video music,

first segment the timeline into:

Opening.

Main section.

Ending.

Then write one sentence for each segment, describing:

how the energy changes.

Why is this method especially useful for Lyria 3.5?

Google DeepMind’s Lyria Prompt Guide already lists:

Dynamics

as a key musical control element.

Dynamics can be understood as:

the music’s energy,

volume,

and density,

changing over time.

For example:

start with quiet piano,

gradually add rhythm,

build into a stronger chorus,

then quiet down at the end.

In other words,

don’t just describe:

“I want this style.”

You can also describe:

“How the music changes in the middle.”

This fits perfectly for:

video scoring.

Step 1|Don't touch music yet—watch the video first

Assuming you have a:

60-second product video.

Don’t ask yet:

Lo-fi? Jazz? Acoustic?

Just watch:

the footage itself.

For example:

0–10 seconds:

The product appears for the first time.

10–50 seconds:

Production process, details, usage scenes.

50–60 seconds:

Finished product, logo area, ending.

You naturally get three segments:

Opening.

Main section.

Ending.

Note:

This doesn’t mean you must cut:

10/40/10 seconds every time.

The real point is:

to find the three phases naturally in your video.

Step 2|Write energy levels only for each segment

Don’t start with music theory.

Use plain language first.

Opening:

Simple.

Main section:

Gradually building.

Ending:

Clean finish.

That’s enough.

Another video might be:

Opening:

Immediately energetic.

Main section:

Maintain the rhythm.

Ending:

Suddenly stop cleanly.

Also fine.

You’re not:

composing yet.

You’re deciding:

what the audience should feel at each stage.

Step 3|Add musical features next

Say it’s a:

handmade food video.

You want it to be:

warm,

natural,

with some rhythm,

and not overpower the voiceover.

You can first define:

Genre:

Acoustic/Light Folk.

Instruments:

acoustic guitar,

soft piano,

light percussion.

Vocal:

None.

Only then combine:

music style

with

the three-part energy flow.

The prompt evolves from vague to detailed

From:

“Make a 60-second warm background music.”

To:

“Create instrumental background music for a 60-second handmade food video. Overall warm and natural, featuring acoustic guitar, soft piano, and light percussion. The opening should be simple with fewer instruments; the main section gradually adds rhythm and layers; towards the end, energy decreases gradually and finishes cleanly. No vocals, and avoid being overly dramatic.”

This prompt has no:

complex chords,

advanced theory,

or mixing jargon.

Yet it’s much more useful than:

“Make me some nice music.”

Why?

Because it answers two different questions.

The first:

What style of music?

The second:

How should the music move with the visuals?

For real video scoring,

the second question is often more important than the first.

The opening is not for throwing all the best sounds at once

Many AI-generated tracks start off:

full of sound.

Drums,

bass,

pads,

melodies,

all at once.

It may sound nice on its own,

but in the video,

a problem arises.

The images:

haven’t yet established anything.

The product:

has just appeared.

The music is already:

100% intense.

Leaving no room to grow.

If the video needs:

gradually building emotion,

the prompt can specify:

Opening should be sparse and restrained.

Simply put:

fewer instruments,

less rhythm,

and space to breathe.

Especially important if there’s voiceover at the start

Say the opening narration is:

“Do you spend a lot of time organizing the same data every day?”

If the music starts with:

loud drums,

busy melodies,

and vocals,

viewers have to process:

voiceover,

music,

images,

and subtitles simultaneously,

which is exhausting.

So the opening can be requested to be:

simple,

vocal-free,

with lower instrument density,

leaving space for content.

The main section ramps up

When the video enters:

demonstrations,

work steps,

product details,

before-and-after shots,

the music starts to:

add rhythm,

add layers,

and raise energy.

This gives viewers a sense that:

things are progressing.

This is where dynamics become truly useful.

It’s not to make music flashy,

but to:

create a sense of progress for the video.

Example: a coffee video

First 8 seconds:

shopfront,

coffee beans,

cup.

Quiet.

8–45 seconds:

grinding, brewing, steam, pouring milk, latte art.

Rhythm gradually builds.

45–60 seconds:

finished cup is served, final frame freezes.

Music:

begins to fade.

Ends cleanly.

So even without lyrics,

viewers feel:

the story

beginning,

progressing,

and completion.

The last segment—ending—is often forgotten

Many people tell AI only about:

what they want at the start,

and in the middle,

but not how to end.

So when the video reaches the end card,

the music is still:

building up,

or abruptly cut.

This conveys a feeling of:

“not really finished, just file ran out.”

So add a closing line in your prompt:

End naturally.

Lower energy gradually.

Leave a clean ending.

This is very useful.

Even for 15-second ads, keep the three segments

Don’t assume a short video

doesn’t have structure.

For example:

First 3 seconds:

Hook.

3–12 seconds:

Product.

12–15 seconds:

Call to action.

You can write:

Opening:

Immediately rhythmic.

Main section:

Keep driving forward.

Ending:

Quick closure with a clear stop.

The segments don’t need to be equal in length,

just serve different functions.

What about a 3-minute video?

Same principle.

For example:

0–20 seconds:

Opening.

20 seconds–2 minutes 30 seconds:

Main section.

Last 30 seconds:

Ending.

You can further divide within the main section:

start steady,

build up,

then wind down.

Google DeepMind encourages users to describe:

how dynamics change in different music sections.

If you’re working on longer content,

you can get more detailed.

But that’s not necessary today.

Start with the:

three segments,

done well.

One important note: don’t treat seconds like exact video editing commands

This point is crucial.

Lyria 3.5 lets your prompt influence:

track length.

Google’s API docs clearly state:

the full model generates several minutes of music,

and length depends on prompt.

But this doesn’t mean:

if you write,

“kick drum exactly at 10.000 seconds,”

it behaves like DAW automation

precise down to the grid.

Lyria is:

a generative model,

not a timeline-based editor.

So what we call today is a:

music energy line,

not a

precise cue sheet.

If you really need a specific sound at an exact second

For example:

ad explosion at 8 seconds,

stop at 23 seconds,

a more reliable approach might be:

first let Lyria generate:

roughly the music you need,

then use:

a proper editing software to

adjust,

cut,

fade,

and line up precisely.

AI helps you create musical raw materials,

while the editor locks down timing.

Don’t ask one tool to do everything perfectly at once.

So for your first test, don’t write too many segments

It’s tempting to see that Lyria supports:

Genre,

Tempo,

Dynamics,

Instruments,

Vocals,

and immediately write:

0–3 seconds like this,

3–7 seconds like that,

7–11 seconds switching again,

resulting in 20 lines.

That’s not necessarily better.

For your first try,

keep it to three segments:

Opening,

Main section,

Ending.

See if AI can grasp the big energy directions.

If it’s not right, change only one thing at a time

For example, first version:

Main section feels flat.

Don’t rewrite the whole prompt.

Just revise:

“Main section should gradually build rhythm and instrument layers more clearly.”

Generate again.

Second version:

Ending is still too sudden.

Change to:

“The last few seconds gradually reduce instruments and end naturally.”

Adjust one thing at a time.

You’ll better understand:

which prompt line

really matters.

This is important because Lyria API is currently single-turn

Google’s API docs clearly say,

music generation with Lyria 3.5 is a:

single-turn process.

Meaning:

You can’t generate a track, then keep chatting with the same piece to say:

“Just fix the section at 35 seconds,”

“Leave everything else unchanged,”

“Lower the drum volume.”

No full multi-turn audio editing yet.

If the first version is not right,

usually you modify the prompt,

and regenerate a new version.

So starting with a clear structure in your prompt

is more critical.

And the same prompt can generate different results each time

Google also reminds:

even with the same prompt,

variation exists between outputs.

If you find:

a version with just the right groove,

save it first.

Don’t think:

“I have the perfect prompt; I’ll generate the same again later.”

Generative music is not:

a fixed search result.

This means a better testing strategy is to generate 2–3 versions at once

Don’t just create one and immediately doubt:

the prompt is wrong.

Stick to the big direction,

try a few versions,

and ask:

Which has the cleanest opening?

Which has the most progressive main section?

Which ends best?

You’re:

choosing a music draft,

not expecting a perfect master on the first try.

If your video has voiceover, add a small reminder

The main section can build energy,

but don’t overshadow the voiceover.

Add to prompt:

“Keep the arrangement supportive under spoken narration.”

This means:

music provides mood,

but main melody or instruments:

don’t compete with vocals.

Especially for:

tutorials,

interviews,

product demos,

where voiceover is:

the star.

If no voiceover, music can carry more story

Examples:

travel,

cooking,

handcraft,

fashion,

product aesthetics videos.

Music can be:

more prominent.

Opening:

set atmosphere.

Main section:

drive forward.

Climax:

build intensity.

Ending:

release.

Because there is no voiceover to accommodate,

the three-part method adjusts

depending on presence of spoken content.

This method also works for photo-to-music generation

Lyria can use:

images

as generation context.

Say you have a:

sunset beach photo.

The photo helps AI grasp:

mood,

location,

and atmosphere.

But you can still add:

“Quiet at first, gradually adding warm rhythm in the middle, ending with simple piano.”

Photo sets:

mood,

energy lines handle:

movement.

Together,

this is more controllable

than just saying:

“Make music based on this image.”

This is actually a key method for all generative AI

Many users give prompts describing only:

what the result looks like.

But good prompts also describe:

how the result evolves.

For writing:

start by explaining the problem,

then provide solutions,

and finish with a call to action.

For videos:

first the hook,

then demonstration,

finally the CTA.

For music too:

build up,

progress,

then close.

AI hasn’t changed the need for:

good content structure,

just helped you:

produce structured output faster.

Today, just remember the three parts

Next time you use Lyria to score a video,

don’t just write:

“Make a good song.”

First jot down on paper:

Opening: _______

Main section: _______

Ending: _______

For example:

Opening:

simple, quiet.

Main section:

gradually build rhythm.

Ending:

lower energy, clean finish.

Then add:

Genre,

instruments,

vocals or not,

total length.

This is already a very complete first version prompt.

Your first test can take only a minute

Grab a short video you’ve already edited.

Watch it once.

Ignore the original music.

Ask yourself:

For the opening:

what feeling do viewers need when they first arrive?

In the middle:

when the image gets active, should music increase?

At the end:

how to let viewers know things are done?

Write these three answers

into your prompt.

You’re not:

asking AI to randomly make:

“music somewhat fitting the theme.”

You’re telling it:

what role the music plays in your content.

Remember one last thing

What video music really needs to match,

is not the video’s “theme,”

but the video’s “time.”

The same coffee video can be:

warm,

luxurious,

or upbeat.

But if the music’s

start,

build,

transition,

end

are completely out of sync with the video,

it’s just a standalone nice song.

Not a good soundtrack.

So when using Lyria 3.5,

don’t rush for the perfect prompt.

Tell the simplest three things clearly first:

How to start.

How to build.

How to end.

The AI has a better chance to truly:

work with the video.

Today, improve a little with AI.

Learn one AI trick a day.

Save some time daily.

Level up bit by bit.

SasaDaily, growing with you.

Recommended Reading

Today's AI Tool|2026/07/21: Google Vids Transforms Docs and Presentations into Video Drafts, Perfect for Teaching, Pitching, and Work Explanations

AI Can Imitate a Style in Seconds, But Creators May Have Spent a Decade to Get There