All opinions are my own - just me talking here, not my employer. This article is homegrown, unfiltered, and fully my responsibility.
I write music (occasionally now) and I build on AWS, so I set myself a challenge: take one of my original songs - specifically Journey, which I wrote in 2017 and which has never had a music video - and let AWS generative AI produce a full, lyric-synced music video for it.
The end result is a five-minute film generated entirely by AWS services, and the project went through quite a few generations as I learned what the stack does brilliantly and where it needs a human in the loop. Here is the real retrospective, warts and all.
I actually really loved the simplicity of the raw first run with Nova Reel 1.0 - it basically generated a pretty literal scene for each lyric. The only issue was those creepy AI-generated humans (we've all seen them before) which was entirely disconcerting.
The Stack
| Service | Role |
|---|---|
| Amazon Bedrock | Managed platform - every model is one async API call, straight to S3. No GPUs to run, pay per generation. |
| Amazon Nova Reel | Text-to-video; generated all the 6-second cinematic shots (Nova Reel 1.0 and eventually 1.1). |
| Amazon Nova Pro | Chained in front of Nova Reel to write a full, meaning-driven prompt for each lyric. |
| Amazon S3 | Automatic storage for every generated clip. |
| Kiro.dev | The Agentic Code platform and IDE. |
How It Evolved Across 3 Distinct Runs
Run 1 - v1.0
Nova Reel 1.0 - lyric wrapped in a fixed style tag. People appeared and were uncanny; scenes generic, but I liked the simplicity and how scenes matched the lyrics.
Run 2 - v2
Nova Reel 1.1 + Nova Pro - writing a one-sentence scene per lyric; environment-only, "no people" still in the prompt but we got silhouettes instead of AI faces.
Run 3 - v5 (Final Cut)
Nova Pro writes the ENTIRE per-lyric prompt - told to visualise the MEANING of each line through objects and places. Hard cuts, trimmed closing lyrics for breathing room, URL title card, and a closing message that fades to black.
How the Pipeline Works
The whole thing runs as a clean, repeatable pipeline. Bedrock's consistent async API made each stage genuinely simple to orchestrate:
- Provide a lyrics sheet with timings to get exact timestamps (lyrics stay the source of timing truth).
- Amazon Nova Pro expands each lyric line into a distinct cinematic scene prompt.
- Amazon Nova Reel generates a 6-second clip per unique scene, writing straight to S3.
- A de-duplication step reuses clips for repeated choruses, saving real money.
- Clips are stitched to the song with crossfades, a title card, and a closing message - all locally with FFmpeg.
The xfade timing was a real challenge: a simple change affected the length of clips and the overall video timing.
Testing and Iteration
This was wonderfully iterative - because Bedrock bills per generation, I could experiment freely and cheaply. Highlights:
- Side-by-side comparison cuts between Nova Reel 1.0 and 1.1 to evaluate quality.
- A forced
style.jsonfor each video creation request to ensure a consistent theme. - Prompt-engineering experiments that dramatically improved scene variety and composition.
- Resumable generation with automatic retry and pacing - transient capacity blips never cost a full re-run.
- A job-status poller that told me exactly when every async Bedrock job had finished.
The Actual Output
| Run 1 - Reel 1.0 | Run 2 - Reel 1.1 + updates | Run 3 - Dialled to 11 | |
|---|---|---|---|
| Watch | Watch on YouTube | Watch on YouTube | Watch on YouTube |
What Changed Run to Run
| Aspect | Run 1 | Run 2 | Run 3 |
|---|---|---|---|
| Video model | Nova Reel 1.0 | Nova Reel 1.1 | Nova Reel 1.1 |
| Prompt method | Lyric + fixed style tag | Nova Pro writes a scene sentence | Nova Pro writes the full per-lyric prompt |
| style.json | All fields used | All fields used | style_bible and seed only |
| People / faces | Frequent, uncanny | Silhouettes (not faces) | None - objects imply people. Win. |
| Scene relevance | Loose | Better | Each shot a visual metaphor for its line |
| Variety | Repetitive | Varied | Highly varied, meaning-led |
| Camera | Drifting pans | Natural motion | Calm, mostly static |
| Transitions | Hard cuts | Crossfades | Hard cuts (deterministic) |
| Ending | Abrupt tail | Video died at 04:04 | Beach to black, UNHCR message held |
The Same Lyric, Three Ways
The clearest way to see the evolution - here is exactly what clip 1 (opening line: "I'm going on a journey") was asked to generate in each run.
Run 1 - Reel 1.0
The lyric itself, wrapped in a fixed style tag:
I loved the simplicity here, but the backgrounds became 'samey' and the AI-generated humans were too much.
Run 2 - Reel 1.1 + Nova Pro
A one-sentence scene, then wrapped in the style scaffold:
Less creepy human figures, but creepy silhouettes instead. Win..? Lots of forests and mist. The video also cuts/pauses around 04:04 as sync between the video and the fading got out of step.
Run 3 - Dialled to 11
Nova Pro writes the whole prompt, visualising MEANING through objects:
In the tradition of Mythbusters, I wanted to throw in a tonne of C4 and crank it up to 11. I absolutely loved that during this run, the lyrics of "To Be Something We're Not" generated an empty stage with a red, flowing backdrop/curtain. There were some really clever uses of metaphor - intended or not.
What Worked Really Well
What Didn't Work, and How to Work Around It
The Numbers
More math: the song is ~5 minutes (304s), but that's not what you're billed for. I had 46 lyric lines; de-duplication collapsed repeated choruses to 37 unique clips.
37 clips × 6 seconds = 222 seconds of generated video. 222 × $0.08 = $17.76. Nova Pro prompt-writing across all lyrics: under $0.05 per run - effectively free. De-duplication saved about $4 versus generating all 46 lines.
Note: the $0.08/sec figure is from public pricing. Confirm against your own AWS bill for exact numbers.
Lessons Learned
- Keep lyrics as the source of truth. Amazon Transcribe struggles with sung vocals over a full music mix, so I used my known lyrics and created a manual timing sheet.
- Prompt positively. With no negative-prompt support, describe what should appear - not what shouldn't.
- Direct the MEANING, not just the scene. Nova Pro turning each lyric into a visual metaphor was the difference between pretty wallpaper and a film that started to say something.
- Model chaining beats single-prompt generation for variety and control.
- For long sequences, use hard cuts. Deterministic plain concatenation is far more reliable than chained crossfades.
- Design for failure. Pacing, retries, resumability and a status poller turn a flaky long run into a calm one.
- Text-to-video excels at landscapes, objects and atmosphere. It is still weak at believable humans - design around that strength.
What's Next
- Stand up EC2 GPU instances (g5 / g6e / p-family) to trial open-weight models (FLUX, Wan, HunyuanVideo) - these support negative prompts and may handle people far better.
- Generalise the pipeline into a reusable template (
novadreamstate) so any song with timed lyrics can have a music video scored on AWS. - Test other available 3rd-party models (when I can pull myself away from the PS5).
The Takeaway
I absolutely loved this whole process. Nova Reel and Pro are amazing, reasonably priced, and iterating on AWS is fantastic. This combination lets me tick the weird combination of technology and creativity that I haven't been able to merge before (outside of using DAWs like Logic). Can't wait to build and expand on it.
One original song, a few sittings, and the AWS generative-AI stack: a finished, lyric-synced music video. Amazon Bedrock, Nova Reel and Nova Pro made the happy path genuinely fast, and the rough edges - creepy humans, the no-negative-prompt quirk, landscape sameness, crossfade drift - all had sensible engineering answers. That balance of powerful managed AI plus someone willing to iterate is exactly what makes this fun.
On a more serious note - regardless of the many fun and creative video snippets brought to life with the latest GenAI technology, the underlying message across all three videos remains the same. This song will always be for anyone who has been forcibly driven from their home by conflict, persecution or violence. May they find comfort, peace and, above all, safety.
- James Scanlon