Back to Blog
A weathered suitcase rests against ancient tree roots in a misty forest clearing - AI generated still from Journey
October 2, 2026

I Built an AI
Music Video on AWS

All opinions are my own - just me talking here, not my employer. This article is homegrown, unfiltered, and fully my responsibility.

I write music (occasionally now) and I build on AWS, so I set myself a challenge: take one of my original songs - specifically Journey, which I wrote in 2017 and which has never had a music video - and let AWS generative AI produce a full, lyric-synced music video for it.

The end result is a five-minute film generated entirely by AWS services, and the project went through quite a few generations as I learned what the stack does brilliantly and where it needs a human in the loop. Here is the real retrospective, warts and all.

I actually really loved the simplicity of the raw first run with Nova Reel 1.0 - it basically generated a pretty literal scene for each lyric. The only issue was those creepy AI-generated humans (we've all seen them before) which was entirely disconcerting.

The Stack

Service Role
Amazon Bedrock Managed platform - every model is one async API call, straight to S3. No GPUs to run, pay per generation.
Amazon Nova Reel Text-to-video; generated all the 6-second cinematic shots (Nova Reel 1.0 and eventually 1.1).
Amazon Nova Pro Chained in front of Nova Reel to write a full, meaning-driven prompt for each lyric.
Amazon S3 Automatic storage for every generated clip.
Kiro.dev The Agentic Code platform and IDE.

How It Evolved Across 3 Distinct Runs

Run 1 - v1.0

Nova Reel 1.0 - lyric wrapped in a fixed style tag. People appeared and were uncanny; scenes generic, but I liked the simplicity and how scenes matched the lyrics.

Run 2 - v2

Nova Reel 1.1 + Nova Pro - writing a one-sentence scene per lyric; environment-only, "no people" still in the prompt but we got silhouettes instead of AI faces.

Run 3 - v5 (Final Cut)

Nova Pro writes the ENTIRE per-lyric prompt - told to visualise the MEANING of each line through objects and places. Hard cuts, trimmed closing lyrics for breathing room, URL title card, and a closing message that fades to black.

How the Pipeline Works

The whole thing runs as a clean, repeatable pipeline. Bedrock's consistent async API made each stage genuinely simple to orchestrate:

  1. Provide a lyrics sheet with timings to get exact timestamps (lyrics stay the source of timing truth).
  2. Amazon Nova Pro expands each lyric line into a distinct cinematic scene prompt.
  3. Amazon Nova Reel generates a 6-second clip per unique scene, writing straight to S3.
  4. A de-duplication step reuses clips for repeated choruses, saving real money.
  5. Clips are stitched to the song with crossfades, a title card, and a closing message - all locally with FFmpeg.

The xfade timing was a real challenge: a simple change affected the length of clips and the overall video timing.

Testing and Iteration

This was wonderfully iterative - because Bedrock bills per generation, I could experiment freely and cheaply. Highlights:

The Actual Output

Run 1 - Reel 1.0 Run 2 - Reel 1.1 + updates Run 3 - Dialled to 11
Watch Watch on YouTube Watch on YouTube Watch on YouTube

What Changed Run to Run

Aspect Run 1 Run 2 Run 3
Video modelNova Reel 1.0Nova Reel 1.1Nova Reel 1.1
Prompt methodLyric + fixed style tagNova Pro writes a scene sentenceNova Pro writes the full per-lyric prompt
style.jsonAll fields usedAll fields usedstyle_bible and seed only
People / facesFrequent, uncannySilhouettes (not faces)None - objects imply people. Win.
Scene relevanceLooseBetterEach shot a visual metaphor for its line
VarietyRepetitiveVariedHighly varied, meaning-led
CameraDrifting pansNatural motionCalm, mostly static
TransitionsHard cutsCrossfadesHard cuts (deterministic)
EndingAbrupt tailVideo died at 04:04Beach to black, UNHCR message held

The Same Lyric, Three Ways

The clearest way to see the evolution - here is exactly what clip 1 (opening line: "I'm going on a journey") was asked to generate in each run.

Run 1 - Reel 1.0
The lyric itself, wrapped in a fixed style tag:

"scene": "I'm going on a journey" "prompt": "Cinematic dreamlike 35mm ... Scene: I'm going on a journey. Slow graceful camera ..."

I loved the simplicity here, but the backgrounds became 'samey' and the AI-generated humans were too much.

Run 2 - Reel 1.1 + Nova Pro
A one-sentence scene, then wrapped in the style scaffold:

"scene": "A solitary figure steps onto a worn path that winds through a misty forest, the soft amber glow of dawn casting long shadows across the moss-covered ground."

Less creepy human figures, but creepy silhouettes instead. Win..? Lots of forests and mist. The video also cuts/pauses around 04:04 as sync between the video and the fading got out of step.

Run 3 - Dialled to 11
Nova Pro writes the whole prompt, visualising MEANING through objects:

"prompt": "In a secluded forest clearing bathed in soft, hazy light, a lone, weather-beaten suitcase rests against the gnarled roots of an ancient tree, its lid slightly ajar as if hastily packed. Nearby, a worn leather journal lies open, pages fluttering, the quill resting on the last written word. A path winds away into the mist, inviting exploration."

In the tradition of Mythbusters, I wanted to throw in a tonne of C4 and crank it up to 11. I absolutely loved that during this run, the lyrics of "To Be Something We're Not" generated an empty stage with a red, flowing backdrop/curtain. There were some really clever uses of metaphor - intended or not.

What Worked Really Well

Pay-per-generation economics made fearless experimentation possible. The final video costs about $17.80 of generation.
The consistent async Bedrock API meant one orchestration pattern worked for every model and every job.
Chaining Nova Pro into Nova Reel was the single biggest quality unlock - richer, more varied, meaning-driven scenes than raw video prompts.
Letting Nova Pro visualise the MEANING of each lyric (a packed suitcase for "gotta leave you behind") produced imagery far stronger than literal scenery.
Smart de-duplication of repeated choruses cut about $4 and sped the run up, with zero quality loss.
Resumable generation with retry and pacing meant a transient capacity blip never forced a full, costly re-run.

What Didn't Work, and How to Work Around It

Generated humans were genuinely unsettling. Any person or face landed squarely in the uncanny valley. The root cause: Nova Reel takes a single positive prompt and has no negativeText parameter. Writing "no faces, no people" actually made it generate MORE faces. This is documented behaviour - AWS's best-practices guide says to avoid negation words like "no" and "not." Fix: Stop naming humans entirely. Describe only environments.
Samey landscapes, then over-corrected. "No people plus one locked style" trended toward empty vistas. Handing full creative control to Nova Pro fixed variety, but the balance between style consistency and creative freedom is still being found.
Elements Nova Reel renders poorly: cars driving backward, a family room with horror-movie vibes, a "train on multiple tracks" about to run over all the suitcases. Fix: get Kiro to rewrite those specific prompts without regenerating the whole song.
The crossfade saga. A chained xfade kept drifting the timeline. After several attempts, the robust answer was to drop xfade entirely and use deterministic plain concatenation with hard cuts. Correct every time.
Capacity throttling. Bursting async jobs at once tripped ServiceUnavailable errors. Solved cleanly with pacing and exponential-backoff retry - and because failed jobs are not billed, it cost nothing.

The Numbers

$17.80 Final video generation cost
37 Unique Nova Reel clips
6s Per clip
$93 Total Bedrock charges (all runs)
$113 Total project spend inc. Kiro
460 Kiro credits used

More math: the song is ~5 minutes (304s), but that's not what you're billed for. I had 46 lyric lines; de-duplication collapsed repeated choruses to 37 unique clips.

37 clips × 6 seconds = 222 seconds of generated video. 222 × $0.08 = $17.76. Nova Pro prompt-writing across all lyrics: under $0.05 per run - effectively free. De-duplication saved about $4 versus generating all 46 lines.

Note: the $0.08/sec figure is from public pricing. Confirm against your own AWS bill for exact numbers.

Lessons Learned

What's Next

The Takeaway

I absolutely loved this whole process. Nova Reel and Pro are amazing, reasonably priced, and iterating on AWS is fantastic. This combination lets me tick the weird combination of technology and creativity that I haven't been able to merge before (outside of using DAWs like Logic). Can't wait to build and expand on it.

One original song, a few sittings, and the AWS generative-AI stack: a finished, lyric-synced music video. Amazon Bedrock, Nova Reel and Nova Pro made the happy path genuinely fast, and the rough edges - creepy humans, the no-negative-prompt quirk, landscape sameness, crossfade drift - all had sensible engineering answers. That balance of powerful managed AI plus someone willing to iterate is exactly what makes this fun.

On a more serious note - regardless of the many fun and creative video snippets brought to life with the latest GenAI technology, the underlying message across all three videos remains the same. This song will always be for anyone who has been forcibly driven from their home by conflict, persecution or violence. May they find comfort, peace and, above all, safety.

- James Scanlon

#AWS #AmazonBedrock #NovaReel #NovaPro #GenerativeAI #AIVideo #MusicTech #BuildOnAWS