Video Layer
The Video and Voice Pipeline
We film real software screens with a script, not a person clicking and hoping. One steady voice take reads the words, Whisper checks every take before it ships, and captions go on every film.
- Year 1Website
- Year 2Map Listings
- Year 3Booking Tool
- Year 4Front Desk
- Year 5Payments
- Year 6Reports
Same six pieces.
Ten layers, already running.
The Short Version
- Screens are captured by a script, so every click lands the same way every time.
- One continuous voice take keeps the tone and pace even from start to end.
- Every take is transcribed by Whisper and checked against the script.
- Training videos get voice and captions, never music.
| Screen capture | Playwright, scripted |
|---|---|
| Voice | ElevenLabs, one continuous take |
| Take check | Whisper transcript vs script |
| Numbers | Spelled out as words |
| Captions | On every film |
| Music | None on training videos |
- ScriptScenesNarrationNumbers as words
- CapturePlaywrightReal app screens
- VoiceElevenLabs takeWhisper check
- AssembleCut to voiceCaptions
- ShipHelp pageYour site
Why We Script the Screen
Most training videos are someone sharing their screen and talking. The mouse wanders. They click the wrong tab and go back. They say um. Then the software changes and the whole video is out of date.
We build ours like a small program. A script says which screen to open, what to click and how long to hold. Playwright, a browser tool, follows that script and records the screen. Every click lands in the same spot every time.
That means a video is not a one time recording. When a screen changes, we update one line in the script and film it again. The new version matches the software your staff actually see.
import os
from playwright.sync_api import sync_playwright
APP_URL = os.environ["APP_URL"]
with sync_playwright() as p:
browser = p.chromium.launch()
context = browser.new_context(
viewport={"width": 1920, "height": 1080},
record_video_dir="takes",
record_video_size={"width": 1920, "height": 1080},
)
page = context.new_page()
page.goto(APP_URL + "/jobs/")
page.wait_for_timeout(1500)
page.get_by_role("button", name="New job").click()
page.wait_for_timeout(2000)
context.close()
browser.close()The video file is saved when the context closes. We film with demo data only, so no real customer ever shows up on screen.
One Voice Take, Read Steady
The voice comes from ElevenLabs. We tried making one clip per scene and stitching them. It sounded wrong. Each clip starts a little fresh, so the tone and speed jump at every cut. People notice even if they cannot say why.
So the whole script goes in as one continuous take. The voice keeps one pace from start to end. We then cut the screen footage to match the voice, not the other way around.
We also write the script for the ear. Numbers are spelled as words. A voice model can read a price or a year in odd ways, and a written word leaves no room to guess. We keep the settings steady and avoid lines that sound like questions but are not, because the voice lifts at the end and sounds unsure.
Every Take Checked by Whisper
A voice model can skip a word, repeat one or say it wrong. You will not catch it by skimming. So every take is run through Whisper, a speech to text model, and the transcript is compared to the script.
import difflib
import re
def words(text):
return re.findall(r"[a-z']+", text.lower())
def take_matches(script, transcript, floor=0.97):
ratio = difflib.SequenceMatcher(None, words(script), words(transcript)).ratio()
return ratio >= floor, ratioIf the take falls under the floor, it does not ship. We look at where the two differ, fix the line or make a new take, and check again. Whisper also gives us word timings, which is how the captions line up with the voice.
How the Pieces Fit Together
Each film starts as one plain text file. It lists the scenes in order. Each scene has the screen to open, the clicks to make and the lines to say over it.
From that one file, the pipeline does the rest.
- The narration lines are joined and sent for one voice take.
- Whisper transcribes the take and returns a time for every word.
- Those times tell us how long each scene needs to be on screen.
- Playwright films each scene and holds it for that long.
- The footage and voice are joined, and captions are laid on from the word times.
Because the script is the single source, nothing drifts. The words on screen, the words spoken and the words in the captions all come from the same lines. Change a line and all three change together on the next run.
It also means a film can be rebuilt by anyone who can run the script. There is no editing project on one person's machine that only they understand.
Captions and No Music
Many people watch with the sound off, at the counter or in the bay. So every film has captions. They come from the script, timed from the Whisper check, so they say exactly what the voice says.
Training videos get no music. Music fights the voice for attention, and it makes a how-to feel like an ad. Staff learning a screen want the words and the clicks, nothing else.
What It Saves the Owner
You stop training each new hire by hand on the same screens. You stop paying to reshoot a video every time the software changes. You get a library of short films that match what your staff see today.
Client films live on your own site at full 1080p, not on a video platform that squeezes them. This is how we make Shop Desk training films, and it is part of our content and video work.
What We Learned the Hard Way
Listen to the take, and also read the transcript. Ears miss a dropped word when the rest sounds good. The transcript does not.
Film only what is true today. A script can click a button that is still being built. If a feature is not live, it does not go on camera, however good it looks.
Watch the file size. Some upload paths cap the size and quietly crush the picture to fit. A crisp screen turns to mush and text you meant to show can no longer be read. We check the final file at full size before it goes out.
For the step by step, read the AI voiceover training video guide. Next in the stack is automation and scheduling, which runs the jobs that keep all of this current.
Questions People Ask
Can you make training videos for my staff?
Yes. We film your real software screens with a script, add a steady voice and captions, and put the films on your own site.
What happens when my software changes?
We update the script for that screen and film it again. There is no need to reshoot the whole video by hand.
Why is there no music in the videos?
Music competes with the voice and makes a how-to feel like an ad. Staff learning a screen follow better with just the words and the clicks.
Is the voice a real person?
No. It is an ElevenLabs voice. We check every take with Whisper against the script so it says exactly what it should.
We sell time
Get Years of Building Switched On in Hours
A thirty minute call. A written price. Nothing built until you say yes.