WrinkleHQ

Build · Video

The Training Film Pipeline That Films Real Software

Our training films show the real software on screen, driven by a script, with one steady voice and clear captions. Every take is checked by a second machine before anyone watches it, and there is no music.

6 years
  1. Year 1Website
  2. Year 2Map Listings
  3. Year 3Booking Tool
  4. Year 4Front Desk
  5. Year 5Payments
  6. Year 6Reports

Same six pieces.
Built once. Switched on in hours.

A job card with the checklist the tech works through.
A job card with the checklist the tech works through.

The Short Version

  • Playwright drives the real screens, so the film shows the software that ships.
  • ElevenLabs reads the whole script in one continuous take, so the voice stays steady.
  • Whisper transcribes every take and flags a garbled number before it reaches a viewer.
  • Captions come from the voice timing. No music, ever, on a training film.
At a Glance
ScreensReal app, driven by Playwright
Data on screenInvented people only
VoiceElevenLabs, one continuous take
Voice checkWhisper on every take
CaptionsBurned in from voice timing
MusicNone
How a training film is made
  1. ScriptScenesNumbers as words
  2. VoiceOne ElevenLabs takeSteady settings
  3. CheckWhisper transcriptCompare to script
  4. ScreensPlaywrightReal appInvented data
  5. RenderTimelineCaptionsFinal video

The Hours Training Eats

Every new hire asks the same questions. How do I send a quote? Where do the photos go? What do I press when the customer says yes? Someone senior stops what they are doing and shows them. Then the next hire starts, and they show it again.

Say a shop hires four people a year and each one takes ten hours of showing. That is forty hours of your best person's time, spent saying the same things. And it is not the same each time. It depends on who is tired that day.

A short film fixes that. The trouble is that most training films are hard to make and go stale fast. A screen recording with someone saying um, over music, of a screen that changed last month.

Sending an estimate for approval. The customer gets a text with a link to their own page.
Sending an estimate for approval. The customer gets a text with a link to their own page.

What We Built

We built a pipeline that makes the film from a script. The script says what is on screen and what the voice says, scene by scene.

  1. The voice. ElevenLabs reads the whole script in one continuous take. One take means one tone. Stitching short clips together makes the voice speed up and slow down between lines.
  2. The check. Whisper listens to the take and writes down what it heard. We compare that to the script. A number read wrong, or a line that rose like a question, shows up here and not in front of a viewer.
  3. The screens. Playwright opens the real app and clicks through it, scene by scene. The film shows the software that ships, not a mock drawn to look like it.
  4. The render. The timeline cuts the pauses at scene changes, lines the screens up with the voice, and burns in captions from the voice timing.

There is no music. A training film is watched by someone trying to learn a click. Music is noise on top of the thing they came for.

How It Is Wired

The real app runs against a small stand in database full of invented people. Names, cars and plates are made up. The app does not know the difference, so every screen looks and acts the way it does on a live day. Nothing on camera belongs to a real customer.

Captions come from the voice itself. The voice service returns the time of each character, and the pipeline groups those into caption lines. Numbers are spoken as words so the voice reads them cleanly, then the captions turn them back into digits so they are easy to read.

The voice settings are fixed and steady, with no style added and normal speed. Those settings are the only thing that changes the pace of the voice, so we do not let each film pick its own. The same pipeline is covered in more depth on the video and voice pipeline page.

An engine opened up mid repair, heads coming off.
An engine opened up mid repair, heads coming off.

What Broke and What We Changed

A Number the Voice Mangled

An early script had a long number in digits. The voice stumbled on it, and so did others near the end. We now write every number as words before recording. The captions still show digits.

A Voice That Sped Up and Slowed Down

One take used looser settings meant for a more lively read. Over a few minutes, the pace drifted. We moved every training film to the steady settings and kept them there.

A Statement That Sounded Like a Question

A line that opens with if and ends on a short command can come out rising, like a question. We rewrite those lines as two flat statements. Then Whisper and a listen confirm it.

Paths the Shell Rewrote

On one machine, the shell quietly turned app routes into folder paths before the browser saw them. Scenes opened the wrong page. We turned that rewriting off for the pipeline and added a check that each scene lands where the script says.

What the Pipeline Deliberately Does Not Do

  • It does not film a live account. Real customers, real plates and real prices never appear. The stand in data is invented from the start.
  • It does not film what is not built. If a feature is on the roadmap, it is not on camera. A film that shows a screen that does not exist trains people on software they cannot use.
  • It does not add music. Voice and captions only.
  • It does not skip the check. No take goes to render without a Whisper transcript read against the script.

What It Means When You Switch It On

Your team gets a short film for each job they need to learn. It shows the real screen, in a calm voice, with captions they can read with the sound off in a noisy shop. When the software changes, the film is made again from the script. Nobody records their screen at night.

Your senior people stop giving the same tour. New hires send each other a link. And the film for one technician who keeps missing one step is one clip, not a whole course.

To make one yourself, read the guide to an AI voiceover training video.

Questions People Ask

Does the training film show my real customers?

No. Films are made against invented data. Real names, plates and prices never appear on screen.

Why is there no music?

People watch training to learn a click. Music makes the voice harder to follow, so we leave it out.

What happens when the software changes?

The film is made again from its script. Playwright clicks through the new screens and the voice is recorded again.

Can my team watch with the sound off?

Yes. Every film has captions timed to the voice.

We sell time

Get Years of Building Switched On in Hours

A thirty minute call. A written price. Nothing built until you say yes.

Get Your Time Back