B-roll is everything on screen that isn't your face. A three-minute talking-head video with nothing but your face loses people, and everyone who has posted one knows it. So the question of how to add B-roll to a talking-head video comes up constantly, and the market has answered it with a dozen tools that drop stock footage onto your timeline in one click.
Most of that footage makes the video worse. Not because stock is bad, but because a clip of someone typing at a laptop, inserted because you said the word "work," doesn't do either of the two jobs B-roll exists to do. This post is about those two jobs, the three kinds of visual that do them, how to add each one in a few minutes, and a continuity trick that the one-click tools get wrong.
What B-Roll Is Actually For
B-roll has exactly two jobs in a talking-head video.
Job one: show the thing. You're talking about a product, a place, a screen, a before-and-after. The viewer needs to see it, and your face can't show it. This is what footage is for.
Job two: make the claim land. You just said "three out of four of my clients had the same problem." That's a number. Numbers on screen get remembered. Numbers only spoken get lost by the next sentence. Same for a list, a quote, a comparison, a definition. The visual's job is to carry what you said so the viewer doesn't have to hold it.
Everything else is decoration. Decoration isn't free. Every second the viewer is looking at a stock clip is a second they're not looking at you, and you're the reason they're watching. The tools that sell B-roll like to cite retention lifts of 30% or more. Treat that as a vendor number. In my experience, the retention comes from visuals that do one of the two jobs, and the generic ones cost you more than they add.
The Three Kinds of Visual, and When Each One Is Right
1. Footage: stock or your own
Right when you're showing a real thing or a real place. A clip of the actual product. The actual city. The actual screen you're describing. If you shot it yourself, better. If you didn't, stock is fine, as long as it's the specific thing and not a mood.
Wrong when it's standing in for an idea. There is no stock clip of "momentum" or "trust" or "scaling a business." What you get instead is a woman smiling at a laptop, and the viewer's brain files it under "ad."
2. A card
Right when the visual should carry the claim. A stat card for the number. A checklist for the list. A quote card for the line you want them to screenshot. A comparison for the before-and-after. A timeline for the sequence. A definition for the term you just coined.
A card is text and shape over your video, or briefly instead of it. It's the visual equivalent of writing the number on the whiteboard. It costs the viewer nothing to read, and it stays in the room after you've moved on.
3. A punch-in
Right when you want emphasis without leaving your face. A quick zoom on you, on the line that matters. It's the cheapest visual there is, it never breaks the connection, and it's what a good editor reaches for far more often than footage.
There's a fourth that's really a version of the first: a screenshot. If you're explaining something on a screen, show the screen.
How to Add B-Roll to a Talking-Head Video, Step by Step
I'll walk through this in CreateSocial because it's what I use, but the principles hold in any editor.
Step 1: Record first. Decide the visuals after.
The words decide the visual, so you need the words first. Record the video, get the transcript, then look at what you said. If you plan B-roll before recording, you end up recording toward the B-roll and the take goes stiff. Record like you're talking to one person, then dress it.
Step 2: Add footage where you're showing a real thing
In CreateSocial, the stock library lives inside the editor. Search for the specific thing, not the mood: "espresso machine," not "morning routine." Click the clip and it lands on a track above your video. Drag it to the moment you say the thing. Trim it short. Two to four seconds is almost always enough, because the viewer got the point in the first second and now they want you back.
The clip's audio is muted by default, and you should leave it that way. Stock ambient sound competing with your voice is the fastest way to make a video feel cheap. You can also drag in footage you shot yourself, and it goes on the same track.
Step 3: Click Visuals for the cards, punch-ins and screenshots
This is the part that used to be the slow part. In the editor toolbar, Visuals reads your transcript and suggests visuals anchored to the words you said. A stat card lands on the number. A checklist lands on the list. A punch-in lands on the line you leaned on. A screenshot slot lands where you said "here's what it looks like." Where a real clip would help, it suggests one of those too.
Each suggestion is a chip on the timeline. Click it and you can change the text, the look and the size right on the canvas. Drag it if the timing feels off. Delete the ones you don't want. It runs once when you open the editor, it's free, and it doesn't use any credits.
To be clear about what it is and isn't: it suggests. You decide. A video with twelve suggestions applied blindly looks like a slideshow, and that's on you, not the tool. Keep the ones that do one of the two jobs. Kill the rest.
Step 4: Check that each visual lands on its word
Play it back. A card that appears a second before the number, or a clip that shows up after you've already moved on, reads as a mistake even if nobody can say why. The visual should arrive on the word it belongs to. In CreateSocial that's what the anchoring does, but check anyway, especially after you've cut anything, because a cut earlier in the video moves everything after it.
Step 5: Approve and render
The card you see on the canvas is the card that renders in the final video, and so is the footage. Then captions for TikTok, Instagram, YouTube, LinkedIn, Facebook and X get written from what you said, and you can repurpose the same recording into a LinkedIn post, an X thread or a carousel. The B-roll only ever has to be done once.
The Continuity Trick Most Auto-B-Roll Tools Miss
This is the part I'd want someone to have told me before I built this.
Say you're walking through a five-point checklist. You mention the first point a minute in. The second one comes at two minutes, after a story. The third at three minutes. That's how people actually talk. The checklist isn't delivered in one breath. It's spread across the video.
Most tools treat every visual as its own event. Point one gets a card. Point two, a minute later, gets a different card. Or worse, one big card with all five points appears at the first mention, sits there for a minute over your face, and is long gone by the time you say point three. Either way the viewer never sees the checklist as a checklist.
When we first built Visuals, the suggestions looked good straight away. What the beat-timing work added is that a checklist stays one checklist. Items appear as you say them, one at a minute in, the next at two minutes, the next at three, on the same card, in the same place, so the viewer watches the list build across the video and sees the continuity in what you're saying. A timeline does the same thing. A comparison fills in as you cover each side.
It sounds small. It's the difference between visuals that follow your argument and visuals that interrupt it.
What the One-Click Tools Do Instead
To be fair to the field, the automatic B-roll tools are good at what they're for. Submagic's Magic B-Rolls "understands your transcript and contextually adds relevant b-roll footage" from a stock library in one click, and you can add clips by hand on any subtitle line. Opus Clip and Vizard add B-roll while they cut a long video into clips. AutoCut's AutoB-Roll pulls Storyblocks footage straight into a Premiere or DaVinci timeline for about $20 a month.
All of them answer the question "what footage matches these words?" That's job one, and if you're showing a real thing, they'll find it. None of them answer "what does this claim need on screen?" which is job two, and job two is where most of a talking-head video lives. They also treat each insert as independent, which is the continuity problem above.
If you already edit in Premiere and you want stock dropped in, AutoCut is a reasonable buy. If you want the visual to carry what you said, you need cards, and you need them to know what the previous card was.
Five Rules That Keep B-Roll From Making the Video Worse
The words decide. If you can't point to the word a visual belongs to, cut the visual.
One visual per claim, not per sentence. A visual every ten seconds because a template said so is a slideshow. Add one when there's a number, a list, a quote, or a real thing to show.
Don't cover your face for long. Footage over you for two to four seconds is a cutaway. Footage over you for twenty seconds is a different video with you narrating it.
Mute the footage. Always.
Prefer the punch-in. When you're tempted to add footage for emphasis, try a zoom on yourself first. It's usually better and it never looks like an ad.
Where This Fits
B-roll is the last ten percent of a talking-head video, and it only works if the first ninety is there. The video has to have something worth saying, which is what the knowledge base is for. It has to be recorded cleanly, which is why we record in sections with the script above the camera, and why the edit is mostly one-click passes rather than timeline surgery. I wrote about how that edit goes separately, and about looking natural on camera, which no amount of B-roll fixes.
Bottom Line
Add B-roll to a talking-head video for two reasons only: to show a real thing, or to make a claim land. Use footage for the first. Use a card or a punch-in for the second. Anchor every visual to the word it belongs to, and let a list build across the video instead of dumping it on screen at the first mention.
If you want to see what that looks like on your own video, the free trial includes the editor, the stock library and Visuals. Record something, click the button, and delete what you don't like. Start your free trial here.