How to tag B-roll so you can actually find it later

Subject, function and quality: the three axes of a tag system that still works after twenty projects, plus a starter vocabulary.

6 min read

Every editor has been here. You need a cutaway of hands on a keyboard. You know you shot one. You are fairly sure it was the corporate job, either the one in March or the one in June. Twenty minutes later you are still scrubbing, and you shoot a new one instead, because that is faster.

That is a tagging problem, and it is worth solving properly once, because B-roll is the footage with the longest useful life. An interview belongs to its project. A well-shot cutaway of rain on a window is useful for the next ten years.

Why most tagging systems collapse

Before the system, the failure modes, because they are consistent:

The vocabulary drifts. Project one uses interview. Project four uses int. Project nine uses talking-head. Now a search for one misses the other two, and the library quietly stops working.

The tags are too specific. coffee-cup-close-up-morning-light describes exactly one clip. Tags are for grouping. If a tag matches one file, it should have been in the filename.

The tagging happens at the wrong time. Tagging during the edit means you tag the clips you used, which are exactly the ones you already found. The value is in tagging what you did not use.

Coverage is partial. Forty of 300 clips tagged is worse than zero, because the search returns results and you believe them.

All four failures come from the same root: tagging done by hand, by a busy person, with no fixed vocabulary. Fix those three variables and the system holds.

The three axes

A tag set that works has exactly three dimensions. Not more.

1. Subject: what is in frame

The concrete, visible content. hands, keyboard, coffee, skyline, traffic, crowd, machine, whiteboard, rain, forest.

Rules: nouns, singular, lowercase, no compounds unless you need them constantly. This is the axis with the most terms, and the one an automatic description pass can fill in for you, because it is purely visible content.

2. Function: what the shot is for

The axis everyone forgets, and the one you search by most often when you are actually cutting.

  • establisher: sets the location, usually wide
  • insert: a close detail, cuts into a scene
  • texture: atmosphere, no information, for breathing room
  • transition: a whip, a pass-by, something to cut through
  • cutaway: covers a jump cut in an interview
  • action: someone doing the thing the piece is about
  • hero: the good one, the shot you would put on a title card

When you are in the timeline you do not think “I need a coffee cup”. You think “I need something to cut to for two seconds while the audio bridges”. That is a function search.

3. Role and quality: how usable it is

  • select or hero: reviewed, good
  • review: something is off, needs a look before you commit
  • reject: soft, shaky, unusable, but not deleted

Three values, not ten. The point is triage, not grading.

A starter vocabulary

Copy this, cut what you do not shoot, add five terms of your own, and then do not change it for a year.

AxisTags
Functionestablisher insert texture transition cutaway action hero
Rolearoll broll interview ambience
Qualityselect review reject
Lightgolden-hour blue-hour daylight night interior
Motionstatic handheld gimbal pan tilt slider aerial
Subjectproject-specific, 8 to 15 terms per shoot

The first five rows are stable across every project you will ever shoot. Only the last row changes. That stability is the whole trick: a vocabulary you do not renegotiate is a vocabulary you apply consistently.

Where to put the tags

Three places, and the order is deliberate.

Finder tags, on macOS. Indexed by Spotlight, searchable from every application’s open dialog, visible in Finder’s sidebar, and they travel with the file across drives. They are the highest-leverage and least-used feature on the platform. Full treatment in using Finder as a footage browser.

The filename. For the things that identify one specific clip rather than a group. Rooftop-Interview-Marcus-Wide is a name, not a tag. Names and tags do different jobs and both are needed.

A sidecar text file. For notes too long for either: why a take was rejected, what the client said, which lens. Spotlight indexes text files, so this becomes searchable content too.

Notice which is missing: your NLE. Premiere keywords and Resolve smart bins are genuinely good, and they live in a project file that will be closed forever in six weeks. Tag there in addition if you like. Never tag there only. That is the argument in finding clips without Premiere.

When to tag

Immediately after ingest, before the edit starts. Not during, not after.

The reason is unglamorous: after the edit, you are done with the project emotionally, and the tagging will not happen. Before the edit, tagging is part of getting to know the footage, which you have to do anyway.

A realistic pass on a 200 clip shoot:

  1. An analysis pass generates subject tags and descriptions unattended.
  2. You spend ten minutes adding function tags, because a machine cannot reliably tell an establisher from a texture shot. That distinction is about intent, not content.
  3. You mark the obvious select and reject clips while reviewing.

Total human time: around fifteen minutes. That is a number you will still be paying in project twenty, which is the only test that matters.

What can be automated and what cannot

Being precise about this saves disappointment.

Automatable, reliably: subject tags, scene type, time of day, camera movement, talking clip versus cutaway classification, obvious technical faults like sustained blur or shake. These are all visible properties.

Automatable, roughly: the function axis. A model can guess that a wide static shot of a building is an establisher. It will not know that you shot it as a transition.

Not automatable: hero, select, reject in any meaningful sense. Those are judgements about your project, and a model has no access to your project.

The practical split is that the machine fills the subject and technical axes across every clip, and you spend your fifteen minutes on the two axes that require having been there. That is a much better use of a human than typing “coffee cup” 40 times. The general version of this argument is in AI video asset management.

Testing the system

One test, run three months later: think of a shot you know you have, and try to find it in under a minute using only search.

If you cannot, the failure will be one of four things, and it will tell you what to fix:

  • No result at all: coverage gap, some clips were never tagged
  • Too many results: vocabulary too coarse, add a function axis
  • Wrong results: vocabulary drifted, two terms mean the same thing
  • You did not know what to search for: the tags are not the words you actually think in

That last one is the most common and the most fixable. Tag with the words you use out loud when you describe the shot to someone else, not with the words that sound like a taxonomy.


Cliptag runs the automatable part on macOS: it analyzes each clip, writes subject tags as real Finder tags, separates talking clips from cutaways, names files by content and files them into folders. What it leaves to you is the judgement, which is the part worth your time. Free plan, on-device mode free and unlimited on Apple Silicon.

Questions
How many tags should a B-roll clip have?

Three to six. One or two subject tags, one function tag saying what the shot is for, and a quality or role tag. More than six and nobody applies them consistently, which makes the whole set untrustworthy. Fewer than three and the search is too coarse to be useful.

Should I tag B-roll in my NLE or in the file system?

Both, if you can, but the file system one is the one that lasts. Bin structures and NLE keywords disappear when the project closes. Finder tags and descriptive filenames stay attached to the clip on the archive drive, where you will be searching in a year.

What is the difference between a subject tag and a function tag?

A subject tag says what is in frame: coffee, hands, skyline, crowd. A function tag says what the shot does in an edit: establisher, transition, texture, insert, cutaway. You search by subject when you know what you want to see, and by function when you know what the cut needs.

Is it worth tagging old B-roll retroactively?

Usually not all of it. Tag the archive you actually reuse, which for most people is a small fraction, and fix the ingest process so new material arrives tagged. Retroactive tagging of everything is a project people start once and never finish.

Try it on your own footage

Drop in a folder. Get back names, tags and transcripts.

Cliptag reads your video, photos and audio, names every file by what is actually in it and files it where you will find it again. Free plan, no account, and the on-device mode stays free and unlimited.

Download free for Mac Or try it in the browser
macOS 11+ · No account · No card
Keep reading
How to organize Sony footage without losing a single C0001.MP4 The Sony card structure explained, why two bodies overwrite each other, and a naming system that still makes sense after the project is done. Search videos by objects: what actually works, and what does not How object recognition on footage really performs, the difference between tags and descriptions, and why timecode-level search is a different problem. The best Finder alternative for videos might be a better-fed Finder What Finder is genuinely bad at, the four features that fix most of it, and when a dedicated browser actually earns its place.