How to Keyword Stock Video: What Photo Rules Miss
Most keyword advice is written for photographs, and most of it carries over to video. But a clip is not a photo with more frames. Five things change when the picture moves — what counts as true, how you name the camera's motion, how a subject stays consistent across a batch, what date the file carries, and how one list travels between agencies. This guide covers those five, with what we measured on real footage.
The short version
For the technical side — codecs, frame rates, durations, file-size caps and CSV columns for nine agencies — see our stock video metadata requirements. This post is about the words.
1. A keyword has to be true for the whole clip
When you keyword a photo, the question is "what is in this picture?" For a clip it is "what is in this picture from the first frame to the last?" — and the answer is often different at 0:02 and at 0:12.
- Someone enters or leaves. A street that is empty for eight seconds and then gets a cyclist is not a No People clip. If the arrival is the point, Entering or Arrival may be the keyword that sells it.
- The light changes. A clip that goes from dusk to dark is not simply Night; if it is a time lapse, the vocabulary even has Day to Night Time Lapse.
- The action turns over. A cook who chops, then plates, then serves has three actions. Keyword the ones a buyer would cut the clip for, and leave out anything that only happens in the last second.
The practical habit: before you keyword, look at the start, the middle and the end of the clip, not just the thumbnail your editor picked. A list written from the thumbnail is the most common way a video ends up with a keyword that is false for most of its running time.
2. Camera movement is a keyword — and you cannot see it in a still
Buyers filter for motion. An editor who needs a slow push past a product, or a static shot to lay text over, searches for exactly that. And agencies use it too: Shutterstock's own rejection guidance counts the same subject shot at different clip lengths or with different camera movements as similar content — see why stock photos and clips get rejected. So the movement is both something buyers search for and something that sets one clip apart from its siblings.
Here is the catch. Movement exists between frames, and almost every automated tagger, and every person keywording from a contact sheet, looks at single frames. We tested how badly that goes.
We shot six short phone clips, 4–7 seconds each, one deliberate camera move per clip: a pan, a tilt up, a zoom in, a tracking shot, a handheld shot and a locked-off static shot. Then we asked a vision model to keyword each clip from three stills taken across it, the way most tools do.
| What we shot | What the model said from 3 stills |
|---|---|
| Pan | nothing |
| Tilt up | nothing |
| Zoom in | nothing |
| Tracking shot | Handheld Shot (half right) |
| Handheld | nothing |
| Static | nothing — correct |
With five stills instead of three it was no better — and it invented a Zoom In on the pan clip. When we traced where that came from, it was not the picture at all: a text step had turned the word "close-up" in the description into a zoom. A close-up is a framing; a zoom is a movement. Nothing that reads words can tell them apart.
What did work was measuring instead of looking: comparing a run of small frames across the clip and tracking how the whole picture shifts or scales between them. That named the pan, the tilt and the zoom correctly, stayed silent on the static shot, and — just as importantly — stayed silent on the handheld and tracking clips it could not confidently call. A slight pan that moved only a few pixels per frame was still detected, because its direction was consistent, where handheld shake is not.
The rule we took from it
3. Use the vocabulary's names for motion, not the film-set ones
If you submit to Getty or iStock, movement terms go through the same controlled vocabulary as everything else — and its names are not the ones a camera operator uses. We checked each one against the vocabulary itself:
| What you would call it | What the vocabulary has |
|---|---|
| Pan | Panning |
| Tilt | Tilt Up / Tilt Down |
| Zoom | Zoom In / Zoom Out |
| Handheld | Handheld Shot |
| Gimbal / stabiliser | Stabilized Shot |
| Tracking / follow shot | Tracking Shot |
| Orbit | Orbiting |
| Time-lapse | Time Lapse |
| Real time / normal speed | Real Time Video |
| Slow-mo | Slow Motion |
| Drone shot | Drone Point of View |
| Dolly shot, push-in, locked-off | no entry |
Two things to take from the last row. First, when there is no entry, leave the movement out rather than forcing a neighbour: a dolly push-in is not a Zoom In, even though both make the subject bigger — one moves the camera, the other moves the lens, and a buyer who wants one does not want the other. Second, a static shot does not need a keyword; the absence of a movement term already says it.
Agencies that take free text — Adobe Stock, Shutterstock, Pond5 — have no such list, so there you write what a buyer types: "slow motion", "aerial view", "time lapse". Same clip, two spellings, which is the whole problem of keeping one list for several agencies (more on that below).
PixTagger keywords a clip from frames across it, measures camera movement instead of guessing it, and maps every term to the exact Getty vocabulary form — then writes the CSV for each agency.
Stock video keyword tool4. The same person should be the same age in every clip
A shoot often produces a dozen or more clips of one model. Keyword each clip on its own and the descriptions drift, because every clip is judged afresh from a different angle, light and expression.
We saw it on a real batch: 19 clips of one woman at a Christmas market were tagged across three different age brackets — One Mid Adult Woman Only on nine, One Mature Woman Only on four, One Young Woman Onlyon two — plus loose age terms on the rest. Every one of those is a real vocabulary entry, so nothing flags it. But to a buyer browsing the contributor's portfolio, one model has become three people.
The fix is to decide it once. Take the age bracket and the other people terms from the model release, apply them to every clip of that person, and only then keyword what differs from clip to clip: the action, the framing, the movement. Getty's combined terms make this easy to check — One Mid Adult Woman Only carries the count, the age and the gender in one entry, so a batch should show exactly one of them per model.
5. The date on a clip is the day you exported it
A photograph carries the moment it was taken in its EXIF data. A rendered video file usually does not: the timestamp it carries is when the file was written — which, for anything that went through an editor, is the day you exported it.
On that same Christmas batch, all 19 clips were dated almost a month after the shoot — the evening they were exported, minutes apart. A late-December market filmed and then dated late January is exactly the kind of mismatch that is easy to miss and awkward to explain, and agencies read the created date against the model release.
So for video, set the created date by hand from your shoot notes or the camera originals. The Getty ESP CSV guide covers the date format Getty expects and why a blank date beats a wrong one.
One clip, several agencies: the list has to change
Video contributors are more likely than photographers to spread one clip across several agencies, and the keyword rules pull in different directions:
- Getty / iStock— only controlled-vocabulary terms, in the vocabulary's exact form (Panning, One Mid Adult Woman Only).
- Adobe Stock — free text, up to 49 keywords but it prefers fewer, and the first ten carry the most weight.
- Pond5 — up to 50, and it asks for 40–50 on every item.
- Vecteezy — caps at 30.
No single list is right for all of them. Write the concepts once, order them strongest first, then let each export cut and reword: a Getty form like One Mid Adult Woman Onlymeans nothing typed into Adobe's search box, where a buyer writes "woman" and an age. We measured how far apart the two agencies' lists end up for the same picture in one photo, two keyword lists.
Titles follow the same logic. Pond5 wants 40–80 characters in the order subject, action, environment, and no number codes; Adobe wants no commas. The overlap that works everywhere is covered in how to write stock photo titles — it applies to clips unchanged.
A video keywording checklist
- Watch the start, middle and end — not the thumbnail. Drop anything true for only a moment.
- Fix the people terms once per model, from the release, and apply them to every clip of that person.
- Keyword the action and the framing clip by clip — this is where siblings should differ.
- Add camera movement only if you know it, in the vocabulary's form for Getty and in plain words elsewhere. Leave static shots without a movement term.
- State the speed when it is not real time: Slow Motion, Time Lapse.
- Set the created date by hand from the shoot, never from the file.
- Export per agency, letting each list be cut and reworded to that agency's rules.
Where PixTagger fits in
PixTagger keywords a clip from frames spread across it — three by default, at 10%, 50% and 90% of the running time so the head and tail padding are skipped, or five on request for clips where the action turns over. The model sees the frames together, so the list is written for the clip rather than for one thumbnail.
Camera movement follows the rule above. It is an optional measurement, taken from how the pixels move across frames in your own browser, on clips the browser can decode — and it is the only step allowed to add a movement term. The text steps that expand and fill a keyword list are barred from adding one, which is how the false Zoom Inabove was fixed. The date field is editable per clip before export, and the same batch exports to Getty, Adobe, Shutterstock, Pond5 and others, each cut to that agency's rules.
In short
Sources & further reading
- Shutterstock — Why was my content rejected for similar content?
- Pond5 — Preparing your footage files
- Pond5 — Master your metadata
- Shutterstock — Description and keyword best practices
- Our own work: every movement term above checked against Getty's controlled vocabulary (24,631 entries); the camera-movement test was six clips we shot for the purpose, one movement each; the age and date examples come from a real 19-clip contributor batch, described without identifying details.
Frequently asked questions
- How do you keyword stock video differently from photos?
- Every keyword has to be true for the whole clip, not just the thumbnail — check the start, middle and end. Add camera movement only when you know it, keep a model's age and people terms identical across every clip of that person, and set the created date by hand, because a rendered clip carries its export date.
- Should I add camera movement keywords like pan or zoom?
- Yes, when you know them — buyers filter for motion, and Shutterstock even treats different camera movements on the same subject as distinct clips. But movement cannot be read from still frames: in our test a vision model looking at three stills named none of a pan, a tilt or a zoom. State it from what you shot, or from a measurement across frames, never from a guess.
- What are Getty's keywords for camera movement?
- The controlled vocabulary uses its own forms: Panning (not Pan), Tilt Up and Tilt Down, Zoom In and Zoom Out, Handheld Shot, Stabilized Shot, Tracking Shot, Orbiting, Time Lapse, Real Time Video, Slow Motion and Drone Point of View. There is no entry for a dolly shot, a push-in or a locked-off static shot — leave those out rather than forcing a neighbour.
- Why is the date wrong on my stock video?
- A photo stores the moment it was taken in EXIF; a rendered video usually stores when the file was written, which is the day you exported it. In one 19-clip batch we checked, every clip was dated almost a month after the shoot. Set the created date by hand from your notes or the camera originals.
Written by a working stock contributor
NoSystem Images
Getty Images / iStock exclusive contributor since 2007
PixTagger is built by NoSystem Images, an exclusive Getty Images and iStock contributor since 2007, with a live portfolio of over 57,000 photos and 9,700 videos. Every keywording rule in the app comes from nearly two decades of actually selling on Getty, iStock and Adobe Stock — not from guesswork.
Related guides & tools
- Stock video metadata requirements for 9 agencies — codecs, frame rates, durations, caps and CSV columns
- How to keyword for Getty: controlled vocabulary — the forms you cannot guess, and why Women is not a head count
- Stock video keyword generator — keyword clips from frames across them and export per agency
Stop hand-keywording every upload
PixTagger writes buyer-focused titles, descriptions and marketplace-ready keywords for your photos and videos in seconds — with a Getty controlled-vocabulary CSV, an Adobe CSV, and qHero export built in.