Learning how to make instructional videos starts with the action they must perform, show the necessary steps accurately, and test the finished lesson with someone who has not seen the process. Recording and editing come after that decision. A polished video can still fail if the learner cannot tell what to do next.
This guide is for teams creating customer education, product onboarding, and professional training. It covers choosing a format, preparing source material, scripting, recording, editing, and checking comprehension. We also inspect a real TapVid photography lesson: its original request, actual export, useful visual decisions, and gaps that should be fixed before treating it as a finished course.
01
1. Define an action and a visible success condition
Replace “teach the product” with a sentence that can be tested: “After this lesson, a new teammate can prepare a product-video brief that includes the correct asset, approved wording, and intended output.” The learner, starting state, action, and result are now explicit. Do not add a second task unless it is necessary to finish the first.
Write down prerequisites beside the objective. A software tutorial may require an account role and a sample project. A physical demonstration may require particular equipment. A conceptual lesson may require vocabulary. Show these before the first action; discovering a missing prerequisite halfway through a video is not a pacing problem.
For our photography case, “understand a 35mm lens” is too broad. A better assessment is to ask the learner to compare two portraits and explain whether the background supports the subject. That is an example assessment we propose here; we did not run a learner study or establish an improvement in photographic skill.
02
2. Choose the medium from what needs to be shown
| Task | Use | Keep visible |
|---|---|---|
| Operate software | Real screen recording with narration | Starting screen, control being used, resulting state |
| Handle or assemble a product | Camera footage of the actual object | Hands, orientation, connection points, finished result |
| Explain a relationship | Diagram or animated explainer | Labels, changing variable, relationship being explained |
| Look up a short changing procedure | Written steps, optionally supported by video | Searchable terms, current version, exception path |
A diagram can explain why something happens without proving that a product performed it. A mock interface can illustrate a concept without teaching the location of a real button. Label these choices honestly. If the learner must reproduce clicks, keep the real interface available; if the lesson explains a process inside an object, a labeled diagram may be clearer than footage alone.
Keep a short written version even when video is the primary format. It gives the learner a way to search for a term, check a prerequisite, and revisit a step without replaying the introduction. This is especially useful when the procedure changes more often than the concept behind it.
Scribe: instructional-video workflow, maintenance and written alternatives
03
3. How to make instructional videos from a complete source packet
Gather the current procedure, approved wording, real screenshots or footage, and one example of a correct result. Add the most likely failure: missing permission, a wrong file type, an incorrectly connected part, or a misunderstood term. Ask a subject expert to verify factual statements before they become narration. An AI-generated explanation is not a substitute for this source review.
Write the visual action alongside the narration. In a repeatable tutorial, each meaningful action needs an observable result. “Configure the project” hides several decisions; “choose the required ratio, then confirm the preview uses that layout” shows the action and its check. If a step takes time, record the wait and decide whether to shorten it with an explicit indication that time has passed.
| Proposed training beat | Visual | Narration / check |
|---|---|---|
| Identify the source | Show the approved image beside its filename | “Use the approved product image.” Verify the SKU or version. |
| Lock the wording | Highlight the exact sentence in the brief | “Copy the approved sentence without adding a claim.” Compare it character by character. |
| Describe the destination | Show the requested layout and intended placement | “Choose the layout for where the video will be viewed.” Check small-screen readability. |
| Verify the brief | Show the three completed items together | “Before production, confirm image, wording, and destination.” Let the learner find an intentionally missing item. |
The table is a proposed instructional script, not a recording of a customer workflow or a measured training result.
04
4. Record clean actions and understandable audio
For screen capture, use a demonstration account with realistic sample data, clear irrelevant notifications, and position the important interface region at a readable size. Rehearse the complete task before recording. Move the pointer to a control, pause long enough to identify it, perform the action, and keep the resulting state on screen. Do not drag the pointer around while explaining an unrelated point.
For camera footage, keep the object, hands, and result visible. Check framing before a take rather than discovering that a connector was outside the shot. Record a short audio sample and listen for fan noise, clipping, and words that disappear under music. Clear narration matters more than a visually elaborate introduction.
If a step fails, pause and repeat that step from a known starting state. You can record narration separately when performing the action and explaining it together causes mistakes. In editing, match the explanation to the moment the learner sees the action; do not leave the important screen behind while the voice is still describing it.
TechSmith: practical recording, narration and editing guidance
05
5. A real TapVid lesson: useful visuals, unfinished instruction
We inspected “Mastering Portrait Photography with a 35mm Lens,” an existing TapVid case-library project. Its source request asks for instruction on 35mm portrait photography and example photos. The project records automatic style and structure, a female voiceover, then approval of an outline and script before generation. This is a review of an existing production, not a claim that we generated it today or that a customer achieved a particular result.
Watch the original TapVid photography lesson
The opening contrasts an isolated portrait with a wider environmental view. The useful decision is to make “include meaningful surroundings” visible instead of describing it only in narration. Later, photographic examples receive composition overlays. These cues give the eye something specific to inspect, although a learner still needs time to apply the idea to an unmarked photo.
Around the near/far diagram sequence, a wireframe face changes from a warning treatment to a calmer state. This is a useful illustration of a spatial relationship. It is not an optical measurement. The lesson also gives a fixed distance rule; the materials we inspected do not establish that number as universally correct. We would replace an absolute rule with subject-expert-reviewed wording and conditions before using this as technical training.
The export is 142.73 seconds at 1920 × 1080, measured from the original MP4. That is output duration, not generation time. Original credit cost, exact model, and learner outcomes were not available in the inspected record. A playable file establishes that a lesson was produced; it does not establish that someone learned the skill.
To finish the instructional job, we would add three things: a specific learning objective at the start, a pause for the learner to choose between two unannotated examples, and feedback explaining the choice. These are proposed revisions, not changes already made to this library video. The distinction matters: an attractive explanation can be a useful component of training without being a complete lesson.
06
6. Edit for the learner’s next action
Cut repetition and irrelevant motion first. Keep cues next to the object they explain. Add chapter names that describe actions rather than “Part 1” and “Part 2.” Where a screen is dense, emphasize the relevant area while keeping enough context to locate it. Give captions enough room; check that they do not cover the control or measurement being explained.
Read every caption and on-screen label against the source. Play once with sound and once without it. The silent pass reveals whether the action can be located; it does not replace the spoken explanation or an accessible transcript. Watch the final export, not only the editor preview, because cropping, subtitle timing, and image loading can differ.
07
7. Test with a fresh viewer, then publish where the task happens
Give the viewer the same starting conditions named in the lesson. Ask them to perform the task without coaching. Note the step where they hesitate, the instruction they misread, and whether the final state is correct. Do not ask only whether they “liked the video.” A friendly rating cannot tell you which missing step caused the failure.
Use each observed failure to choose the smallest repair. Missing context means adding a prerequisite. A hidden control means a better crop or cue. Wrong sequencing means rerecording the transition. An unsupported explanation means correcting the source before re-rendering. Test again after that change; adding more background information to every failure usually makes the lesson harder to navigate.
Place the lesson beside the task: in onboarding, a help article, or the relevant training module. Include the written steps, transcript, prerequisites, version date, and a named owner. When the interface or process changes, update the affected lesson and verify that links still open the current version. Monitor repeated support questions alongside completion; neither metric alone proves learning.
08
Questions that affect production
How long should it be? Long enough to show the complete task and its result, with no unrelated steps. Split genuinely separate tasks into chapters or separate lessons. A universal duration target can make a simple lesson bloated and a difficult one incomplete.
Do you need to appear on camera? Only when a person, physical action, or interpersonal context helps teach the task. Clear screen capture and narration may be enough. Can AI help? Yes, with explanatory assembly and visual organization, provided the source, labels, and final result are reviewed. TapVid is an Explainer Video Engine; our inspected example demonstrates an explanatory treatment, not automatic subject-matter verification.
What should you make first? Choose one recurring question that currently requires a teammate to explain the same process repeatedly. Document the correct result, produce the shortest complete lesson, and check whether a new viewer can reproduce it. That gives the next revision a concrete purpose.




