video skill (coreyhaines31/marketingskills)
- Install
- SKILL.md (verbatim)
- Before Starting
- 1. Video Goal
- 2. Production Approach
- 3. Technical Context
- Choosing Your Approach
- Programmatic Video
- Hyperframes (HTML/CSS — recommended for agents)
- Remotion (React)
- When to Pick Which
- AI Video Generation
- Model Comparison
- Prompting for Video Models
- When to Use AI Generation vs. Stock
- AI Avatars
- HeyGen (recommended — has MCP server)
- Synthesia
- When to Use Avatars vs. Other Approaches
- Editing & Repurposing Tools
- Repurposing Workflow
- Reverse-Engineer a Viral Edit
- Video Production Workflows
- Product Demo Video
- Explainer Video
- Batch Social Clips
- Agent-Native Video Pipeline
- Common Mistakes
- Task-Specific Questions
- Tool Integrations
- Related Skills
- Other files in this skill
- references/ai-video-prompting.md (verbatim)
- Prompt Structure
- Example Prompts by Use Case
- Camera Movement Vocabulary
- Style Keywords
- Cinematic
- Commercial/Corporate
- Documentary
- Social/Trendy
- Model-Specific Tips
- Veo (Google)
- Runway Gen-4
- Kling
- Pika
- Common Prompt Mistakes
- Prompting Workflow
- Aspect Ratios
- Cost Optimization
- references/edit-anatomy.md (verbatim)
- When to use it
- Step 1 — Pull the reference so you can actually read the edit
- Step 2 — Extract the anatomy, beat by beat
- Step 3 — Write the beat sheet
- Step 4 — Review once, then execute
- Originality guardrail
- Common mistakes
What it does. When the user wants to create, generate, or produce video content using AI tools or programmatic frameworks. Also use when the user mentions 'video production,' 'AI video,' 'Remotion,' 'Hyperframes,' 'HeyGen,' 'Synthesia,' 'Veo,' 'Sora,' 'Runway,' 'Kling,' 'Seedance,' 'Hailuo,' 'MiniMax,' 'Pika,' 'Hunyuan,' 'Wan,' 'video generation,' 'AI avatar,' 'talking head video,' 'programmatic video,' 'video template,' 'explainer video,' 'product demo video,' 'video pipeline,' 'copy this edit,' 'match this video style,' 'reverse-engineer this video,' 'edit like this reference,' or 'make me a video.' Use this for video creation, generation, and production workflows. For video content strategy and what to post, see social. For paid video ad creative, see ad-creative. Part of coreyhaines31/marketingskills (marketing skills for agents) (coreyhaines31/marketingskills).
| Upstream | coreyhaines31/marketingskills |
| Skill file | skills/video/SKILL.md |
| License | MIT |
| Author | Corey Haines |
| Fetched | 2026-09-10 |
Install
npx skills add coreyhaines31/marketingskills --skill video, or copy the skill folder into~/.claude/skills/video/.- Raw file:
curl -sL https://raw.githubusercontent.com/coreyhaines31/marketingskills/HEAD/skills/video/SKILL.md
SKILL.md (verbatim)
name: video
description: "When the user wants to create, generate, or produce video content using AI tools or programmatic frameworks. Also use when the user mentions 'video production,' 'AI video,' 'Remotion,' 'Hyperframes,' 'HeyGen,' 'Synthesia,' 'Veo,' 'Sora,' 'Runway,' 'Kling,' 'Seedance,' 'Hailuo,' 'MiniMax,' 'Pika,' 'Hunyuan,' 'Wan,' 'video generation,' 'AI avatar,' 'talking head video,' 'programmatic video,' 'video template,' 'explainer video,' 'product demo video,' 'video pipeline,' 'copy this edit,' 'match this video style,' 'reverse-engineer this video,' 'edit like this reference,' or 'make me a video.' Use this for video creation, generation, and production workflows. For video content strategy and what to post, see social. For paid video ad creative, see ad-creative."
metadata:
version: 2.1.0
Video
You are an expert video producer who helps create marketing videos using AI generation models, AI avatars, and programmatic video frameworks. Your goal is to help users produce professional video content efficiently — from product demos and explainers to social clips and ads.
Before Starting
Check for product marketing context first:
If .agents/product-marketing.md exists (or .claude/product-marketing.md, or the legacy product-marketing-context.md filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Gather this context (ask if not provided):
1. Video Goal
- What type of video? (Product demo, explainer, testimonial, social clip, ad, tutorial)
- What's the target platform? (YouTube, TikTok/Reels/Shorts, website, ads, sales deck)
- What's the desired length?
2. Production Approach
- Do you need a human presenter? (AI avatar vs. voiceover vs. screen recording)
- Do you have existing footage or assets? (Screenshots, logos, product UI)
- Do you need generated footage? (AI-generated scenes, B-roll)
- Is this a one-off or a template for repeated use?
3. Technical Context
- What's your tech stack? (Node.js, Python, etc.)
- Do you have API keys for any video tools?
- Budget constraints? (Some tools charge per minute of video)
Choosing Your Approach
Pick the right tool for the job:
| Approach | Best For | Tools | When to Use |
|---|---|---|---|
| Programmatic | Templated, data-driven, batch video | Remotion, Hyperframes | Product updates, personalized videos, recurring content |
| AI Generation | Original footage from text/image prompts | Veo 3, Sora 2, Runway, Kling, Seedance | B-roll, hero shots, creative visuals you can't film |
| AI Avatars | Talking-head presenter without filming | HeyGen, Synthesia | Explainers, tutorials, multilingual content |
| Editing/Repurposing | Cutting long-form into short clips | Descript, Opus Clip, CapCut | Podcast/webinar → social clips |
Programmatic Video
Build videos with code. Best for repeatable, templated, or data-driven video at scale.
Hyperframes (HTML/CSS — recommended for agents)
Open-source, Apache 2.0, from HeyGen. Uses plain HTML/CSS/JS — no framework DSL to learn. LLM-native: AI models generate better HTML than React components.
npm install hyperframes
Key concept: Each frame is an HTML document. Compose frames into a timeline, render to MP4.
import { render } from "hyperframes";
await render({
frames: [
{ html: "<h1>Welcome to Acme</h1>", duration: 3 },
{ html: "<h2>Here's what we built</h2>", duration: 3 },
{ html: "<p>Try it free →</p>", duration: 2 },
],
output: "intro.mp4",
width: 1080,
height: 1920, // 9:16 for vertical
});
Best for: Product announcements, changelogs, data-driven reports, personalized outreach videos.
Why agents prefer it: Plain HTML/CSS means any coding agent can generate frames without learning a framework. Deterministic rendering — same input always produces identical output.
Remotion (React)
Mature open-source framework. More powerful than Hyperframes but requires React knowledge.
npx create-video@latest
Key concept: React components are frames. Props drive content. Render locally or via Remotion Lambda (AWS) for scale.
export const ProductDemo: React.FC<{ title: string; features: string[] }> = ({
title, features
}) => {
const frame = useCurrentFrame();
return (
<AbsoluteFill style={{ background: "#000", color: "#fff" }}>
<h1>{title}</h1>
{features.map((f, i) => (
<Sequence from={i * 30} key={i}>
<p>{f}</p>
</Sequence>
))}
</AbsoluteFill>
);
};
Best for: Complex animations, interactive previews, large-scale batch rendering (Lambda).
When to Pick Which
| Factor | Hyperframes | Remotion |
|---|---|---|
| Agent compatibility | Better (plain HTML) | Good (React) |
| Animation complexity | Basic (CSS transitions) | Advanced (Spring, interpolate) |
| Batch rendering | Local | Lambda (AWS) for scale |
| Learning curve | Minimal | Moderate (React + Remotion API) |
| License | Apache 2.0 | Company license for commercial use |
AI Video Generation
Generate original footage from text or image prompts. Use for B-roll, hero visuals, and scenes you can't practically film.
Model Comparison
| Model | Resolution | Max Duration | Best For | Cost |
|---|---|---|---|---|
| Veo 3 (Google) | Up to 1080p (4K varies) | Variable | Top overall quality, synced audio | API-based |
| Sora 2 (OpenAI) | Up to 1080p | Up to ~20 sec | Cinematic + synced audio, ChatGPT/API integration | API + ChatGPT |
| Runway Gen-4 | Up to 4K | ~10 sec/gen | Motion control, temporal consistency, edit-style workflows | $12-76/mo |
| Kling 2.5/3.0 (Kuaishou) | Up to 1080p | Up to 2 min | Long-take generation, lower per-second cost | ~$0.03/sec |
| Seedance (ByteDance) | Up to 1080p | Short clips | Fast generation, strong motion fidelity at low cost, batch-friendly | Per-credit |
| Hailuo / MiniMax | Up to 1080p | Short clips | Character consistency across shots | Per-credit |
| Pika 2.x | 1080p | Short clips | Quick effects, image-to-video, lower bar to entry | Per-credit |
| Hunyuan Video / Wan 2 | 720p–1080p | Variable | Open-source self-hosted; full control, no API fees | Free (GPU) |
Quick picks:
- Highest quality + audio: Veo 3 or Sora 2
- Batch / volume / cost: Kling, Seedance
- Character consistency across multiple shots: Hailuo
- Self-hosted, brand-controlled: Hunyuan Video or Wan 2 (open weights)
- Storyboard → video workflow: Runway, LTX Studio
- Image-to-video from a still you already have: Kling, Pika, Runway
Prompting for Video Models
Good video prompts specify: subject + action + camera + style + mood
A close-up shot of hands typing on a laptop keyboard,
shallow depth of field, warm office lighting,
camera slowly pulls back to reveal a modern workspace,
cinematic color grading, 4K
Common mistakes:
- Too vague ("a person working") — add specifics
- Ignoring camera movement — specify dolly, pan, static
- Forgetting style — "cinematic," "documentary," "commercial"
- Requesting text in video — AI models struggle with readable text
For detailed prompting guides: See references/ai-video-prompting.md
When to Use AI Generation vs. Stock
| Use Case | AI Generation | Stock Footage |
|---|---|---|
| Exact scene you imagined | Yes | Rarely matches |
| Consistent style across clips | Yes | Hard to match |
| Recognizable real locations | No (hallucinations) | Yes |
| Specific products/brands | No (use programmatic) | No |
| Quick B-roll | Either works | Faster |
AI Avatars
Create talking-head videos without filming. An AI avatar delivers your script with realistic lip-sync, expressions, and gestures.
HeyGen (recommended — has MCP server)
Best lip-sync and micro-expressions. 230+ avatars, 140+ languages.
Agent integration: HeyGen has an official MCP server — AI agents can generate avatar videos directly.
| Plan | Videos | Duration |
|---|---|---|
| Free | 3/mo | 3 min max |
| Creator | Unlimited | 5 min |
| Business | Unlimited | 20 min |
Check heygen.com/pricing for current prices.
Best for: Product explainers, feature announcements, personalized sales outreach, multilingual content.
Custom avatars: Upload a 2-5 min video of yourself to create a digital twin. Looks and sounds like you, generates videos from text scripts.
Synthesia
Full-body avatars with expressive body language. Built-in script generation from URLs/docs.
Best for: Corporate training, compliance videos, enterprise presentations where professional tone > realism.
When to Use Avatars vs. Other Approaches
| Scenario | Use Avatar | Use Instead |
|---|---|---|
| Recurring content (weekly updates) | Yes | — |
| Multilingual versions | Yes | — |
| Personalized outreach at scale | Yes | — |
| Authentic founder content | No | Film yourself |
| Product UI walkthrough | No | Screen recording |
| Creative/artistic video | No | AI generation |
Editing & Repurposing Tools
Turn existing content into multiple video formats.
| Tool | What It Does | Best For |
|---|---|---|
| Descript | Transcript-based editing — edit video by editing text | Cleaning up interviews, podcasts, webinars |
| Opus Clip | Auto-clips long videos, scores virality potential | Long-form → short-form at scale |
| CapCut | Visual effects, captions, platform-native styling | TikTok/Reels polish |
| Captions.ai | Auto-captions, eye contact correction, AI dubbing | Solo talking-head content |
Repurposing Workflow
Long-form content (podcast, webinar, demo)
↓
Descript: Clean up, remove filler, polish
↓
Opus Clip: Auto-extract 5-10 best moments
↓
CapCut: Add captions, effects, platform styling
↓
Distribute: TikTok, Reels, Shorts, LinkedIn
Reverse-Engineer a Viral Edit
To replicate the style of a video edit you admire — the cut rhythm, caption treatment, punch-ins, on-screen text, sound design — decompose it into a reusable edit spec (a beat sheet) and apply it to your own footage. Pull the reference with watch-video (visual/multimodal mode extracts frames at the cut points) or social-fetch, extract the edit anatomy beat by beat, and output a per-beat table plus the 3–5 signature moves that make the edit recognizable. Review the beat sheet once before executing it (in Remotion/Hyperframes, CapCut, or an AI restyle tool). Copies the editing grammar, never the reference's footage/script/music. Full method: references/edit-anatomy.md.
Video Production Workflows
Product Demo Video
- Script the key features and value props (use copywriting skill)
- Screen record the product flow
- Programmatic overlay — use Hyperframes/Remotion for titles, callouts, transitions
- AI B-roll — generate establishing shots or lifestyle scenes with Veo/Runway
- Voiceover — record yourself or use AI avatar for narration
- Export at platform-appropriate specs
Explainer Video
- Script the problem → solution → CTA arc
- Choose presenter — AI avatar (HeyGen) or voiceover + visuals
- Build visuals — programmatic slides, screen recordings, AI-generated scenes
- Add captions — always, for accessibility and engagement
- Export — landscape for YouTube/website, vertical for social
Batch Social Clips
- Create master template in Hyperframes/Remotion
- Feed data — product features, testimonials, stats
- Render batch — one template, many variations
- Add platform-specific captions via CapCut or Captions.ai
- Schedule across platforms
Agent-Native Video Pipeline
The most powerful setup combines tools that agents can control directly:
Agent writes script (from product context)
↓
Hyperframes: Generate templated video (HTML → MP4)
and/or
HeyGen MCP: Generate avatar video from script
and/or
Veo/Runway API: Generate B-roll footage
↓
Agent assembles final cut
↓
Output: Ready-to-publish video
What makes this agent-native:
- Hyperframes uses HTML — any coding agent can generate it
- HeyGen MCP server — agents call it directly
- Video model APIs — standard HTTP requests
- No manual editing step required
Common Mistakes
- Starting with tools, not strategy — decide what video you need before picking tools
- AI-generated text in video — models can't reliably render readable text; use programmatic overlays instead
- Uncanny valley avatars — if avatar quality matters, invest in HeyGen Creator+ tier
- No captions — 85% of social video is watched without sound
- Wrong aspect ratio — 9:16 for social, 16:9 for YouTube/website, 1:1 for feeds
- Over-producing — authentic often outperforms polished, especially on TikTok
Task-Specific Questions
- What type of video do you need? (Demo, explainer, social clip, ad, tutorial)
- Do you need a human presenter or can it be voiceover/text?
- Is this a one-off or a repeatable template?
- What platform is it for? (This determines aspect ratio and length)
- Do you have existing assets to work with? (Screenshots, footage, scripts)
- What's your budget for video tools?
Tool Integrations
| Tool | Type | MCP | Guide |
|---|---|---|---|
| HeyGen | AI avatars | Yes | heygen.md |
| Hyperframes | Programmatic video | - | hyperframes.md |
| Remotion | Programmatic video | - | remotion.dev |
| Runway | AI generation | - | runwayml.com/docs |
Related Skills
- social: For video content strategy, hooks, and what to post
- ad-creative: For paid video ad creative and iteration
- copywriting: For video scripts and messaging
- marketing-psychology: For hooks and persuasion in video
Other files in this skill
references/ai-video-prompting.md (verbatim)
AI Video Prompting Guide
How to write effective prompts for AI video generation models (Veo, Runway, Kling, Pika).
Prompt Structure
A strong video prompt follows this formula:
[Subject] + [Action] + [Camera movement] + [Visual style] + [Lighting/mood] + [Technical specs]
Example Prompts by Use Case
Product hero shot:
A sleek laptop on a minimal white desk, screen glowing with a dashboard UI,
camera slowly orbits 180 degrees around the desk,
soft volumetric lighting from the left, shallow depth of field,
cinematic commercial aesthetic, 4K
Lifestyle B-roll:
A woman in a modern co-working space smiling while looking at her phone,
natural window light, candid documentary feel,
camera handheld with subtle movement, warm color grading
Abstract/brand:
Flowing liquid gold particles forming the shape of a network graph,
dark background, particles catch light as they move,
slow-motion macro photography style, dramatic rim lighting
SaaS explainer scene:
An overhead shot of a team around a conference table pointing at charts,
camera slowly pushes in, bright modern office,
clean corporate style, even lighting, 1080p
Camera Movement Vocabulary
Use these terms — video models understand them:
| Term | Effect |
|---|---|
| Static | Locked camera, no movement |
| Pan left/right | Camera rotates horizontally |
| Tilt up/down | Camera rotates vertically |
| Dolly in/out | Camera moves toward/away from subject |
| Orbit | Camera circles around subject |
| Tracking shot | Camera follows moving subject |
| Crane/aerial | Camera rises or descends |
| Handheld | Subtle shake, documentary feel |
| Zoom | Lens zoom (different from dolly) |
| Slow push | Gradual dolly in — builds tension/focus |
Style Keywords
Cinematic
- "cinematic color grading"
- "anamorphic lens flare"
- "shallow depth of field"
- "film grain"
- "35mm film"
Commercial/Corporate
- "clean commercial lighting"
- "bright and airy"
- "professional corporate aesthetic"
- "even, diffused lighting"
Documentary
- "handheld documentary style"
- "natural lighting"
- "candid, unposed"
- "observational camera"
Social/Trendy
- "vertical 9:16"
- "fast-paced cuts"
- "bold text overlays"
- "high contrast, saturated colors"
Model-Specific Tips
Veo (Google)
- Excels at photorealism and complex scenes
- Supports audio generation synced to video
- Best with detailed, descriptive prompts
- Specify "high resolution" or "1080p" for best quality
- Can handle multiple subjects and scene transitions
Runway Gen-4
- Strong motion control — specify camera movements precisely
- Best temporal consistency (subjects stay consistent across frames)
- Use motion brush for specific area animation
- Image-to-video works well — provide a reference frame
- Keep prompts under 100 words for best results
Kling
- Can generate up to 2 minutes (much longer than others)
- Good for longer narrative sequences
- More affordable for bulk generation
- Quality drops slightly at longer durations
- Best with simpler scenes and fewer subjects
Pika
- Fastest generation time (under 2 minutes)
- Good for quick iterations and experimentation
- Effects mode adds motion to still images
- Best for short clips (5-15 seconds)
- Less control over camera movement
Common Prompt Mistakes
| Mistake | Why It Fails | Fix |
|---|---|---|
| "A person using our app" | Too vague, no visual detail | Describe the person, setting, lighting, camera |
| Including text/logos | AI can't render readable text | Add text in post via Hyperframes/CapCut |
| "Make it viral" | Not a visual instruction | Describe the visual style you want |
| Extremely long prompts (200+ words) | Models lose focus | Keep to 50-100 words, be specific |
| No camera direction | Random/static camera | Always specify movement or "static" |
| "Realistic" alone | Not specific enough | "Photorealistic, natural lighting, shot on RED camera" |
Prompting Workflow
- Reference first — find a real video that looks like what you want
- Describe it — break down: subject, action, camera, style, mood
- Generate 3-4 variations — same concept, different angles or styles
- Iterate on the best — refine the prompt based on results
- Composite — combine AI footage with programmatic text/overlays
Aspect Ratios
Always specify in your prompt or generation settings:
| Platform | Ratio | Resolution |
|---|---|---|
| YouTube | 16:9 | 1920x1080 or 3840x2160 |
| TikTok/Reels/Shorts | 9:16 | 1080x1920 |
| Instagram Feed | 1:1 or 4:5 | 1080x1080 or 1080x1350 |
| Website hero | 16:9 | 1920x1080 |
| 16:9 or 1:1 | 1920x1080 |
Cost Optimization
- Iterate at low resolution — upscale only the final version
- Use Kling for drafts — cheapest per second, switch to Veo/Runway for finals
- Image-to-video — providing a reference frame saves generation credits and gives better results
- Batch similar prompts — models often offer volume discounts
- Cache and reuse — B-roll clips can be reused across multiple videos
references/edit-anatomy.md (verbatim)
Reverse-Engineering an Edit (The Beat Sheet)
A viral short-form video usually isn't winning on the footage — it's winning on the edit: the cut rhythm, the caption style, the punch-ins, the on-screen text landing on the exact word, the b-roll cutaways, the sound design. This reference turns a reference edit you admire into a reusable edit spec — a beat sheet you (or an editing tool) can execute against your own footage — without copying a single frame of theirs.
This is the tool-agnostic half of "copy any viral edit": the decomposition. The generation is whatever you edit with afterward — CapCut, Premiere, Remotion/Hyperframes, or an AI restyle tool. The spec is the deliverable.
When to use it
- A competitor's or creator's edit keeps stopping your scroll and you want to understand why and replicate the technique
- You have raw footage (a talking-head clip, a demo) and a reference edit whose style you want to match
- You're briefing an editor or a template and need the edit decisions written down, not vibes
Don't use it to copy someone's actual creative — this extracts the editing grammar (structure, rhythm, caption treatment), not the script, footage, or brand. Same rule as mining organic content for vocabulary in the hook system: take the technique, never the creative.
Step 1 — Pull the reference so you can actually read the edit
You cannot decompose an edit from a description of it. Get the frames and the timing:
- watch-video (visual or multimodal mode) — extracts the transcript and samples frames at the cut points, so you can read on-screen text, caption style, and shot changes. This is the primary tool.
- social-fetch — pull the post for the caption, engagement, and the media URL when the reference is a specific tweet/Reel/TikTok.
- Screenshots of key frames also work if the user supplies them — you need the visual, not just the words.
Note the total duration and roughly how many cuts there are before you start — cuts-per-second is the single most telling number about an edit's energy.
Step 2 — Extract the anatomy, beat by beat
Walk the reference from 0:00 and log every editing decision. The dimensions that define a short-form edit:
| Dimension | What to read off the reference |
|---|---|
| Shot & framing | Talking head / screen recording / b-roll / text card; close-up vs. wide; headroom, rule-of-thirds, or dead-center |
| Cut rhythm | Where each cut lands and how fast (cuts-per-second); is it on the beat, on the word, or on the breath? |
| On-screen text | The words, when each appears/disappears, and where on the frame (top-third caption vs. big centered statement) |
| Caption style | Font, weight, color, outline/box, and animation (word-by-word pop, karaoke highlight, whole-line) |
| Motion | Punch-ins / zoom pushes, shakes, whip-transitions, speed ramps — where and how aggressive |
| B-roll & overlays | Cutaways, stickers, arrows, emoji, screenshots, meme inserts — what's laid over the base footage and when |
| Sound design | Music choice and where it hits, SFX (whooshes, dings, risers), and deliberate silence before a beat |
| Hook (first 2s) | The single most-copied element — what's on screen and said in the opening two seconds, before anyone's committed |
| Pacing curve | Does it stay frantic, or fast-hook → slower-body → fast-CTA? Map the energy over the runtime |
Read the pattern, not just the instances: "a hard cut + punch-in on every new sentence," "caption is one word at a time, yellow, karaoke-highlighted, bottom third," "a whoosh SFX on every scene change." Patterns are what make an edit replicable; a list of 40 individual cuts is not.
Step 3 — Write the beat sheet
Two artifacts: a per-beat table and a short style summary.
The beat sheet — one row per beat (a beat = a cut or a distinct edit event):
| Beat | Time | Shot | On-screen text | Caption style | Transition / motion | Audio |
|------|-----------|-----------------|-----------------------|----------------------|-----------------------|------------------|
| 1 | 0:00–0:02 | CU talking head | "STOP doing this" | word-pop, yellow, ctr| hard in, slow push | music in + riser |
| 2 | 0:02–0:04 | screen record | (caption only) | karaoke, white, btm | hard cut + whoosh | click SFX |
| … | | | | | | |
The style summary — the 3–5 signature moves that make this edit recognizable, stated so they're reusable:
- e.g. "Every sentence gets a hard cut + a 5% punch-in." / "Captions are one word at a time, bottom-third, karaoke-highlighted." / "A whoosh SFX on every cut; music drops out for 0.5s before the CTA." / "The hook is a bold centered statement on frame 1, no logo."
The signature moves are the real deliverable — someone can apply those five rules to any footage and get the style. The table is the detailed backup.
Step 4 — Review once, then execute
Show the beat sheet before anyone edits anything — the same review-once gate as the ad-creative creative review page. The reviewer checks two things:
- The on-screen text says what you want (mapped to your message, not the reference's)
- The scene changes land where you want them (your footage's beats, not a blind copy of the reference's timing)
Approve, then execute the spec with your footage:
- Remotion / Hyperframes — when you want the edit templated and data-driven (see the programmatic-video section in SKILL.md); the beat sheet is the composition spec.
- CapCut / Premiere / an editor — hand off the beat sheet + style summary as the brief.
- An AI restyle tool — feed the style summary as the target style.
Originality guardrail
You are copying the edit, not the content. The beat sheet describes technique (cut rhythm, caption treatment, motion, sound design) applied to your footage and your message. General editing techniques and style cues are usually reusable — U.S. copyright protects expression, not procedures or methods (17 U.S.C. §102(b)) — but the reference's specific creative expression is not, and closely reproducing a finished video's exact selection and arrangement of choices can still create risk. So copy the grammar, not the finished work: use your own footage, message, script, voiceover, licensed music/SFX/samples, and brand elements. If the reference's "style" is really a specific bit or sketch, that's their creative — draw inspiration, don't reproduce it.
Common mistakes
- Describing instead of reading — you can't extract caption style or cut timing from the transcript alone; pull the frames (watch-video).
- Logging instances, not patterns — 40 cut timestamps isn't a spec; "hard cut + punch-in per sentence" is.
- Copying the reference's timing onto different footage — beats land on your words and your cuts; the reference gives you the grammar, not the calendar.
- Skipping the hook — the first 2 seconds carry most of the retention; decode them in the most detail.
- Reproducing the creative — matching the edit is fine; re-shooting their exact bit, script, or using their footage/music/SFX is not.
Back to coreyhaines31/marketingskills (marketing skills for agents) or Agent skills.