How To Make Shorts From Long Videos With Claude And DaVinci Resolve
How one screen and camera recording becomes ten vertical shorts. Layer 1 is the cut in Tella by transcript. Layer 2 is the finished reel rebuilt inside DaVinci Resolve by your agent, with the hook and captions as editable Text+ nodes. Every setting below came from a real run tonight.
I record once. One take, my screen and my face, forty to sixty minutes, no script. That recording became ten shorts this week and I did not open a timeline by hand for any of them.
The trick is not the AI. It is the order. Cut by transcript first, because words are the edit. Then lay the picture out so the hook and the captions live in the gap between the screen and my head, never on my face and never on the screen. Then add the graphics last, in a place where you can change them.
That last part is what most AI editing skips. A baked-in title is a title you cannot fix on Sunday night when a word is wrong. A professional editor I follow put it plainly: he uses AI to feed him ideas and write his scripts, then builds the graphics inside Resolve where a change is thirty seconds, not a new render. That is the version below.
I also went through the four other setups people are publishing for Claude inside Resolve: Matt Penny's template-driven pre-editor, Andy Diep's marker-and-approve workflow, and Jason Cooperson's seven stage folder and his Resolve MCP build. Every one of them lands on the same three truths. The folder is the system. Templates are what make the output reliable. The human keeps the judgment. Those are built into the version below, with credit where the idea came from.
Two layers. Layer 1 needs Tella and a chat window. Layer 2 needs DaVinci Resolve, the free one is enough, and an agent that can talk to it.
The three levels
You scrub through the recording, guess where the good part starts, cut it by eye, and add a title in CapCut. One short takes an evening.
Tella cuts by transcript from a list of word ranges. A script stacks the screen over the head with a clean band between them. The hook holds the band for three seconds, then the captions take it. Ten shorts in an afternoon.
Your agent reads the transcript, proposes the cuts, builds them in Tella, rebuilds each one in Resolve with editable hook and captions, renders, and hands you a review page. You read transcripts, not timelines.
The mental model
Scrub the video, cut by eye, drop a title on top, export, repeat.
Words are the edit. Cut by transcript, lay out the picture around a band that belongs to the text, add graphics where you can still change them.
| Role | Talks to you | Job |
|---|---|---|
| Transcript | First | Decides what the shorts are. Word ranges, not timestamps |
| Tella | Cut | Cuts by word index and exports camera and screen as separate tracks |
| The band | Layout | Screen on top, head below, a 360 pixel gap the text owns |
| Resolve | Finish | Hook and captions as Text+ nodes. Editable, then rendered |
Layer 1: the cut and the stack
1. Record once, both tracks
Screen plus camera in Tella, one take, talk through the thing you actually did. Do not edit while recording. The transcript will find the shorts.
2. Pick the moments from the transcript
Read the transcript, not the video. A short is a stretch of 30 to 65 seconds with one idea, a reason to keep watching in the first line, and a line that ends it. Mark each one as a start word and an end word. Ten candidates from an hour is normal. Cut in Tella by word index: duplicate the recording, keep only the ranges, and let Tella trim the silences to about a fifth of a second. Word to word cuts leave dead air at the seams, so trim silences after the cut, not before.
3. Export the tracks, not the flattened video
Every Tella layout puts the head right under the screen with no gap, so a title lands on the screen and captions land on the face. Export the camera and the screen as separate tracks and build the picture yourself.
4. Stack with the band
Canvas 1080 by 1920. Screen scaled to 1080 by 608 at the top, 40 pixels down. Camera cropped around the head, scaled to 1080 by 911, bottom edge on the floor. That leaves a band from 648 to 1009 that nothing sits in except text. The factory script does this with ffmpeg in one command.
5. Hook, then captions, in the band
Two lines, all caps, Anton, white over red, black outline, fitted to 995 pixels wide. The hook owns the band for the first three seconds and fades. Only then do the captions appear, three words at a time, directly above the hat. Name the tool in the hook. Every hook gets scored before it renders.
6. Read every short before it posts
A review page with the transcript of each finished cut and the length of every pause at every seam. If it stumbles when you read it, it stumbles on Instagram. Then splice a spoken call to action onto any cut that lost it.
Starter prompts
Paste these as written. They are short on purpose, because the long ones drift.
You are cutting a screen and camera recording into vertical shorts. Here is the transcript with word indices: [PASTE]. Find every stretch of 30 to 65 seconds that carries one idea, opens with a line that earns the next second, and ends on a line that closes it. For each, give me: the start word and end word with their indices, a working title, and the one tool or thing being named. Then write two on-screen hook lines per short, all caps, 2 to 4 words each, line two names the tool, and score each against: a number, a contrarian claim, direct YOU. Keep only hooks scoring 2 of 3. Flag any short that does not contain a spoken call to action so I can splice one. Do not invent lines that are not in the transcript.
DaVinci Resolve is open with CursorBridge listening on 127.0.0.1:9876. Build a 1080 by 1920, 30 fps project from camera.mp4 and screen.mp4 in [FOLDER]. Insert the Fusion hook clip first on V1 at 01:00:00:00, then screen video on V2, camera video on V3 and camera audio on A1, all at record frame 108000. Screen: Tilt 1949. Camera: Zoom 1.5, crop a 720 fitted pixel window around the head found from an exported neutral frame, Pan to centre it, Tilt negative 1595. Hook lines: [LINE ONE] over [LINE TWO], Anton, white over red, black outline, fade at frame 100. Then compound screen, camera and audio and import a captions comp from [WORDS JSON] onto it. Export uniquely named check frames at 1s, 6s and 20s and show me before rendering. Never overwrite the project called [NAME].
Here are the transcripts of ten finished cuts with the pause length at every seam: [PASTE]. Read each one aloud in your head. Mark Clean, Watch or Fix. For every Watch or Fix, quote the exact words on either side of the seam and say what a listener would hear. Do not suggest re-recording.
Open the Resolve project [NAME]. Walk the red seam markers with me one at a time: for each, tell me the words on either side of the join and the pause length. Then for each graphic on V1 apply my notes: [NOTES]. Re-render only the ranges you changed, named by range, and show me a frame from each before and after.
Layer 2: the same reel, rebuilt inside DaVinci Resolve
Layer 1 gives you a finished MP4. Layer 2 gives you a timeline: the same layout, but the hook and the captions are Text+ nodes you can open and change. Your agent drives Resolve through the CursorBridge script (Workspace, Scripts, CursorBridge), which listens on localhost and works on the free version with no external scripting. These are the settings that had to be measured because the Inspector's units are not what they look like on a portrait timeline.
The folder is the system
Your agent opens one folder: the style card, the hook and caption templates, the two scripts, and the word file. Everything it needs to be consistent lives there, so a new session starts with the same rules as the last one. Matt Penny and Jason Cooperson both build their editors this way, and it is why their outputs repeat. Do not prompt the look. Point at the template.
Project
New project, timelineResolutionWidth 1080, timelineResolutionHeight 1920, timelineFrameRate 30. Set these before creating the timeline. Import camera.mp4 and screen.mp4. Timelines start at 01:00:00:00, which is record frame 108000 at 30 fps. Insert at 0 and the clips land before the timeline starts and the viewer is black.
Order of insertion
Insert the Fusion composition for the hook first, on V1, at the playhead. Inserting a Fusion composition ripples every track, so anything already on the timeline moves five seconds. Then insert screen video only on V2, camera video only on V3, and camera audio only on A1. If you let the screen's audio onto the timeline your render is silent under the real voice.
Screen transform
A 16 by 9 clip fits to 1080 by 608, centred. Tilt moves it up. On a 1080 by 1920 timeline one Tilt unit is 0.316 pixels, which is (1080 divided by 1920) squared. To put the top edge at 40 pixels: Tilt 1949. Zoom stays 1.0.
Camera transform
Zoom 1.5 makes the fitted 1080 by 608 clip 1620 by 911. Crop is in fitted pixels before zoom and cropped pixels stay where they are, so crop a 720 pixel window around the head, then Pan to centre it. One Pan unit is one pixel, positive moves right. Tilt negative 1595 puts the head band on the floor. Find the head by exporting a neutral frame and reading it, never by guessing.
The hook as a Fusion clip
Two Text+ nodes over a transparent Background, merged. Font Anton, Size 0.204 for line one and 0.183 for line two, black outline from the second shading element at 0.08. Centres at 0.568 plus and minus 0.032. Two BezierSplines on the merge blends: 1 at frame 99, 0 at frame 100. Resolve rewrites GlobalOut on import, so fade with blends, not tool ranges. The clip sits on V1 under the media and shows through because the band is transparent in both layers.
Templates, not taste
The hook comp and the captions comp are templates. The agent swaps the text and the timings and nothing else. That is the single biggest reliability lever in every setup I studied: an agent asked to design a title will give you a different title every time, an agent asked to fill a template gives you yours. New looks get added as new template files, never as prompts.
Captions
A Fusion composition clip is capped at the five second generator default, so captions cannot live in one. Make a compound clip of screen, camera and audio, then import a captions comp onto the compound: one Text+ per three word chunk from the word timings, Size 0.094, centred at 0.508, each with its own blend spline that opens on the chunk's first frame and closes on its last. Sixty nodes for a fifty second short. The bridge reports the compound as failed and creates it anyway, so check the track, not the message.
Render
mp4, H.264, 1080 by 1920, 30 fps, audio aac 48 kHz. Add the job, start it, poll the status. Frame exports refuse to overwrite an existing file, so name every check frame uniquely or you will measure a stale picture. Sixty Text+ nodes make the render about ten times slower than the plain stack; that is the price of editable captions.
Review in the timeline, then approve
The script drops a red marker at every seam. Play through the markers and listen to each join before anything else happens. Andy Diep's rule is the right one: treat the agent like a junior editor on their first day and approve the first cut before it makes another. A marker you delete is a cut you rejected. That review takes two minutes and it is the whole difference.
Two passes on graphics, partial re-renders
First pass builds every graphic. Second pass goes one by one: move this down, it is on my forehead, make it smaller, change the word. Render only the range you changed with the range flag, not the whole short. Jason Cooperson's build is shaped around this loop and it is what makes the second hour feel like editing instead of waiting.
Give it to your agent, three ways
Same skill, three worlds. Pick the one you actually use. The skill file at the bottom of this page is the instructions in every case.
An agent with connectors (Claude Desktop, Claude Code, Grok Bot)
The factory scripts and the Resolve bridge are plain files and a local HTTP port, so any agent with a shell can run the whole thing. Claude Code and Grok Bot do this directly.
- Clone the Short Form Content Factory repo and read styles/blade-agent-demo.md, the locked look every short follows.
- In Resolve: Workspace, Scripts, CursorBridge. It prints the port it listens on.
- Give the agent the skill from the bottom of this page. Say the trigger with the folder that holds camera.mp4 and screen.mp4 and the words file.
- Layer 1 runs stack.py and the render script; Layer 2 runs resolve_stack.py with --hook and --words.
cd ~/Short\ form\ content\ factory/scripts/tella python3 stack.py <slug> # Layer 1: ffmpeg band stack python3 resolve_stack.py <tracks_dir> "L2 short 1" --hook "CHATGPT ANSWERS YOU|GROK BOT DOES THE WORK" --words <slug>.words.json --render
ChatGPT (a project or a custom GPT)
ChatGPT cannot drive Tella or Resolve, but it can do the thinking half well: pick the cuts and write the hooks from a transcript, and write the Resolve scripts you then run.
- Layer 1: paste the transcript with word indices into a project that carries the cut prompt above. Take its word ranges into Tella by hand or hand them to an agent that can call the Tella API.
- Layer 2: ask it for a Python or Lua script for Resolve that does one step, run it from Workspace, Scripts, and when it fails open Workspace, Console and paste the error back. That loop is how the settings on this page were found.
- Ask it for Lua you can paste into the Fusion console to build the Text+ nodes, then edit the text in the Inspector.
# a script dropped here shows up under Workspace > Scripts ~/Library/Application Support/Blackmagic Design/DaVinci Resolve/Fusion/Scripts/Utility/
Anything with an API (a token and a curl call)
Tella has a REST API for the cut and the export. The Resolve bridge is JSON over localhost. A token in an env variable, never in a chat.
- Export the Tella key once as TELLA_API_KEY. Duplicate the recording, cut by transcript word indices, then trim silences over 700 ms to 220 ms, then export with tracks granularity to get camera.mp4 and screen.mp4.
- Every Resolve step is a POST to the bridge: /projects/create, /project/setting, /media/import, /timeline/create, /fusion/insert, /media/insert, /clip/properties, /clip/fusion/import, /timeline/compound-clip, /render/start.
- Check frames come back from /project/export-frame. Read them before you render.
export TELLA_API_KEY=... # never paste it into a prompt curl -s https://api.tella.com/openapi.json | head -c 400 # Resolve bridge, read only curl -s http://127.0.0.1:9876/status curl -s "http://127.0.0.1:9876/timeline/clips?track_type=video&track_index=2"
Failure modes
Every one of these has happened to me or to someone I set this up for.
| Failure | Fix |
|---|---|
| Title sits on the screen, captions sit on the face | Export tracks, stack with the band, put all text in the band |
| Dead air at every cut | Trim silences after the word cut, leave about a fifth of a second |
| Black viewer in Resolve | Clips inserted at record frame 0; timelines start at 108000 |
| Media jumped five seconds | A Fusion clip was inserted after the media; insert it first |
| Silent render | Screen audio was on the timeline; only camera audio goes on A1 |
| Screen ended up too low | Tilt is 0.316 pixels per unit on a portrait timeline, not one |
| Head off to one side | Pan is one pixel per unit; find the head from a real frame |
| Hook still visible at five seconds | Merge blend fades the foreground only; put both lines over a transparent Background and fade both merges |
| Captions cut off at five seconds | A Fusion clip is capped at the generator default; put captions on a compound clip |
| Check frame never changes | Frame export will not overwrite; use a new file name every time |
| Every short comes out looking a little different | Point the agent at a template file, never describe the look in a prompt |
| Agent trimmed dead air in the wrong places | Have it mark, not cut. You approve the markers, then it cuts |
| A one word fix re-renders ten minutes | Render the changed range only |
The tools I use for this
| Tool | What it is for here | |
|---|---|---|
| Tella | Records screen and camera together and cuts by transcript through its API. | Get it |
| DaVinci Resolve | The free version is enough for Layer 2. Text+ nodes are the editable hook and captions. | no link, just use it |
| ffmpeg | Does the Layer 1 stack in one command. | no link, just use it |
| Claude | The agent that runs both layers here. | Open |
The free skill
It is the cut and stack procedure my agents run: read the transcript, pick the moments, cut in Tella by word index, export the camera and screen as separate tracks, stack them with the band, place the hook and captions in the band, render, and report the pauses at every seam so a human reads the result before it posts. It carries the Resolve settings that were measured, not guessed.
--- name: one-recording-ten-shorts description: Cut one screen and camera recording into vertical shorts by transcript in Tella, stack screen over head with a text band, place the hook and captions in the band, then rebuild the same reel inside DaVinci Resolve with editable Text+ nodes. Read only until the human approves the cut list. --- # One recording, ten shorts You are cutting a screen and camera recording into 30 to 65 second vertical shorts. The words are the edit. The text lives in the band between the screen and the head. Graphics go where they can still be changed. ## Inputs - The recording in Tella (id or share link) and its transcript with word indices. - The style card: styles/blade-agent-demo.md in the Short Form Content Factory repo. It is locked. Do not restyle. - The hook ledger. Never repeat a shipped hook. ## Layer 1: cut and stack 1. Read the transcript, not the video. Propose 6 to 10 shorts as start word and end word with indices, one idea each, a first line that earns the next second, a last line that closes. Stop and show the list. Wait for approval. 2. For each approved short: duplicate the recording, cut by word index (indices are not contiguous; snap to existing ones; zero length tokens at the boundaries are not audio), then trim silences over 700 ms down to about 220 ms. Leading silence to 150 ms. 3. Export with tracks granularity so you get camera.mp4 (1920 by 1080) and screen.mp4 (3840 by 2160). Never composite on a flattened Tella export: every Tella layout leaves no gap between the screen and the head. 4. Stack: canvas 1080 by 1920, screen scaled 1080 by 608 at y 40, camera cropped 1280 by 1080 around the head and scaled 1080 by 911 at y 1009. The band is y 648 to 1009 and only text goes in it. 5. Hook: two lines, all caps, Anton, line one white, line two red, black outline, fitted to 995 px wide, ink below y 655, at most 250 px tall. Line two names the tool. Score each hook: a number, a contrarian claim, direct YOU; keep 2 of 3 or better. Hook holds 0 to 3.18 s and fades by 3.32 s. 6. Captions: three word chunks from the word timings, Anton 70 px, top 905, black outline, hidden until 3.32 s. The hook owns the band alone. 7. Render, then build a review page: every finished cut's transcript with the pause length at every seam. Mark Clean, Watch, Fix. Flag any cut without a spoken call to action so one can be spliced from a take that has it. ## Layer 2: the same reel inside DaVinci Resolve Requires Resolve open and the CursorBridge script running (Workspace, Scripts, CursorBridge; JSON on 127.0.0.1:9876; free version is fine). Measured on Resolve 21 on a 1080 by 1920, 30 fps timeline. 1. Create a new project. Never load or modify an existing one. Set timelineResolutionWidth 1080, timelineResolutionHeight 1920, timelineFrameRate 30 before creating the timeline. Import camera.mp4 and screen.mp4. 2. Timelines start at 01:00:00:00, which is record frame 108000. Everything is inserted there. 3. Insert the Fusion composition first on V1 (it ripples every track). Import the hook comp into it: two Text+ nodes over a transparent Background, Anton, Size 0.204 and 0.183, black outline (shading element 2, thickness 0.08), centres 0.568 plus and minus 0.032, both merge blends keyframed 1 at frame 99 and 0 at frame 100. Resolve rewrites GlobalOut on import; fade with blends. 4. Insert screen as video only on V2, camera as video only on V3, camera as audio only on A1. Screen audio never goes on the timeline. 5. Screen: Tilt 1949 (one Tilt unit is 0.316 px here, which is (1080/1920) squared). Camera: Zoom 1.5, crop a 720 fitted pixel window around the head (crop is in fitted pixels before zoom and does not recentre), Pan to centre (one unit is one pixel, positive is right), Tilt negative 1595. Find the head from an exported neutral frame with the other layers disabled. Name every exported frame uniquely; export refuses to overwrite. 6. Compound screen, camera and audio into one clip, then import the captions comp onto the compound: one Text+ per chunk, Size 0.094, centre 0.508, each with a blend spline that opens on its first frame and closes on its last. A Fusion clip alone is capped at the five second generator default. 7. Export check frames at 1 s, 6 s and 20 s. Show them. Only then set render: mp4, H.264, 1080 by 1920, 30 fps, aac 48 kHz. Add the job, start, poll to completion, verify the file has audio above silence. ## Report Cut list with indices, hook per short with its score, seam pauses, the check frames, the render path and duration. Anything you could not verify is stated as such. Never post anything. ## Rules learned from the other Claude-in-Resolve setups (credited: Matt Penny, Andy Diep, Jason Cooperson, Greg) - The folder is the system. Open the factory folder every session: style card, templates, scripts, word file. Never prompt the look. - Templates, not taste. Fill the hook and captions templates. Never design a new look in a prompt; a new look is a new template file, approved first. - Mark, then cut. On a first pass, place red markers at proposed cuts and seams and stop. The human approves the markers before any cut is made. Treat yourself as a junior editor on day one. - Two passes on graphics. Build all, then refine one by one from the human's notes. Re-render only the changed range. - Plan on the strongest model, execute on the cheaper one. Ask for voice notes when the brief is thin; more context beats a shorter prompt. - Anything Resolve cannot draw (mockups, animated UI) is a HyperFrames overlay composited on top; text stays as Text+ so it remains editable.
What done looks like at thirty days
- Every recording is cut from its transcript, not scrubbed
- Hooks and captions sit in the band on every short, never on the face
- One short has been rebuilt in Resolve and a title edited in the Inspector
- The review page is read before anything posts
- The hook ledger has every hook and none repeat
- Recording is the only step you do by hand
Want the whole factory, not just the cut?
The AI CEO Lab has the full short form content factory: the style cards, the review page, the hook ledger, and the Resolve layer, with the sessions where I run it on a real recording.