CUDA Video Editor
The CUDA Video Editor turns uploaded clips, generated video, narration, slides, subtitles, music, and transitions into finished MP4 or GIF deliverables — trimmed, graded, watermarked, subtitled, or fully restyled. You steer it from chat, and the finished timeline appears in Live -> Video Studio so you can preview, trim, reorder, and rebuild without starting over.
What It Is For
Use Video Editor when the job is about assembling or changing video:
- trim a long recording into a highlight reel
- add subtitles or voice-over
- convert landscape video into vertical short-form format
- stitch several generated clips into one story
- mix music, narration, and original audio
- add title cards, explainers, transitions, and watermarks
- restyle an entire video into a new art style while keeping faces and on-screen text
- finalize a polished deliverable for download or sharing
Reach for Video Lab to create or transform footage (generate clips, re-light a shot, upscale to 4K, rip a YouTube source). Reach for Video Editor to cut and assemble footage into a deliverable. They share one Studio view, and most real jobs use both.
See also: Video Lab Works On Your Footage — the release note on what Video Lab can now do with footage you already have.
How The Workflow Feels
- Upload or generate the source media.
- Ask Alfrada OS for the edit you want.
- Alfrada OS runs the editing steps on a dedicated GPU rendering server.
- The session's Live -> Video Studio panel shows the build steps while work is running.
- After Alfrada OS finalizes the video, Video Studio unlocks preview plus a horizontal timeline.
- You can trim, reorder, adjust transitions, move audio blocks, and click Update preview to rebuild.
You do not need to name internal operations in normal use. Describe the result and constraints, and Alfrada OS will choose the right editing path.
Main Capabilities
Ingest and inspect
Alfrada OS can probe uploaded videos for duration, resolution, frame rate, codecs, and audio. It can also ingest a recording into a timeline-style canvas so later edits refer to scenes rather than raw timestamps.
Ingest does more than measure the file. It transcribes the audio with word-level timing, labels who is speaking when the recording has more than one voice, lifts any subtitle track already embedded in the file, and describes the picture frame by frame. That transcript is what makes "add captions" a one-line request — ask for automatic subtitles and Alfrada OS uses the transcript it already has instead of asking you to type the cues.
Analysis cost scales with length, so for a long recording say which part you care about — "ingest 12:00 to 14:30" — and only that window is analysed. Ingesting anything over about twenty minutes returns a warning saying so.
Reframe and transform
One request covers the whole cleanup pass, and Alfrada OS applies it in a single render rather than one per change (each render costs a little quality, so doing them together matters):
- Reframe — crop to a target shape without letterboxing. "Make this 9:16 around the speaker on the left" keeps the full height and trims the sides, anchored where you say.
- Rotate and flip — fix footage shot sideways on a phone.
- Speed, reverse, and loop — speed ramps, boomerangs, slow motion. Ask for true slow motion and Alfrada OS synthesises the in-between frames rather than just holding each one twice.
- Fades — open from black as well as close to it.
- Frame rate — 24, 25, 30, 50, or 60.
- Cleanup — stabilise shaky handheld footage, denoise, sharpen, or deinterlace old material.
You can also hand Alfrada OS a raw FFmpeg filter chain when you want something outside the presets. Filters that read files from disk are refused.
Overlays and lower-thirds
Beyond the corner watermark, Alfrada OS can composite things onto a video: a logo or second clip as picture-in-picture, two videos side by side or stacked for comparison, a text card, or a lower-third name badge that appears at one timestamp and leaves at another. Overlay text is measured and wrapped to fit the frame, so unlike a watermark it will not run off the edge.
Trim and reformat
Use this for cutdowns, clips, social formats, and aspect-ratio conversion. The four resize modes are contain (letterbox), cover (crop to fill), contain-blur (blurred background behind a sharp center), and stretch.
You can name a size three ways: exact pixels ("1080x1920"), a platform preset (Instagram post, Instagram Reel, TikTok, YouTube Short, YouTube landscape, Twitter landscape, LinkedIn landscape, or 2K cinema — each carrying its own frame rate), or "keep the source size", which leaves 4K footage at 4K and vertical footage vertical instead of forcing everything through a 1920x1080 default.
Trims can start anywhere, not just at the beginning: "thirty seconds starting at 1:10" cuts that window.
Stitch and transition
Video Editor can join clips with any of these transitions: fade, crossfade, slide (left, right, up, or down), zoom, wipe (left or right), dissolve, pixelize, radial, and circle (open or close).
The two lists above are complete — this guide is the canonical reference for the editor's resize modes and transitions. Other pages give a summary and link here.
A hard cut is a hard cut: asking for no transition (or a transition of zero seconds) joins the clips without any crossfade. If you want breathing room between scenes, ask for a pause — "half a second of black between each scene" — and the editor inserts it without touching the clips on either side.
Color grades
Ask for a one-shot color look and Alfrada OS applies it across the whole video. Twelve presets are available: cinematic, black-and-white, matrix green, optimus blue, SpaceX black, Instagram warm, Instagram cool, vintage 8mm, cyberpunk, documentary, noir, and sketch. You can ask Alfrada OS to list them, describe the mood you want and let it pick, or change the color look later from the Live Video Studio timeline.
The twelve presets are a starting point, not the ceiling. Describe a look they do not cover and Alfrada OS can build the grade directly — the underlying colour, contrast, and detail controls are all available, and a custom look can be layered on top of a preset.
Subtitles
It can burn styled subtitles (font, size, position) into the video, with control over background opacity as well.
Burning in is not the only option:
- Caption files — export the cues as an
.srtor.vttsidecar to upload alongside the video or hand to a translator. Ask for automatic captions and the transcript from ingest is used directly. - Switchable captions — mux the caption file into the finished MP4 as a real subtitle track the viewer can turn off, instead of painting it onto the picture.
- Word-by-word highlighting — the karaoke look short-form social video uses, where each word lights up as it is spoken. This needs the word-level timings that ingest produces, so it works on transcribed audio rather than hand-typed cues.
Audio mixing
Layer narration, music, and source audio into one video. If a generated video already has audio, Alfrada OS checks before overlaying new narration so you do not accidentally bury dialogue or sound effects.
Narration keeps its timing when scenes are stitched. Each scene's soundtrack is locked to its picture — a line that finishes early leaves a real pause, and that pause survives the join and any music laid over the top. You can also place narration inside a scene ("start the voice-over one second in"), hold the last frame or add black to give a scene room, and turn per-line speech files into one narration track with a set gap between lines.
Slides and explainers
The editor can render title cards, bullet slides, full-frame image cards, and programmatic animated slides, then stitch them between live clips. The animated slides run on Manim, so mathematical notation and diagrams that build up step by step are both available — ask for an equation and it is typeset properly rather than drawn as text.
Image slides do not have to sit still. Ask for a slow push, a drift, or a Ken Burns move and the still gets motion, which is what keeps a six-second photo from reading as a frozen frame in a narrated explainer.
(To put a chart on screen, generate the chart image first and use it as an image slide.)
Scenes and jump cuts
Alfrada OS can find the shot boundaries in a recording and either report the timings or cut one file per shot, which is the fast way into "find the strongest moments" on footage you have not watched.
For talking-head recordings it can also strip the dead air — detecting the silences, removing them, and keeping a little breathing room on each side so words are not clipped at the joins. It reports how many seconds it actually removed.
Audio
Alfrada OS can lift the soundtrack out of a video as an MP3, WAV, AAC, or FLAC file — for transcription, for reuse under different picture, or just to keep the audio when the visuals are being replaced.
Restyle a whole video
Ask for a full visual restyle — "make this trailer Ghibli-style" — and Alfrada OS repaints the video frame by frame. It takes the key frames from the ingested video, restyles each one with Image Lab's identity-preserving edit (faces, composition, and on-screen text survive the new look), reassembles the restyled frames into video, and lays the original audio back over the result.
Two safeguards come built in:
- Every restyle is automatically audited by the local Image Consistency engine, which samples the restyled frames and attaches an advisory verdict on whether the look drifted between them.
- A single bad frame does not mean starting over — point at its timestamp and Alfrada OS redoes just that frame.
Restyle this trailer in watercolor style but keep the faces and on-screen text; if any frame drifts, redo just that frame.It names the target style, states what must be preserved, and pre-authorizes single-frame fixes instead of a full re-run.
Finalize
Finalizing is the last step. It applies final polish, marks the finished file (MP4 or GIF) as the deliverable, and makes the timeline editable in Live Video Studio.
Live Video Studio
Live Video Studio is the visual editing surface inside the Work panel. In short: while Alfrada OS builds, you watch a live step list; once the video is finalized, you get a preview player plus a horizontal timeline where you can reorder and trim clips, change transitions, apply a color look, move audio blocks, and click Update preview to rebuild.
For the full Studio walkthrough — the two phases, hands-on timeline editing, and how follow-up messages sync into the timeline — see Live Workspace. That page is the canonical guide to the Studio; this one covers what the editor itself can do.
Good Prompt Patterns
I uploaded a long recording. Find the strongest 3-4 moments, cut them into a 60-second highlight reel, add clean subtitles, smooth crossfades, and finalize it for download.This gives Alfrada OS the source, duration, editorial goal, subtitle requirement, transition style, and final deliverable request.
Turn this landscape demo video into a 9:16 Reel under 45 seconds. Use contain-blur so nothing important is cropped, add burned-in subtitles, and keep the pacing fast.The prompt names the target format and avoids accidental cropping.
Create a 60-second narrated explainer from these six images. Write the script, generate voice-over, add soft background music, use crossfades between sections, then finalize. I will tweak the timeline in Live Video Studio after.This chains script, speech, music, image timing, transitions, and the Live editing handoff.
Use the three Video Lab clips we just generated. Stitch them into one 20-second product teaser with a title card, consistent color look, background music, and a final logo watermark.Video Editor is the right follow-up once Video Lab has produced separate clips.
Best Practices
- Upload source files before asking for edits.
- Name the target format: horizontal, square, vertical, GIF, or MP4.
- State the duration limit up front.
- Say whether original audio should be kept, lowered, removed, or replaced.
- Ask Alfrada OS to finalize when you want the output to appear as an editable Live Video Studio timeline.
- After finalizing, make small visual timing changes in Video Studio instead of re-prompting the whole edit.
Heads Up
- Video Lab and Video Editor are not a strict generate-vs-edit split. Video Lab also works on footage you already have — prompt-edit a clip, upscale it to 4K, transfer an acting performance, or pull a YouTube video into the session — while Video Editor remains the assembly and finishing tool (trim, stitch, subtitles, audio mix, finalize). Both surface in the same Live -> Video Studio panel.
- Some AI-generated clips include native audio. Alfrada OS checks for that before mixing narration or music, but you should say whether to preserve, lower, or replace the source audio.
- Timeline durations can be approximate until Alfrada OS probes or trims the source file.
- A dedicated GPU rendering server does the heavy rendering. Very large videos can take longer than text or document workflows.
- Editing operations re-encode the whole file, so the time they take scales with the length of the source, not with how simple the edit sounds. Long renders are handed to a background job and Alfrada OS notifies you in the same conversation when they finish, so the chat is not frozen while a restyle runs. For a long recording, cut out the window you need before editing it.
- GIF output is capped at 640px wide and 12 frames per second, and carries no audio. It is produced by the build step, not by the finishing pass.
For the launch note, see Live Video Studio. For all tool capabilities, see the Tool Library.