THE SHORT ANSWER
A video is a sequence of frames. Claude can't paint frames, but it can write a program that paints them: a web page where every element's position, colour and opacity is a function of time. Play that page in a browser and you get animation. Capture the browser 30 or 60 times a second and you get a video file.
So when someone says "Opus 5.5 made this video", what usually happened is: they described the video they wanted, Claude wrote an HTML file (or a React component, or a Three.js scene, or a Blender Python script), and a second tool turned that code into an MP4. Sometimes that second tool is just a screen recorder.
- 1You write a brief
One line ("make a 15-second motion graphics showreel, go all out") or a multi-page shot list with timings, colours and copy.
- 2Claude writes the code
HTML with CSS or GSAP animation, SVG, canvas, WebGL shaders, a Three.js scene, or a composition for a framework like Remotion or HyperFrames.
- 3A renderer captures frames
A headless browser steps through the animation frame by frame and saves each one, or you simply screen-record the page playing.
- 4FFmpeg assembles the file
Frames plus any audio (voiceover, music) become an MP4. Rendering frameworks do this step for you.
THE THREE PIPELINES
Almost every video in the index fits one of three setups. They differ in how much control you get over timing, and how much work it is.
1 · One-shot page + screen recording
Ask Claude for a self-contained animated HTML file, open it, record the screen. Fastest route, and the natural fit for the viral one-line showreel prompt. Timing depends on your machine, and there's no audio unless you add it.
2 · Coding agent + rendering framework
Claude Code (or another agent) builds a project in HyperFrames (HTML → MP4) or Remotion (React → MP4). Frames are captured deterministically, so a 30-second video is exactly 30 seconds every time, and you can re-render after edits.
3 · Multi-model production
Claude writes the script and the animation code; other models supply the voice (ElevenLabs, Kokoro, Gemini TTS), music, images or live-action shots; FFmpeg or an editor stitches it together. This is how the longer explainers and music videos were made.
Also: Blender and native tools
A few creators had Claude write Blender Python to model and animate 3D scenes, or drive After Effects. Same principle: Claude writes the instructions, the tool renders.
PIPELINE 1: ONE PROMPT, ONE PAGE
The prompt that started the wave was a single sentence. 86 of the 277 prompts in this index are a variation of it: "make a dynamic 15-second motion graphics video that shows what an incredible motion designer you are, like it's your showreel for a résumé. go all out."
Claude answers with one HTML file: a stage, a timeline of scenes, and animation written in CSS keyframes, the Web Animations API, or a canvas render loop. Open it in a browser and it plays. The simplest way to turn that into a post is to record the screen and trim the start and end.
Why this works better than you'd expect: motion design is mostly layout, typography, easing curves and timing, which are exactly the things CSS and JavaScript express precisely. Claude has seen an enormous amount of front-end code. What changed with Opus 5.5, by the account of the people making these, is taste: the pacing, the easing, and the restraint look designed rather than generated.
Ask for "a single self-contained HTML file, 1920×1080 stage, 15 seconds, no external assets, loops once". Constraints like a fixed stage size and duration make the result far easier to record cleanly.
PIPELINE 2: RENDER FRAMEWORKS
Screen recording is fragile: a dropped frame, a notification, a slow machine and the timing is off. Rendering frameworks fix this by controlling the clock. Instead of playing the animation in real time, they ask the page to draw frame 1, capture it, draw frame 2, capture it, and so on, then encode the result.
| HyperFrames | Remotion | |
|---|---|---|
| You write | HTML, CSS and seekable JS animation | React components |
| Output | Deterministic MP4 | Deterministic MP4 (and more) |
| Agent support | Ships skills for coding agents | Official agent skills for best practices |
| Good for | Motion graphics, launch videos, anything that starts as a web page | Data-driven and templated video, longer explainers |
In practice the workflow is: open Claude Code in an empty project, install the framework's skill, describe the video, and let the agent scaffold the composition, preview it, and render. Because the output is code, revisions are edits: "make scene three two seconds shorter" is a one-line change and a re-render, not a trip back to a timeline.
Creators in the index who named their stack used HyperFrames for product launch videos and Remotion for a narrated history of AI and a character-driven short. See the step-by-step guide.
PIPELINE 3: MULTI-MODEL PRODUCTION
Claude writes text and code; it doesn't speak, sing or photograph. The longer, more produced videos in the index bring in other models for those parts and use Claude as the director and the animator.
- Voice: ElevenLabs, Kokoro, Gemini TTS and other text-to-speech models read a script Claude wrote.
- Music: generated tracks, or in one case music produced by a Python script Claude wrote.
- Images and footage: image models or video models for elements code can't draw well, such as faces and photographic scenes.
- Assembly: FFmpeg, driven by a script Claude wrote, or a traditional editor.
The hard part here isn't any single model; it's synchronisation. Captions have to land on the words, cuts on the beat. Creators who did this well had Claude generate the voiceover first, read the timestamps back, and then write the animation against those timings.
WHAT CLAUDE OPUS 5.5 IS
Claude Opus 5.5 is Anthropic's flagship model, released on 22 September 2026. It's available in the Claude apps on paid plans (Pro, Max, Team and Enterprise; free users don't have it yet) and through the API as claude-opus-5-5, with a 1M-token context window, at $4 per million input tokens and $20 per million output tokens (Anthropic docs, TechCrunch).
Anthropic's launch material talks about coding, agents and professional work. It doesn't mention video. The motion-graphics wave was discovered by users in the first days after release, and every clip here is their work, not a demo from Anthropic.
WHAT IT COSTS
You pay for tokens, not seconds of video. By our rough estimate, a self-contained 15-second animation is a few thousand to a few tens of thousands of output tokens, which at $20 per million is cents, not dollars. On a Claude subscription, it's part of your usage allowance.
Costs climb with iteration and with agents. A Claude Code session that scaffolds a Remotion project, reads documentation, renders previews and takes several rounds of notes can use hundreds of thousands of tokens. Third-party voice, music and image models are billed separately by those providers.
WHERE IT BREAKS
- No real footage. Anything photographic, such as people, products on a table or real places, has to come from elsewhere. Code draws shapes, type and 3D geometry.
- Timing drift. Screen-recorded pages play at whatever speed your machine manages. Use a render framework when timing matters.
- Audio sync. Claude can't hear. Lining up visuals with speech or music needs timestamps fed back in.
- Not reproducible. The same prompt produces different code, and a different video, every run. Shared prompts are starting points.
- Long form. Past a minute or two, coherence depends on structure you impose: scene lists, a script, a style guide.
EXAMPLES FROM THE INDEX






RELATED OPEN-SOURCE SKILLS
- ★ 53,518heygen-com/hyperframesHTML/CSS-driven deterministic video rendering
- ★ 4,738remotion-dev/skillsOfficial Remotion best-practice skills for agents
- ★ 61,488calesthio/OpenMontageEnd-to-end AI video production pipelines
- ★ 2,134digitalsamba/claude-code-video-toolkitScript-to-render explainer and demo videos