What problem does it solve? Raw talking-head, interview, or podcast footage lacks visual structure, making it hard to highlight key takeaways for social or presentation audiences. This Skill packages an existing clip with timed, designed graphic overlay cards — kinetic titles, lower-thirds, data callouts, quotes, side panels, and picture-in-picture — synced to what is actually being said, while the original video plays untouched underneath. ## Core Features & Use Cases - Transcript-driven card design: Extracts audio with ffmpeg, transcribes locally via Whisper (hyperframes transcribe), and derives card timing and content directly from the word-level transcript. - Flexible visual system: Choose output ratio (16:9, 9:16, 4:5), layout (split, stack, pip, overlay), one of 10 card styles, and a video frame; card count auto-infers from video duration and information density. - HTML/GSAP composition rendering: The agent writes each card as an HTML fragment, assembles a single GSAP-driven composition, and renders it to MP4 via the hyperframes CLI. - Use Case: Take a 2-minute founder interview clip, confirm a 9:16 portrait canvas with a stack layout and editorial style, and produce an MP4 where 12 designed takeaway cards appear in rhythm with the speaker's points — ready for TikTok or Reels. ## Quick Start Ask the agent to package your talking-head video with designed graphic overlay cards, for example: "Dress up my interview clip interview.mp4 with on-screen graphic cards for a 9:16 vertical output."