What problem does it solve? Turning online videos and podcast episodes into usable, citable text normally requires manual note-taking or paid cloud transcription APIs. This Skill downloads audio from major platforms and runs ASR transcription entirely on your own machine, producing structured transcripts without any API key. ## Core Features & Use Cases - Multi-platform video transcription: Paste links from Bilibili, Douyin, Xiaohongshu, YouTube, or WeChat Channels and receive a cleaned transcript with semantic section headings and timestamps, powered by the local SenseVoice-Small model. - Podcast speaker diarization: Transcribe Xiaoyuzhou, Ximalaya, or Apple Podcasts episodes with paraformer + CAM++ to automatically separate host and guest, outputting a speaker-blocked Markdown transcript plus an SRT subtitle file. - Local file support and WeChat Channels delivery modes: Transcribe local mp4/m4a/mp3/wav files, and for WeChat Channels choose between an in-chat verbatim transcript, a text-only PDF, or a screenshot-illustrated PDF. - Use Case: Paste a one-hour podcast episode link and receive a speaker-labeled transcript with host and guest names mapped, plus an SRT file ready to burn as subtitles. ## Quick Start Paste a video or podcast link into the chat and ask the agent to transcribe it with the video-transcript skill, for example by saying "transcribe this Bilibili video into a cleaned transcript".