CutClaw - An open-source AI video editing tool developed jointly by the University of Taiwan and Beijing Jiaotong University.
CutClaw is an open-source AI video editing tool developed by the GVC Lab at the Greater Bay Area University and a team from Beijing Jiaotong University. The tool employs a multi-agent architecture and uses a 'music-driven' approach to automatically edit hours-long videos into short, rhythmically precise clips.
What is CutClaw?
CutClaw is an open-source AI video editing tool developed by the GVC Lab at the Greater Bay Area University and a team from Beijing Jiaotong University. Employing a multi-agent architecture, the tool automatically edits hours-long videos into precisely paced short clips using a "music-driven" approach. The system first analyzes the music's tempo and structure, then combines this with user text commands. An AI scriptwriter plans the shots, an editor selects segments, and a reviewer performs quality checks, ultimately rendering a cinematic video compatible with multiple platforms. CutClaw supports one-click material deconstruction and cache reuse, making it suitable for travel photography, marketing, and other scenarios.
CutClaw's main functions
- Music-Driven EditingBy analyzing musical beats, downbeats, and energy curves, and strictly aligning the visual narrative with the musical structure, true audio-visual synchronization can be achieved.
- Multi-agent collaborationSimulates a professional post-production workflow: AI screenwriter (plans story pacing and shots), AI editor (selects clip timings), and AI reviewer (inspects shot length and aesthetics), forming a closed-loop optimization.
- Command-based controlWith just one sentence of text description (such as "showing the protagonist's madness"), the system automatically understands the style and executes it, without the need to manually drag the timeline.
- Intelligent Material DeconstructionIt can break down hours-long videos into a structured shot library with a single click, annotating the camera techniques, characters' emotions, and narrative points; and extract the beat and energy features from the audio, converting it into searchable assets.
- Content-aware croppingAutomatically identifies the main subject of the image and intelligently adjusts the aspect ratio (9:16, 16:9, etc.) to adapt to the publishing needs of multiple platforms such as Douyin and Xiaohongshu.
- caching accelerationThe deconstruction results are cached after the initial processing, and can be reused directly when editing the same material again, greatly improving efficiency.
How to use CutClaw
- Installation EnvironmentAfter cloning the code repository from GitHub, create a Python 3.12 virtual environment and install project dependencies.
- Prepare materials:exist
resource/Place video and audio files in the respective directories. You can also optionally include subtitle files to skip speech recognition. - Start running:implement
streamlit run app.pyLaunch the visual interface, or run it directly by passing the file path and command parameters via the command line. - Configuration ModelIn the configuration file, set the API keys supported by LiteLLM, specifying the large models used for video understanding, audio parsing, and agent inference.
- Get the finished productWait for the system to automatically complete the material deconstruction, shot planning, editing and rendering, and download video files in various aspect ratios adapted to different platforms.
Key information and usage requirements for CutClaw
- Project BackgroundThe AI video editing system, jointly open-sourced by the GVC Lab of the Greater Bay Area University and Beijing Jiaotong University, enables music-driven automatic editing of long videos based on a multi-agent architecture.
- Core MechanismThe process employs a multi-agent pipeline of "screenwriter-editor-reviewer" to deconstruct the source material, generate structured subtitles, plan shots based on the music's beat (emphasis/energy/pitch), and finally render a short film with a precise rhythm and cinematic feel.
- Technology dependenceFor calling large model APIs via the LiteLLM gateway, Gemini-3/Qwen3.5 is recommended for video understanding, Gemini-3 for audio parsing, and MiniMax-2.7/Kimi-2.5 for agent inference.
- Environment configurationPython 3.12, Conda environment, GPU (CUDA) acceleration for video encoding and decoding is strongly recommended.
- Document PreparationYou need to put the video (.mp4/.mkv) and audio (.mp3/.wav) into the...
resource/Table of Contents, Optional .srt Subtitles Skip ASR to save time and API costs. - API ConfigurationYou must configure the API keys for each model provider (OpenAI, Google, Moonshot, etc.) through environment variables or configuration files.
- Operating modeSupports Streamlit visual interface (
streamlit run app.pyAccess localhost:8501 or via the CLI command line (python local_run.pyPass in the path and command parameters.
CutClaw's core advantages
- True music-driven editing Unlike traditional tools that "edit the video first and then add background music", CutClaw first deeply analyzes the music's beat, downbeat, and energy curve, allowing editing decisions to be completely driven by the music structure, achieving true audio-visual harmony.
- Professional-grade multi-agent collaboration Simulates the entire post-production process of film and television: AI screenwriters plan the narrative rhythm, AI editors select precise time points for segments, and AI reviewers conduct quality checks (shot length, main character proportion, aesthetic score), forming a self-correcting closed loop, rather than a single generation.
- End-to-end processing of long videos Optimized specifically for scenarios where "hours of footage are cut into short videos of a few minutes", it can deconstruct massive amounts of footage into structured, searchable assets with one click, and with the caching mechanism, it can achieve a highly efficient workflow of "slow initial cut, fast re-cut".
- Zero-threshold instruction control No professional knowledge is required; a single natural language description (such as "show the clown's madness and elegance") can drive stylized editing, automatically understanding emotions, rhythm, and visual preferences.
- Platform native adaptation Content-aware intelligent cropping automatically recognizes the main subject of the image and generates multiple aspect ratio versions such as 9:16 (Douyin), 16:9 (Bilibili), and 1:1 (Xiaohongshu) with one click, eliminating black borders and image cropping errors.
CutClaw's project address
- GitHub repositoryhttps://github.com/GVCLab/CutClaw
- arXiv technical paper: https://arxiv.org/pdf/2603.29664
Comparison of CutClaw's similar products
| Comparison Dimensions | CutClaw | OpusClip | Mora |
|---|---|---|---|
| Core positioning | Long-form video editing with a cinematic feel, music-driven narrative | Long video to short video, viral clip extraction | Video generation, multi-agent scene coordination |
| Music synchronization method | First, analyze the music structure.(Beat/Energy/Verse/Chorus), then driving visual editing decisions | Supports music beat alignment, focusing on extracting highlights from content and then adding background music. | Prioritizing visual consistency, music synchronization is not a core function. |
| Long video support | Hours(Hours-long) End-to-end processing | Supports converting podcast/live stream replays into short videos. | Support for long sequence generation |
| Architectural features | Multi-agent closed loop(Collaboration of screenwriter, editor, and reviewer) | Single-model algorithm recommendation | Multi-agent(Similar to the CutClaw architecture) |
| open source | yes | no | yes |
| Control method | Natural language command control style | Automatic extraction + manual adjustment of segments | Text prompt control generation |
| Applicable Scenarios | Travel photography/Vlog cinematic production, film and television derivative works | Social media marketing, live stream segments | Creative video generation, virtual scene construction |
Application scenarios of CutClaw
- Travel photography and Vlog productionA few hours of travel footage, combined with background music, can quickly generate short, cinematic films with precise pacing and natural timing, significantly saving post-production time.
- Film and television derivative works and mashupsIt can re-edit movie or TV series clips based on specific music rhythms to automatically generate mashup videos that are character-driven, emotion-driven, or plot-driven.
- Mass production of marketing contentBased on the same set of materials and different music styles, quickly generate multiple versions of promotional videos to suit the brand's needs for placement on different platforms.
- Multi-platform short video distributionAutomatic cropping and generation of various aspect ratios such as 9:16 (Douyin/Video Account), 16:9 (Bilibili), and 1:1 (Xiaohongshu), allowing for one-time production covering all platforms.
- Music Videos and Rhythm-Driven ContentUsing the ability to analyze musical structures, the visuals are precisely aligned with the musical beats to create visually appealing music content or dance videos with a strong sense of rhythm.