Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
-
Updated
Sep 11, 2026 - Python
Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
Turn videos into chapters, keyframes, subtitles, and structured data.
Adaptive video frame extraction for SfM, Gaussian Splatting, and photogrammetry.
Generalized Keyframe Extraction for VideoQA and Video-Guided Agentic Tasks. Accepted by ECCV 2026.
MCP server for video intelligence — analyze any video with AI vision. Extract keyframes, transcripts, and trading strategies from YouTube, Instagram & 1000+ sites.
Local-first MCP server for video: download any video from almost any URL, or get a transcript plus scene-aware deduplicated keyframes. YouTube, TikTok, Instagram, X, WeChat Channels, MP4/HLS. No cloud, no API keys.
A curated, mathematically rigorous survey and resource catalog for unsupervised, self-supervised, and multimodal foundation-model video summarization.
智能视频镜头分析工具:场景切分、关键帧提取、导演向运镜/构图/动作分析,支持成片/线稿双模式。
Fit a video into a token budget for multimodal LLMs. Picks the frames that carry the most information instead of sampling on a timer.
An end-to-end video summarization pipeline implementing and evaluating Compact BiLSTM and Interpretable Transformer architectures for keyframe selection. Features frozen MobileNetV3 feature extraction, custom Top-K duration budget policies, and performance benchmarking on the TVSum and SumMe datasets.
A unified pipeline for video frame reconstruction combining classical CV (HSV keyframes + Farneback optical flow) with deep-learning interpolation via RIFE v4.25 for high-quality, perceptually consistent frames.
Compress long video into LLM-ready artifacts - semantic representative frames, timeline-aligned subtitles, motion segments. Offline file transcription + RTSP live understanding. Python, MIT.
Sanitized HCMAIC 2026 E2E statistics, artifact hashes, execution status, and keyframe coverage research brief
Cinematic Scene Change Detector is a Streamlit + OpenCV tool for detecting video scene cuts using histogram comparison and frame differencing, with real-time threshold tuning and keyframe extraction.
A lightweight, opinionated keyframe extraction pipeline for short-form videos using scene detection and semantic similarity.
FastAPI backend for ReplayRAG - extracts keyframes and transcripts from recorded session segments, indexes them, and answers timestamped Q&A queries via a local Ollama (Llama 3.2) RAG pipeline.
Overview of a multimodal video summarization pipeline for information-dense content developed during a contract project at webAI (Open Project - UC Berkeley).
Turn any video into a timestamped transcript + labelled keyframe contact sheets so Claude, ChatGPT or Cursor can actually understand it. YouTube, TikTok, Facebook, Instagram or local files. Python, yt-dlp, faster-whisper, ffmpeg.
Media analysis pipeline for AI agents — scene-change keyframes and timestamped transcripts from video files
Facial Emotion Recognition from sequential images using ML techniques (thesis) / Αναγνώριση συναισθηματικής κατάστασης προσώπου από ακολουθιακές εικόνες με χρήση τεχνικών μηχανικής μάθησης
To associate your repository with the keyframe-extraction topic, visit your repo's landing page and select "manage topics."