Official pytorch repository for CG-DETR "Correlation-guided Query-Dependency Calibration in Video Representation Learning for Temporal Grounding"
-
Updated
Aug 21, 2024 - Python
Official pytorch repository for CG-DETR "Correlation-guided Query-Dependency Calibration in Video Representation Learning for Temporal Grounding"
[AAAI 2022] Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding
[ICLR 2025] TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that beats Qwen2.5-VL-72B with 73% fewer frames on LVBench.
paper list on Video Moment Retrieval (VMR), or Natural Language Video Localization (NLVL), or Temporal Sentence Grounding in Videos (TSGV))
Pytorch implementation of the paper 'Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding' (AAAI2024).
Official Implementation of Moment Alignment Transformer
[NCA] Official implementation of the paper Motion2Language, Unsupervised learning of synchronized semantic motion segmentation
[BMVC 2024] Official Implementation of the paper guided attention for interpretable motion captioning
Transformer with Controlled Attention for Synchronous Motion Captioning
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
Offline audio QA, transcription, translation, emotion, temporal grounding and timed summaries on Apple Silicon — native GigaChat Audio MLX.
This paper presents the VLMI framework to detect activities in complex videos. It combines Swin Transformer video features with language prompts and an EIoU-based similarity measure, enabling accurate, query-driven activity detection and timestamping, handling visual noise and temporal uncertainty without full manual labeling.
Semantic Augmentation and Gated-adapter Encoding for Video Moment Retrieval — closing the linguistic robustness gap with LLM-driven query augmentation and nested gated-adapters.
Multimodal video analysis — Whisper transcription, LLM chapter summarisation, CLIP key-frame extraction, and natural language Q&A over uploaded video content.
EMCompress: Video-LLMs with Endomorphic Multimodal Compression (ACL 2026 Findings)
Temporal context grounding for LLM sessions. Stops AI from living in the past.
Corpus-scale video moment retrieval benchmark — 75h of rights-cleared video, 500 queries, 4 difficulty tiers. Built on Mixpeek.
Add a description, image, and links to the temporal-grounding topic page so that developers can more easily learn about it.
To associate your repository with the temporal-grounding topic, visit your repo's landing page and select "manage topics."