HyperAIHyperAI

Command Palette

Search for a command to run...

3 months ago

End-to-end Dense Video Captioning as Sequence Generation

Wanrong Zhu Bo Pang Ashish V. Thapliyal William Yang Wang Radu Soricut

End-to-end Dense Video Captioning as Sequence Generation

Abstract

Dense video captioning aims to identify the events of interest in an input video, and generate descriptive captions for each event. Previous approaches usually follow a two-stage generative process, which first proposes a segment for each event, then renders a caption for each identified segment. Recent advances in large-scale sequence generation pretraining have seen great success in unifying task formulation for a great variety of tasks, but so far, more complex tasks such as dense video captioning are not able to fully utilize this powerful paradigm. In this work, we show how to model the two subtasks of dense video captioning jointly as one sequence generation task, and simultaneously predict the events and the corresponding descriptions. Experiments on YouCook2 and ViTT show encouraging results and indicate the feasibility of training complex tasks such as end-to-end dense video captioning integrated into large-scale pretrained models.

Benchmarks

BenchmarkMethodologyMetrics
dense-video-captioning-on-vittE2ESG
CIDEr: 25.0
METEOR: 8.1

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp
End-to-end Dense Video Captioning as Sequence Generation | Papers | HyperAI