HyperAIHyperAI

Command Palette

Search for a command to run...

3 months ago

Completeness Modeling and Context Separation for Weakly Supervised Temporal Action Localization

{ Yizhou Wang Tingting Jiang Daochang Liu}

Completeness Modeling and Context Separation for Weakly Supervised Temporal Action Localization

Abstract

Temporal action localization is crucial for understanding untrimmed videos. In this work, we first identify two underexplored problems posed by the weak supervision for temporal action localization, namely action completeness modeling and action-context separation. Then by presenting a novel network architecture and its training strategy, the two problems are explicitly looked into. Specifically, to model the completeness of actions, we propose a multi-branch neural network in which branches are enforced to discover distinctive action parts. Complete actions can be therefore localized by fusing activations from different branches. And to separate action instances from their surrounding context, we generate hard negative data for training using the prior that motionless video clips are unlikely to be actions. Experiments performed on datasets THUMOS'14 and ActivityNet show that our framework outperforms state-of-the-art methods. In particular, the average mAP on ActivityNet v1.2 is significantly improved from 18.0% to 22.4%. Our code will be released soon.

Benchmarks

BenchmarkMethodologyMetrics
weakly-supervised-action-localization-onCMCS
mAP@0.1:0.7: 32.4
mAP@0.5: 23.1
weakly-supervised-action-localization-on-1CMCS
mAP@0.5:0.95: 21.2
weakly-supervised-action-localization-on-2CMCS
mAP@0.5: 36.8

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp
Completeness Modeling and Context Separation for Weakly Supervised Temporal Action Localization | Papers | HyperAI