Command Palette
Search for a command to run...
Pratt Sarah ; Yatskar Mark ; Weihs Luca ; Farhadi Ali ; Kembhavi Aniruddha

Abstract
We introduce Grounded Situation Recognition (GSR), a task that requiresproducing structured semantic summaries of images describing: the primaryactivity, entities engaged in the activity with their roles (e.g. agent, tool),and bounding-box groundings of entities. GSR presents important technicalchallenges: identifying semantic saliency, categorizing and localizing a largeand diverse set of entities, overcoming semantic sparsity, and disambiguatingroles. Moreover, unlike in captioning, GSR is straightforward to evaluate. Tostudy this new task we create the Situations With Groundings (SWiG) datasetwhich adds 278,336 bounding-box groundings to the 11,538 entity classes in theimsitu dataset. We propose a Joint Situation Localizer and find that jointlypredicting situations and groundings with end-to-end training handilyoutperforms independent training on the entire grounding metric suite withrelative gains between 8% and 32%. Finally, we show initial findings on threeexciting future directions enabled by our models: conditional querying, visualchaining, and grounded semantic aware image retrieval. Code and data availableat https://prior.allenai.org/projects/gsr.
Code Repositories
Benchmarks
| Benchmark | Methodology | Metrics |
|---|---|---|
| grounded-situation-recognition-on-swig | JSL | Top-1 Verb: 39.94 Top-1 Verb u0026 Grounded-Value: 24.86 Top-1 Verb u0026 Value: 31.44 Top-5 Verbs: 67.6 Top-5 Verbs u0026 Grounded-Value: 40.6 Top-5 Verbs u0026 Value: 51.88 |
| grounded-situation-recognition-on-swig | ISL | Top-1 Verb: 39.36 Top-1 Verb u0026 Grounded-Value: 22.73 Top-1 Verb u0026 Value: 30.09 Top-5 Verbs: 65.51 Top-5 Verbs u0026 Grounded-Value: 36.6 Top-5 Verbs u0026 Value: 50.16 |
| situation-recognition-on-imsitu | JSL | Top-1 Verb: 39.94 Top-1 Verb u0026 Value: 31.44 Top-5 Verbs: 67.6 Top-5 Verbs u0026 Value: 51.88 |
| situation-recognition-on-imsitu | ISL | Top-1 Verb: 39.36 Top-1 Verb u0026 Value: 30.09 Top-5 Verbs: 65.51 Top-5 Verbs u0026 Value: 50.16 |
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.