HyperAIHyperAI

Command Palette

Search for a command to run...

3 months ago

Improving Constituency Parsing with Span Attention

Yuanhe Tian Yan Song Fei Xia Tong Zhang

Improving Constituency Parsing with Span Attention

Abstract

Constituency parsing is a fundamental and important task for natural language understanding, where a good representation of contextual information can help this task. N-grams, which is a conventional type of feature for contextual information, have been demonstrated to be useful in many tasks, and thus could also be beneficial for constituency parsing if they are appropriately modeled. In this paper, we propose span attention for neural chart-based constituency parsing to leverage n-gram information. Considering that current chart-based parsers with Transformer-based encoder represent spans by subtraction of the hidden states at the span boundaries, which may cause information loss especially for long spans, we incorporate n-grams into span representations by weighting them according to their contributions to the parsing process. Moreover, we propose categorical span attention to further enhance the model by weighting n-grams within different length categories, and thus benefit long-sentence parsing. Experimental results on three widely used benchmark datasets demonstrate the effectiveness of our approach in parsing Arabic, Chinese, and English, where state-of-the-art performance is obtained by our approach on all of them.

Code Repositories

cuhksz-nlp/SAPar
Official
pytorch
Mentioned in GitHub

Benchmarks

BenchmarkMethodologyMetrics
constituency-parsing-on-atbSAPar
F1: 83.26
constituency-parsing-on-ctb5SAPar + BERT
F1 score: 92.66
constituency-parsing-on-penn-treebankSAPar + XLNet
F1 score: 96.40

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp