ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation-Reference-Cited by-同舟云学术

ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation

Published:2023 Issue: Volume:11 Page:1719-1733
ISSN:2307-387X
Container-title:Transactions of the Association for Computational Linguistics
language:en
Short-container-title:

Author:

Wu Zhengxuan¹,Manning Christopher D.²,Potts Christopher³

Affiliation:

1. Stanford University, USA. wuzhengx@stanford.edu

2. Stanford University, USA. manning@stanford.edu

3. Stanford University, USA. cgpotts@stanford.edu

Abstract

Abstract Compositional generalization benchmarks for semantic parsing seek to assess whether models can accurately compute meanings for novel sentences, but operationalize this in terms of logical form (LF) prediction. This raises the concern that semantically irrelevant details of the chosen LFs could shape model performance. We argue that this concern is realized for the COGS benchmark (Kim and Linzen, 2020). COGS poses generalization splits that appear impossible for present-day models, which could be taken as an indictment of those models. However, we show that the negative results trace to incidental features of COGS LFs. Converting these LFs to semantically equivalent ones and factoring out capabilities unrelated to semantic interpretation, we find that even baseline models get traction. A recent variable-free translation of COGS LFs suggests similar conclusions, but we observe this format is not semantically equivalent; it is incapable of accurately representing some COGS meanings. These findings inform our proposal for ReCOGS, a modified version of COGS that comes closer to assessing the target semantic capabilities while remaining very challenging. Overall, our results reaffirm the importance of compositional generalization and careful benchmark task design.

Publisher

MIT Press

Subject

Artificial Intelligence,Computer Science Applications,Linguistics and Language,Human-Computer Interaction,Communication

Link

https://direct.mit.edu/tacl/article-pdf/doi/10.1162/tacl_a_00623/2200660/tacl_a_00623.pdf

Reference40 articles.

1. Lexicon learning for few shot sequence modeling;Akyurek,2021

2. Abstract Meaning Representation for sembanking;Banarescu,2013

3. Systematic generalization with edge transformers;Bergen,2021

4. Smatch: An evaluation metric for semantic feature structures;Cai,2013

5. Meta-learning to compositionally generalize;Conklin,2021