Pith. sign in

REVIEW 2 major objections 5 minor 33 references

Subjectivity in the Annotation of Bridging Anaphora

T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A double-annotation pilot on the GUM test set finds 401 bridging instances where the original corpus had 222, and shows that recognizing bridging anaphors is the most subjective stage of annotation.

desk verdict Useful pilot with honest numbers, but the under-annotation claim reads as a stretch; the definitional shift and author-led adjudication make the 401 vs 222 comparison too weak to carry it. read the letter →

arxiv 2506.07297 v1 pith:J4J2ZPKV submitted 2025-06-08 cs.CL

classification cs.CL
keywords bridginganaphoraannotationsubjectivityinformationstatusinter-annotatoragreementsubtypescorpusGUMdiscourse
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that identifying bridging anaphora—expressions like 'the door' understood via an earlier 'house'—is an inherently subjective task, and that existing annotated corpora may be severely under-annotated for it. The authors run a double-annotation pilot on the 26k-token test set of the GUM corpus, adding a new subtype taxonomy and adjudicating disagreements to produce a reference set (GUMBridge v0.1). The adjudicated set contains 401 bridging instances, nearly double the 222 originally annotated in the same data, at a density per 1k tokens comparable to the densest existing resources. Agreement analysis shows the main bottleneck is recognizing which mentions are bridging anaphors at all, with F1 of only 0.38 between annotators, while antecedent selection accuracy is 72% and subtype labeling reaches Cohen's kappa 0.58.

What carries the argument

The central object is the GUMBridge annotation pilot: a double annotation of the GUM test set using an information-status definition of bridging (an accessible first mention whose interpretation depends on a prior, non-identical entity), followed by classification into a newly proposed 11-category subtype taxonomy (COMPARISON, ENTITY, SET, and OTHER relations) and an adjudication step that merges the two annotators' judgments into a single reference set. The comparison of the adjudicated density against previous resources (ISNotes, BASHI, ARRAU RST, GUM v10) is what carries the under-annotation claim, while the per-stage agreement metrics (anaphor recognition F1, antecedent accuracy, subtype kappa) carry the subjectivity claim.

What would settle it

Take a random sample of documents from BASHI or ISNotes, re-annotate them with the GUMBridge guidelines and adjudication, and compare bridging density per 1k tokens to the original resource; if density does not rise substantially above the original 7.9 (BASHI) or 16.6 (ISNotes), the claim that these resources are severely under-annotated would be contradicted.

Watch

Extended reading notes

Core claim

The central discovery, stated on the paper's own terms, is that a careful information-status-based reannotation with adjudication finds substantially more bridging than previous resources captured: 401 instances in the GUM test set vs 222 in the original GUM v10, a density of 15.4 per 1k tokens on par with ISNotes (16.6) and ARRAU RST (16.5) and well above BASHI (7.9) and GUM v10 (8.5). At the same time, the study quantifies how subjective this task is: annotator overlap on anaphor recognition is low (F1 0.38), antecedent resolution is right 72% of the time when the anaphor is agreed, and bridging subtype agreement is moderate (kappa 0.58). The paper concludes that recognition, not subtype labeling, is the limiting stage, and recommends high-recall annotation with manual adjudication plus grounding world-knowledge judgments in a knowledge base.

Load-bearing premise

The under-annotation conclusion assumes that bridging density per 1,000 tokens is comparable across corpora with different genres and annotation definitions, so that the higher density in the pilot reflects greater coverage rather than a broader definition of bridging.

Editorial extensions

If this is right

  • Existing bridging resolution systems evaluated on BASHI or GUM may have been trained and tested against gold data missing roughly half the bridging instances, so reported performance may not reflect real-world recall.
  • Future annotation efforts for bridging should favor high recall with adjudication rather than trying to achieve high initial agreement, since low recognition agreement is expected.
  • Annotation guidelines should allow multiple bridging subtype labels per pair, because many disagreements (e.g., meronymy vs. set-member) reflect genuinely overlapping relations.
  • World-knowledge accessibility should be tied to an explicit knowledge base to reduce annotator variation in anaphor recognition.
  • The GUMBridge subtype taxonomy maps cleanly onto ARRAU's relation set, enabling direct comparison of resources.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the under-annotation finding generalizes, published bridging-resolution evaluation scores may be optimistic on precision and pessimistic on recall, since systems were tuned to recover fewer instances than actually exist.
  • The density comparison conflates definitional scope (lexical vs. referential bridging, mediated information status) with coverage; a targeted re-annotation of BASHI under GUMBridge guidelines is the direct way to separate those factors.
  • The low anaphor-recognition agreement suggests that bridging resolution should be cast as two sub-tasks—instance detection and link resolution—with separate evaluation, rather than a single end-to-end metric.
  • Since GUM is multi-genre while the comparison corpora are all news text, the density gap could partly reflect genre effects; testing on matched genres would clarify this.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper reports on an annotation pilot for bridging anaphora on the GUM test set, introducing a new annotation resource, GUMBridge v0.1, with a new subtype taxonomy. It measures inter-annotator agreement at three stages—anaphor recognition, antecedent resolution, and subtype classification—and finds that recognition agreement is very low (F1=0.38), antecedent resolution accuracy is 72%, and subtype agreement is moderate (kappa=0.58). Based on the adjudicated GUMBridge data containing 401 bridging instances (15.4 per 1k tokens) compared to 222 in GUM v10 on the same texts, the paper claims that some previous resources are likely to be severely under-annotated.

Significance. The subjectivity analysis is a useful contribution: the paper quantifies disagreement at each annotation stage, provides concrete examples of each type of subjectivity, and releases an adjudicated pilot dataset, which is a valuable resource for the community. The proposed subtype taxonomy and the recommendations for handling ambiguity are also potentially useful. However, the headline conclusion about severe under-annotation in previous resources is not supported by the evidence as presented, because the density comparison conflates definitional scope, genre, and adjudication with coverage error. The paper's strength lies in its careful documentation of subjectivity; the under-annotation claim needs to be reframed or supported with additional analysis.

major comments (2)
  1. [Section 2, Table 1] The cross-corpus density comparison cannot support the conclusion that previous resources are 'likely to be severely under-annotated.' The paper itself notes that BASHI deliberately excludes lexical bridging, ISNotes includes mediated comparative anaphora in its bridging count, ARRAU RST annotates entity-coherence-based relations, and GUMBridge adopts an information-status definition that 'greatly widens the scope' of bridging. The differences in instances per 1k tokens (7.9–16.6) are therefore confounded with annotation definition and corpus genre. The conclusion should be limited to 'under-annotated relative to GUMBridge's definition' or supported by a controlled comparison that applies the same guidelines to samples from the other corpora.
  2. [Section 3.5, Table 1 and Table 3] The inference that GUM v10 is under-annotated because GUMBridge adjudicated 401 instances versus 222 on the same test set is not secure. The adjudicated set was produced under a broader definition that includes subtypes such as COMPARISON-RELATIVE/TIME/SENSE and OTHER, which total 147 instances in Appendix A and may fall outside the original GUM v10 bridging definition; additionally, adjudication was led by Annotator A, one of the authors. Given that inter-annotator F1 for anaphor recognition is only 0.38 (Table 2), the 179 additional instances could plausibly arise from annotator variation and definitional expansion rather than from systematic under-annotation in GUM v10. The paper should provide an overlap analysis showing how many of the additional instances would qualify under the original GUM v10 guidelines, or it should temper the under-annotation claim accordingly.
minor comments (5)
  1. [Section 2] The text states that ISNotes distinguishes information that is 'known to the hearer and/or has been refereed to previously'; 'refereed' should be 'referred'.
  2. [Section 3.1] The phrase 'reader/header' appears in the definition of accessibility; it should likely be 'reader/hearer'.
  3. [Section 2] The sentence 'There have additional been efforts in areas closely related to bridging' contains a word-order error; it should read 'There have also been additional efforts'.
  4. [Table 3] The counts in Table 3 do not obviously reconcile with the final total of 401 GUMBridge annotations; the authors should specify how the adjudicated set was derived from the 562 raw agreement and disagreement annotations.
  5. [Appendix A] The count table would benefit from percentages alongside the raw counts to facilitate direct comparison with the density figures in Table 1.

Circularity Check

2 steps flagged · score 4.0 of 10

Under-annotation claim is partly definitional: broader GUMBridge schema and BASHI's referential-only scope guarantee density differences the paper reads as missed annotations.

  1. self definitional [Section 2, Table 1 discussion]
    "The corpus specifically contains annotations only for referential bridging, not lexical bridging. ... As we can see in Table 1, there has been considerable variation in the frequency of bridging annotations in previous resources, with ARRAU RST (counting both lexical and referential bridging) and ISNotes identifying bridging instances with approximately twice the rate per 1k tokens as the annotations in BASHI and GUM v10. This suggests that some previous bridging resources, such as BASHI and GUM, have likely been under-annotated for bridging instances."

    BASHI's lower density (7.9 per 1k) is the direct result of its annotation definition, which excludes lexical bridging, while ARRAU RST is counted by the paper as 'counting both lexical and referential bridging.' The under-annotation conclusion is therefore not an independent empirical discovery but a consequence of comparing unlike definitions; the paper provides no evidence that BASHI omitted instances its own referential-bridging definition would require.

  2. self definitional [Section 2 and Section 3.5 (density doubling)]
    "However, this information status based approach also greatly widens the scope of what should be considered bridging, which in turn increases the influence of subjective judgment by annotators. ... Notably, the number of instances in the test set of the GUM (v10) annotations nearly doubles, going from 222 instances of bridging to 401 in GUMBridge test, suggesting a significant improvement in coverage of bridging instances in this new annotation effort."

    The jump 222→401 is cited as evidence that GUM v10 was under-annotated, but GUMBridge adopted a deliberately broadened information-status definition that, in the paper's own words, 'greatly widens the scope' of bridging. Under the broadened definition, additional instances are expected by construction. The paper does not show that the extra 179 instances would have been annotated under GUM v10's original guidelines; absent that equivalence, the 'improvement in coverage' restates the input definitional choice rather than demonstrating missed instances.

full rationale

The annotation pilot's agreement measurements, confusion matrix, and qualitative subjectivity analysis are empirical and self-contained; they do not reduce to the paper's assumptions. The circularity is confined to the abstract's strongest claim ('some previous resources are likely to be severely under-annotated'). That inference rests on density comparisons across corpora whose definitions the paper itself acknowledges diverge: BASHI deliberately excludes lexical bridging, and GUMBridge adopts an information-status view that 'greatly widens the scope' of bridging. The 222→401 increase and the BASHI/ARRAU density gap are therefore partly guaranteed by definitional choices, not by evidence of missed annotation under each resource's own guidelines. The paper does not perform the overlap/definitional-equivalence analysis needed to separate coverage error from scope widening. No fitted parameters are renamed as predictions and no load-bearing self-citation chain is used; the subjectivity findings remain independent. Overall circularity score 4.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claims rest on a definitional choice for bridging scope, the use of existing GUM annotations as a base, and the treatment of one author as the reference annotator. These are domain assumptions that shape both the density comparison and the agreement measurements.

assumptions (4)
  • domain assumption Information status definition of bridging (Accessible via inference from a previous, non-identical entity) is the correct scope for bridging annotation.
    Adopted from ISNotes and BASHI; Section 3.1. The under-annotation comparison with BASHI and GUM depends on this definitional choice.
  • domain assumption Existing GUM entity span and coreference annotations used as the base are accurate enough to not distort bridging judgments.
    Section 3.3: annotators identify bridging among pre-existing entity annotations in GUM v10.
  • domain assumption Annotator A (an author) serves as a valid reference standard for computing precision and recall of Annotator B.
    Section 3.4, Table 2: PRF of Annotator B relative to Annotator A; no independent gold standard.
  • standard math Cohen's kappa and PRF are appropriate agreement measures for this setting.
    Standard tools; used in Section 3.4.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Subjectivity in the Annotation of Bridging Anaphora." pith.science (2026). https://pith.science/paper/J4J2ZPKV

@misc{pith2026250607297,
  author       = {Pith},
  title        = {Pith review of: Subjectivity in the Annotation of Bridging Anaphora},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/J4J2ZPKV}},
  note         = {Machine review of arXiv:2506.07297}
}
read the original abstract

Bridging refers to the associative relationship between inferable entities in a discourse and the antecedents which allow us to understand them, such as understanding what "the door" means with respect to an aforementioned "house". As identifying associative relations between entities is an inherently subjective task, it is difficult to achieve consistent agreement in the annotation of bridging anaphora and their antecedents. In this paper, we explore the subjectivity involved in the annotation of bridging instances at three levels: anaphor recognition, antecedent resolution, and bridging subtype selection. To do this, we conduct an annotation pilot on the test set of the existing GUM corpus, and propose a newly developed classification system for bridging subtypes, which we compare to previously proposed schemes. Our results suggest that some previous resources are likely to be severely under-annotated. We also find that while agreement on the bridging subtype category was moderate, annotator overlap for exhaustively identifying instances of bridging is low, and that many disagreements resulted from subjective understanding of the entities involved.

Figures

Figures reproduced from arXiv: 2506.07297 by the authors.

Figure 1
Figure 1. Bridging Subtype Classification in GUMBridge v0.1. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Confusion matrix of bridging subtypes for [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Counts of bridging relation types by genre in [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

33 extracted references · 23 canonical work pages

  1. [1]

    Tatsuya Aoyama, Shabnam Behzad, Luke Gessler, Lauren Levine, Jessica Lin, Yang Janet Liu, Siyao Peng, Yilun Zhu, and Amir Zeldes. 2023. https://doi.org/10.18653/v1/2023.law-1.17 GENTLE : A genre-diverse multilayer challenge set for E nglish NLP and linguistic evaluation . In Proceedings of the 17th Linguistic Annotation Workshop (LAW-XVII), pages 166--178...

  2. [2]

    Herbert H. Clark. 1975. https://aclanthology.org/T75-2034/ Bridging . In Theoretical Issues in Natural Language Processing

  3. [3]

    Kerstin Eckart, Arndt Riester, and Katrin Schweitzer. 2012. https://doi.org/10.1007/978-3-642-28249-2_7 A Discourse Information Radio News Database for Linguistic Analysis , pages 65--76. Springer Berlin Heidelberg, Berlin, Heidelberg

  4. [4]

    Biaoyan Fang, Timothy Baldwin, and Karin Verspoor. 2022. https://doi.org/10.18653/v1/2022.findings-acl.275 What does it take to bake a cake? the R ecipe R ef corpus and anaphora resolution in procedural text . In Findings of the Association for Computational Linguistics: ACL 2022, pages 3481--3495, Dublin, Ireland. Association for Computational Linguistics

  5. [5]

    Yulia Grishina. 2016. https://doi.org/10.18653/v1/W16-0702 Experiments on bridging across languages and genres . In Proceedings of the Workshop on Coreference Resolution Beyond O nto N otes ( CORBON 2016) , pages 7--15, San Diego, California. Association for Computational Linguistics

  6. [6]

    Kian Kenyon-Dean, Eisha Ahmed, Scott Fujimoto, Jeremy Georges-Filteau, Christopher Glasz, Barleen Kaur, Auguste Lalande, Shruti Bhanderi, Robert Belfer, Nirmal Kanagasabai, Roman Sarrazingendron, Rohit Verma, and Derek Ruths. 2018. https://doi.org/10.18653/v1/N18-1171 Sentiment analysis: It`s complicated! In Proceedings of the 2018 Conference of the North...

  7. [7]

    Hideo Kobayashi, Yufang Hou, and Vincent Ng. 2023. https://doi.org/10.18653/v1/2023.acl-long.383 P air S pan BERT : An enhanced language model for bridging resolution . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6931--6946, Toronto, Canada. Association for Computational Linguistics

  8. [8]

    Hideo Kobayashi and Vincent Ng. 2020. https://doi.org/10.18653/v1/2020.coling-main.331 Bridging resolution: A survey of the state of the art . In Proceedings of the 28th International Conference on Computational Linguistics, pages 3708--3721, Barcelona, Spain (Online). International Committee on Computational Linguistics

Show all 33 references
  1. [9]

    Lauren Levine and Amir Zeldes. 2024. https://doi.org/10.18653/v1/2024.crac-1.5 Unifying the scope of bridging anaphora types in E nglish: Bridging annotations in ARRAU and GUM . In Proceedings of The Seventh Workshop on Computational Models of Reference, Anaphora and Coreferen...

  2. [10]

    Jessica Lin and Amir Zeldes. 2021. https://doi.org/10.18653/v1/2021.law-1.18 W iki GUM : Exhaustive entity linking for wikification in 12 genres . In Proceedings of the Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop, pa...

  3. [11]

    Katja Markert, Yufang Hou, and Michael Strube. 2012. https://aclanthology.org/P12-1084/ Collective classification for fine-grained information status . In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 795...

  4. [12]

    Natalia N Modjeska. 2004. Resolving other-anaphora

  5. [13]

    Anna Nedoluzhko, Ji r \' M \' rovsk \'y , and Petr Pajas. 2009. https://aclanthology.org/W09-3017 The coding scheme for annotating extended nominal coreference and bridging anaphora in the P rague dependency treebank . In Proceedings of the Third Linguistic Annotation Workshop...

  6. [14]

    Malvina Nissim, Shipra Dingare, Jean Carletta, and Mark Steedman. 2004. https://aclanthology.org/L04-1402/ An annotation scheme for information status in dialogue . In Proceedings of the Fourth International Conference on Language Resources and Evaluation ( LREC `04) , Lisbon,...

  7. [15]

    Tim O ' Gorman, Kristin Wright-Bettner, and Martha Palmer. 2016. https://doi.org/10.18653/v1/W16-5706 Richer event description: Integrating event coreference with temporal, causal and bridging annotation . In Proceedings of the 2nd Workshop on Computing News Storylines ( CNS 2...

  8. [16]

    Maciej Ogrodniczuk and Magdalena Zawis awska. 2016. https://doi.org/10.18653/v1/W16-0703 Bridging relations in P olish: Adaptation of existing typologies . In Proceedings of the Workshop on Coreference Resolution Beyond O nto N otes ( CORBON 2016) , pages 16--22, San Diego, Ca...

  9. [17]

    Cecilia Ovesdotter Alm. 2011. https://aclanthology.org/P11-2019/ Subjective natural language problems: Motivations, applications, characterizations, and implications . In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Te...

  10. [18]

    Barbara Plank, Dirk Hovy, and Anders S gaard. 2014. https://doi.org/10.3115/v1/P14-2083 Linguistically debatable or just plain wrong? In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 507--511, Baltimore,...

  11. [19]

    Massimo Poesio. 2004. https://aclanthology.org/W04-0210/ Discourse annotation and semantic annotation in the GNOME corpus . In Proceedings of the Workshop on Discourse Annotation, pages 72--79, Barcelona, Spain. Association for Computational Linguistics

  12. [20]

    Massimo Poesio and Ron Artstein. 2008. https://aclanthology.org/L08-1091/ Anaphoric annotation in the ARRAU corpus . In Proceedings of the Sixth International Conference on Language Resources and Evaluation ( LREC `08) , Marrakech, Morocco. European Language Resources Associat...

  13. [21]

    Ellen F. Prince. 1981. Toward a taxonomy of given-new information. Radical pragmatics, pages 223--255

  14. [22]

    Ant \`o nia Mart \'i

    Marta Recasens, Eduard Hovy, and M. Ant \`o nia Mart \'i . 2010. https://aclanthology.org/L10-1103/ A typology of near-identity relations for coreference ( NIDENT ) . In Proceedings of the Seventh International Conference on Language Resources and Evaluation ( LREC '10) , Vall...

  15. [23]

    Ina Roesiger. 2016. https://aclanthology.org/L16-1275 S ci C orp: A corpus of E nglish scientific articles annotated for information status analysis . In Proceedings of the Tenth International Conference on Language Resources and Evaluation ( LREC '16) , pages 1743--1749, Port...

  16. [24]

    Ina R \"o siger. 2018. https://aclanthology.org/L18-1058/ BASHI : A corpus of W all S treet J ournal articles annotated with bridging links . In Proceedings of the Eleventh International Conference on Language Resources and Evaluation ( LREC 2018) , Miyazaki, Japan. European L...

  17. [25]

    Paul R \"o ttger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022. https://doi.org/10.18653/v1/2022.naacl-main.13 Two contrasting data annotation paradigms for subjective NLP tasks . In Proceedings of the 2022 Conference of the North American Chapter of the Association...

  18. [26]

    a rtner, Agnieszka Falenska, Arndt Riester, Ina R \

    Katrin Schweitzer, Kerstin Eckart, Markus G \"a rtner, Agnieszka Falenska, Arndt Riester, Ina R \"o siger, Antje Schweitzer, Sabrina Stehwien, and Jonas Kuhn. 2018. https://aclanthology.org/L18-1457 G erman radio interviews: The GRAIN release of the SFB 732 silver standard col...

  19. [27]

    Olga Uryupina, Ron Artstein, Antonella Bristot, Federica Cavicchio, Francesca Delogu, Kepa Joseba Rodr \'i guez, and Massimo Poesio. 2019. https://api.semanticscholar.org/CorpusID:164858637 Annotating a broad range of anaphoric phenomena, in a variety of genres: the ARRAU corp...

  20. [28]

    Zeerak Waseem. 2016. https://doi.org/10.18653/v1/W16-5618 Are you a racist or am I seeing things? annotator influence on hate speech detection on T witter . In Proceedings of the First Workshop on NLP and Computational Social Science , pages 138--142, Austin, Texas. Associatio...

  21. [29]

    Ralph Weischedel, Sameer Pradhan, Lance Ramshaw, Martha Palmer, Nianwen Xue, Mitchell Marcus, Ann Taylor, Craig Greenberg, Eduard Hovy, Robert Belvin, et al. 2011. Ontonotes release 4.0. LDC2011T03, Philadelphia, Penn.: Linguistic Data Consortium, 17

  22. [30]

    Juntao Yu, Sopan Khosla, Ramesh Manuvinakurike, Lori Levin, Vincent Ng, Massimo Poesio, Michael Strube, and Carolyn Ros \'e . 2022. https://aclanthology.org/2022.codi-crac.1/ The CODI - CRAC 2022 shared task on anaphora, bridging, and discourse deixis in dialogue . In Proceedi...

  23. [31]

    Amir Zeldes. 2017. https://doi.org/http://dx.doi.org/10.1007/s10579-016-9343-x The GUM corpus: Creating multilayer resources in the classroom . Language Resources and Evaluation, 51(3):581--612

  24. [32]

    Amir Zeldes. 2022. https://doi.org/10.5210/dad.2022.102 Can we fix the scope for coreference? P roblems and solutions for benchmarks beyond OntoNotes . Dialogue & Discourse, 13(1):41--62

  25. [33]

    Shuo Zhang and Amir Zeldes. 2017. https://aaai.org/papers/619-flairs-2017-15451/ GitDOX : A linked version controlled online XML editor for manuscript transcription . In Proceedings of the Thirtieth International Florida Artificial Intelligence Research Society Conference (FLA...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.