Pith. sign in

REVIEW 3 cited by

AmbigDocs: Reasoning across Documents on Different Entities under the Same Name

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.12447 v3 pith:4YY6WVT2 submitted 2024-04-18 cs.CL

classification cs.CL
keywords differentdocumentsentitiesambiguousnameanswersacrossambigdocs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Different entities with the same name can be difficult to distinguish. Handling confusing entity mentions is a crucial skill for language models (LMs). For example, given the question "Where was Michael Jordan educated?" and a set of documents discussing different people named Michael Jordan, can LMs distinguish entity mentions to generate a cohesive answer to the question? To test this ability, we introduce a new benchmark, AmbigDocs. By leveraging Wikipedia's disambiguation pages, we identify a set of documents, belonging to different entities who share an ambiguous name. From these documents, we generate questions containing an ambiguous name and their corresponding sets of answers. Our analysis reveals that current state-of-the-art models often yield ambiguous answers or incorrectly merge information belonging to different entities. We establish an ontology categorizing four types of incomplete answers and automatic evaluation metrics to identify such categories. We lay the foundation for future work on reasoning across multiple documents with ambiguous entities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies

    cs.CL 2026-07 conditional novelty 6.0 of 10

    RegretBench evaluates LLM clarification as a sequential policy under hidden intent, showing that final accuracy alone misses large differences in interaction efficiency and stopping quality.

  2. MDBench: A Synthetic Multi-Document Reasoning Benchmark Generated with Knowledge Guidance

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MDBench is a synthetically generated, knowledge-guided benchmark for multi-document QA on which frontier LLMs achieve only about 60% exact match.

  3. Resolving Conflicting Evidence in Automated Fact-Checking: A Study on Retrieval-Augmented LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new dataset of claims with conflicting web evidence shows retrieval-augmented LLMs are fragile, and source-credibility cues help only modestly.

Pith tools