Pith. sign in

REVIEW 1 cited by

NLCTables: A Dataset for Marrying Natural Language Conditions with Table Discovery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.15849 v1 pith:FY6ORHDT submitted 2025-04-22 cs.IR

classification cs.IR
keywords tablediscoverynlctablesdatasettableschallengingconditionslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the growing abundance of repositories containing tabular data, discovering relevant tables for in-depth analysis remains a challenging task. Existing table discovery methods primarily retrieve desired tables based on a query table or several vague keywords, leaving users to manually filter large result sets. To address this limitation, we propose a new task: NL-conditional table discovery (nlcTD), where users combine a query table with natural language (NL) requirements to refine search results. To advance research in this area, we present nlcTables, a comprehensive benchmark dataset comprising 627 diverse queries spanning NL-only, union, join, and fuzzy conditions, 22,080 candidate tables, and 21,200 relevance annotations. Our evaluation of six state-of-the-art table discovery methods on nlcTables reveals substantial performance gaps, highlighting the need for advanced techniques to tackle this challenging nlcTD scenario. The dataset, construction framework, and baseline implementations are publicly available at https://github.com/SuDIS-ZJU/nlcTables to foster future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Something's Fishy In The Data Lake: A Critical Re-evaluation of Table Union Search Benchmarks

    cs.IR 2025-05 conditional novelty 6.0 of 10

    Simple baselines rival or surpass specialized table union search models on current benchmarks, indicating these benchmarks reward surface overlap and general embeddings rather than isolating semantic understanding.

Pith tools