Ref-Long is a new long-context referencing benchmark on which all 13 tested LCLMs perform poorly, revealing a capability gap that simple retrieval benchmarks miss.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
Ref-Long is a new long-context referencing benchmark on which all 13 tested LCLMs perform poorly, revealing a capability gap that simple retrieval benchmarks miss.