A proposed pipeline shows LLMs introduce detectable race and gender biases when summarizing life narratives, creating potential for representational harm in research.
Novelqa: A benchmark for long-range novel question answering.arXiv preprint arXiv:2403.12766, 2024a
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CL 3representative citing papers
MemoryAgentBench is a multi-turn benchmark covering four memory competencies, and current memory agents fail at selective forgetting and long-range understanding.
Empirical study claiming to be the first broad comparison of chunking methods in RAG, highlighting effectiveness, cost, and generalization limitations across scenarios.
citing papers explorer
-
Whose Story Gets Told? Positionality and Bias in LLM Summaries of Life Narratives
A proposed pipeline shows LLMs introduce detectable race and gender biases when summarizing life narratives, creating potential for representational harm in research.
-
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
MemoryAgentBench is a multi-turn benchmark covering four memory competencies, and current memory agents fail at selective forgetting and long-range understanding.
-
Chunking Methods on Retrieval-Augmented Generation - Effectiveness Evaluation Against Computational Cost and Limitations
Empirical study claiming to be the first broad comparison of chunking methods in RAG, highlighting effectiveness, cost, and generalization limitations across scenarios.