RISC reformulates self-consistency answer selection as a ranking task solved by a lightweight LambdaRank model with five hand-designed features, yielding better accuracy-efficiency trade-offs than majority voting on QA benchmarks.
Wildhallucinations: Evaluating long-form factuality in llms with real-world entity queries
4 Pith papers cite this work, alongside 5 external citations. Polarity classification is still indexing.
representative citing papers
LoVeC uses RL to train LLMs to output verbalized numerical confidence scores for statements in long-form text, achieving better calibration than self-consistency baselines on QA datasets while being 20x faster.
Per-Entity Bias Mapping claims aggregate visibility metrics fail because large brands exhibit higher fabricated citation rates than smaller ones in AI responses, attributed to the Brand Hallucination Paradox.
citing papers explorer
-
Boosting Self-Consistency with Ranking
RISC reformulates self-consistency answer selection as a ranking task solved by a lightweight LambdaRank model with five hand-designed features, yielding better accuracy-efficiency trade-offs than majority voting on QA benchmarks.
-
LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations
LoVeC uses RL to train LLMs to output verbalized numerical confidence scores for statements in long-form text, achieving better calibration than self-consistency baselines on QA datasets while being 20x faster.
-
Per-Entity Bias Mapping for AI Visibility: Why Brand Mentions Require Entity-Specific Calibration
Per-Entity Bias Mapping claims aggregate visibility metrics fail because large brands exhibit higher fabricated citation rates than smaller ones in AI responses, attributed to the Brand Hallucination Paradox.
- To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling