Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:11:34.287353Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 2 inbound Pith citation observations for arXiv:2509.04499.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:11:34.287353Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T00:05:16.851179Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T03:25:57.760381Z
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7823c93a-edcc-402f-b2f2-e680d0563bcf · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3013c672-c55a-4d1c-9682-a485aa7943df · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Ragas: Automated Evaluation of Retrieval Augmented Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9dc44bc-a79a-4578-b5c2-3e942b1c687e · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c798c6-88fc-41ce-bb7e-2f2f82e5404c · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be94324d-07da-420b-b4b8-c615ede5ab49 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Evaluating large language models for health-related queries with presuppositions
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b852ff80-9a9c-4d66-8c7b-3dfe52d4cc33 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence URL https://aclanthology.org/2024.findings-acl
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 781b68b5-d02b-423b-b110-12113439cfb4 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence LLMs as Factual Reasoners: Insights from Existing Benchmarks and Beyond
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c21a836-6d0a-4fe0-a99c-ae75ed72e686 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Evaluating verifiability in generative search engines
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126f7511-dd2e-421b-a71c-92957b87db53 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence WebGPT: Browser-assisted question-answering with human feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dccbcbb-204b-4e8d-bbc9-9b3b80d17c0b · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Towards a holistic approach: Understanding sociodemographic biases in nlp models using an interdisciplinary lens
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 64492b73-914c-40e4-a48a-e3624a884437 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Search engines in the ai era: A qualitative understanding to the false promise of factual and verifiable source-cited responses in llm-based search
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9593ae03-a769-400c-bd17-8ebad3f56a97 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10320c5b-8d04-4d97-b52c-8017b84917ad · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence AMRFact: Enhancing summarization fac- tuality evaluation with AMR-driven negative samples generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93a6344b-313e-430d-8bab-483dbfadaad6 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence doi: 10.18653/v1/2024.naacl-long.33
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ff17efd-2ccb-40aa-a291-e7281088edcc · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Evaluation of RAG Metrics for Question Answering in the Telecom Domain
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21e73924-ae84-4953-8327-dd62c7313cd2 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d1e1ee4-d3d4-4691-8a57-091e900ec4fe · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence An Audit on the Perspectives and Challenges of Hallucinations in NLP
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 739d0dda-4089-4998-8ef1-903c676ed2fd · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b060272-c084-446e-b192-9712ed71c677 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence ClashEval: Quantifying the tug-of-war between an LLM's internal prior and external evidence
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5779e042-20b3-4fdb-a671-1223b1e58183 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 825b2d9e-1750-4537-8cc2-a6e3fe9e482e · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4378976d-0391-4a1b-8fb9-0f15f5861088 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 168f026a-924d-4c0f-a961-74dde6bd3edc · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence why should we ban bottled water?
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 701066d6-a89f-4555-a29c-f495092c50dd · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence FABLES: Evaluating faithfulness and content selection in book-length summarization
Reference 850
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1817579f-0e06-4273-a8b9-e37e8552ba73 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Do LVLMs understand charts? analyzing and correcting factual errors in chart captioning
Reference 1973
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4225e9eb-2dfb-4a23-9860-9a1279076526 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Adaptive-rag: Learning to adapt retrieval-augmented large language models through question complexity
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3e402624-09d6-4fba-bd94-dfd4655d5470 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a57a8e9-b3f6-4ac9-9bd9-67b214770385 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Deep Research Agents: A Systematic Examination And Roadmap
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8587929c-a81f-45af-a2bb-4ff14eaefd61 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Deep Research Bench: Evaluating AI Web Research Agents
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49c24558-4356-4933-b435-8cc352dab517 · outbound
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Evaluating top- k rag-based approach for game review generation
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d4cf8085-492e-4ea4-8a3c-709cf9896ef3 · inbound
What if AI systems weren't chatbots? DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
Reference 185
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0d4dc21a-3bc8-4671-9efc-deb5df358022 · inbound
HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.