Pith. sign in

Paper Citation Record · LEDGER

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

As of 10 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 2 inbound Pith citation observations for arXiv:2509.04499.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04499 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:11:34.287353Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T00:05:16.851179Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T03:25:57.760381Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7823c93a-edcc-402f-b2f2-e680d0563bcf · outbound

This paper cites BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:33.925590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:33.925590Z digest=sha256:29a24f98c719c985b79dbea4f447df75cf73b1be405d22a6ee83440cefbeb51c

Observation 3013c672-c55a-4d1c-9682-a485aa7943df · outbound

This paper cites Ragas: Automated Evaluation of Retrieval Augmented Generation.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Ragas: Automated Evaluation of Retrieval Augmented Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.144349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.144349Z digest=sha256:0cf346eab8c76c0009f567bc19091e0de10d0fae23b8681a3b7755d70ebcbe5e

Observation b9dc44bc-a79a-4578-b5c2-3e942b1c687e · outbound

This paper cites RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.167576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.167576Z digest=sha256:f6b757fecb30780f16f993cabdbdf6897eb50e4467bca58bae900d416db0f08a

Observation 11c798c6-88fc-41ce-bb7e-2f2f82e5404c · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 8

Resolution
malformed identifier
no resolver link, observed 2026-08-05T12:11:34.177271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.177271Z digest=sha256:900811066a067a08da3716d677a00106364a6aef2ec6153444c76e0d5dfc39f8

Observation be94324d-07da-420b-b4b8-c615ede5ab49 · outbound

This paper cites Evaluating large language models for health-related queries with presuppositions.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Evaluating large language models for health-related queries with presuppositions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:11:34.795469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T12:11:34.193117Z digest=sha256:afee9223b3f42ee05d96b862465dd1973ff7c148dec1d2d2b940eb7163e26441

Observation b852ff80-9a9c-4d66-8c7b-3dfe52d4cc33 · outbound

This paper cites URL https://aclanthology.org/2024.findings-acl.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence URL https://aclanthology.org/2024.findings-acl

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:11:34.777440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T12:11:34.198673Z digest=sha256:56352f974cf0330dcd955961fa510a8f7bb1a7090e18d470469e20722f2b37cd

Observation 781b68b5-d02b-423b-b110-12113439cfb4 · outbound

This paper cites LLMs as Factual Reasoners: Insights from Existing Benchmarks and Beyond.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence LLMs as Factual Reasoners: Insights from Existing Benchmarks and Beyond

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.208721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.208721Z digest=sha256:e4c1305b6c1beae4204dcf1ea2aa9c4fbece0cc1a080df6735b97d23b2cef26e

Observation 0c21a836-6d0a-4fe0-a99c-ae75ed72e686 · outbound

This paper cites Evaluating verifiability in generative search engines.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Evaluating verifiability in generative search engines

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.213637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.213637Z digest=sha256:0b0190f9cec5b92ed0fd34764b19e87a02aa48a658eafc41b846fa89e06659b6

Observation 126f7511-dd2e-421b-a71c-92957b87db53 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence WebGPT: Browser-assisted question-answering with human feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.218030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.218030Z digest=sha256:2976bf8de48eb51332c5028f505f0ce759cf3e8eb96d81b96ff5e461a95a6dce

Observation 6dccbcbb-204b-4e8d-bbc9-9b3b80d17c0b · outbound

This paper cites Towards a holistic approach: Understanding sociodemographic biases in nlp models using an interdisciplinary lens.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Towards a holistic approach: Understanding sociodemographic biases in nlp models using an interdisciplinary lens

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:11:34.749138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T12:11:34.222910Z digest=sha256:80c6a71af189e4d1fd7cec5556fdf339eeebbb7b03507d0a9152a8a8a9ec370b

Observation 64492b73-914c-40e4-a48a-e3624a884437 · outbound

This paper cites Search engines in the ai era: A qualitative understanding to the false promise of factual and verifiable source-cited responses in llm-based search.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Search engines in the ai era: A qualitative understanding to the false promise of factual and verifiable source-cited responses in llm-based search

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:11:34.732373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T12:11:34.227365Z digest=sha256:3a25b1a1c4b9e86be4ad4cdf9628e517579b12364de1d9921f7e233bdac8c86c

Observation 9593ae03-a769-400c-bd17-8ebad3f56a97 · outbound

This paper cites MLGym: A New Framework and Benchmark for Advancing AI Research Agents.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.231652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.231652Z digest=sha256:9dc10655fdb9295bb1bddfa69e4bc2e43bb52192decf85285fec0b45db01eb41

Observation 10320c5b-8d04-4d97-b52c-8017b84917ad · outbound

This paper cites AMRFact: Enhancing summarization fac- tuality evaluation with AMR-driven negative samples generation.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence AMRFact: Enhancing summarization fac- tuality evaluation with AMR-driven negative samples generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:11:34.715479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T12:11:34.237189Z digest=sha256:67cb4b019d30438227e083d4d772a85568d6d915c7bd70f18b8673afb195ff44

Observation 93a6344b-313e-430d-8bab-483dbfadaad6 · outbound

This paper cites doi: 10.18653/v1/2024.naacl-long.33.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence doi: 10.18653/v1/2024.naacl-long.33

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.241721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.241721Z digest=sha256:bd09cfab2a66c0f851c6e1419165300a08874d5b4bcc5c852f5c9cf0b6c27a44

Observation 6ff17efd-2ccb-40aa-a291-e7281088edcc · outbound

This paper cites Evaluation of RAG Metrics for Question Answering in the Telecom Domain.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Evaluation of RAG Metrics for Question Answering in the Telecom Domain

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.247219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.247219Z digest=sha256:19199401a59690caf991982d9fea0e73c4c0f004ab7036f0b8fc03f19e9ac67b

Observation 21e73924-ae84-4953-8327-dd62c7313cd2 · outbound

This paper cites MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.251807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.251807Z digest=sha256:9ef9942225b4114fb7e60227ded9e86a4fc7924c35ec74271590652870d77de2

Observation 3d1e1ee4-d3d4-4691-8a57-091e900ec4fe · outbound

This paper cites An Audit on the Perspectives and Challenges of Hallucinations in NLP.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence An Audit on the Perspectives and Challenges of Hallucinations in NLP

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.257317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.257317Z digest=sha256:a291396dd7a65b46605ac10d26eb0a595a6dce77e765b385c6da9dbfdc2e3b0d

Observation 739d0dda-4089-4998-8ef1-903c676ed2fd · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.262789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.262789Z digest=sha256:c1cfeab42ab34720a771d7c2cc870fa1a492c979bcdd4ad6344ea0fb53115469

Observation 0b060272-c084-446e-b192-9712ed71c677 · outbound

This paper cites ClashEval: Quantifying the tug-of-war between an LLM's internal prior and external evidence.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence ClashEval: Quantifying the tug-of-war between an LLM's internal prior and external evidence

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.268788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.268788Z digest=sha256:610b37acfbfc2a8275665736380ada506d1696de59b0ed00d926da2ee35f4f2b

Observation 5779e042-20b3-4fdb-a671-1223b1e58183 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.273162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.273162Z digest=sha256:7954b74fc6cd1b86458bcc77a18204287678ea3bb92e04e32acb1bfe195e6ac0

Observation 825b2d9e-1750-4537-8cc2-a6e3fe9e482e · outbound

This paper cites RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.278284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.278284Z digest=sha256:2772ed341208f1e239625f7baf962d12403c89ba97714b908c9e82bf016353e8

Observation 4378976d-0391-4a1b-8fb9-0f15f5861088 · outbound

This paper cites an unresolved cited work.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:11:34.696479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T12:11:34.283087Z digest=sha256:e77b787e07ebd85af09b8a0c5357b1a3ebb2e060428ac151728a51cc56ada8c8

Observation 168f026a-924d-4c0f-a961-74dde6bd3edc · outbound

This paper cites why should we ban bottled water?.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence why should we ban bottled water?

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:11:34.679820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T12:11:34.287353Z digest=sha256:b1727d973f8ea72a1e9cbc74b2df71d85a97c56ef0fef41db90f8b000da30416

Observation 701066d6-a89f-4555-a29c-f495092c50dd · outbound

This paper cites FABLES: Evaluating faithfulness and content selection in book-length summarization.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence FABLES: Evaluating faithfulness and content selection in book-length summarization

Reference 850

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.203975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.203975Z digest=sha256:8730569effd925a5c48ff161be9165e0623754cb458e259aa718d7b1dfa4df70

Observation 1817579f-0e06-4273-a8b9-e37e8552ba73 · outbound

This paper cites Do LVLMs understand charts? analyzing and correcting factual errors in chart captioning.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Do LVLMs understand charts? analyzing and correcting factual errors in chart captioning

Reference 1973

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:11:34.830141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T12:11:34.172678Z digest=sha256:44f4f4609d21827c54ec4baa46e465da2b94d6d758911c184713d2070898846e

Observation 4225e9eb-2dfb-4a23-9860-9a1279076526 · outbound

This paper cites Adaptive-rag: Learning to adapt retrieval-augmented large language models through question complexity.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Adaptive-rag: Learning to adapt retrieval-augmented large language models through question complexity

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:11:34.814135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T12:11:34.187938Z digest=sha256:6858fdfb0b7290bbe62e255f249c4a07642e88362b7fcd16ad47e2d76e60d844

Observation 3e402624-09d6-4fba-bd94-dfd4655d5470 · outbound

This paper cites DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.040150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.040150Z digest=sha256:356a9ee4ef70525dd14ff608834ba94f0321e92848a928210b4cc031a0fadcb2

Observation 5a57a8e9-b3f6-4ac9-9bd9-67b214770385 · outbound

This paper cites Deep Research Agents: A Systematic Examination And Roadmap.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Deep Research Agents: A Systematic Examination And Roadmap

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.182534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.182534Z digest=sha256:c6bc72df125f6c8d7320174be7b9fa683648a2ab59b4c55ba5e4206181172418

Observation 8587929c-a81f-45af-a2bb-4ff14eaefd61 · outbound

This paper cites Deep Research Bench: Evaluating AI Web Research Agents.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Deep Research Bench: Evaluating AI Web Research Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:33.637116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:33.637116Z digest=sha256:9ca0dd07ed8504d3540e106efd3a5d26febd050cfb44ab9177f98d245f9348b3

Observation 49c24558-4356-4933-b435-8cc352dab517 · outbound

This paper cites Evaluating top- k rag-based approach for game review generation.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Evaluating top- k rag-based approach for game review generation

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:11:34.845478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T12:11:33.766309Z digest=sha256:b9103663103767f30f522a362121dfd4aa81b7660f5e0858787579dddab9400c

Pith citing papers

Observation d4cf8085-492e-4ea4-8a3c-709cf9896ef3 · inbound

What if AI systems weren't chatbots? cites this paper.

What if AI systems weren't chatbots? DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:25:57.763934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T03:21:58.498789Z digest=sha256:f51ac723f3263e82eacce465639714a5fdbd69cd856e7b835937be5e1f443787

Observation 0d4dc21a-3bc8-4671-9efc-deb5df358022 · inbound

HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research cites this paper.

HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T00:05:16.851179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:05:16.851179Z digest=sha256:c99695c396339a6b955edb8b8bc2426bb02a625b5ce812c179aa09c06ea529c1