Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T10:15:37.280308Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 3 inbound Pith citation observations for arXiv:2605.19196.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T10:15:37.280308Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:10:34.573818Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-11T13:10:34.802140Z
81 of 81 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 356886a0-2e85-4089-9810-848988859a24 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? gpt-oss-120b & gpt-oss-20b Model Card
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d591a02f-3dd0-42a1-8bd3-8eff9f8892d4 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Introducing claude haiku 4.5
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1e2b9691-4bd0-49f3-ae10-d59cab5c1d98 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Introducing claude opus 4.7
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b9fccac4-766c-4e27-b441-3c6d35673f0f · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Benchmarking large language mod- els in retrieval-augmented generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 969b4227-fb9d-4029-990a-5fb841b7973d · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a097571f-6501-418a-b96b-aebcfc860649 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b01fbb24-2c84-4cfb-afd5-eed3f595cc59 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Alpacafarm: A simulation framework for methods that learn from human feedback
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c3584117-ebb2-4b22-911c-1f129064559a · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? RAGAS: Automated evaluation of retrieval augmented generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6233498b-3c8d-4fac-9938-7f12b5a55f94 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Are we on the right way to assessing LLM-as-a-judge?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0caa09fe-b545-48a3-9d4b-96c4ce779b96 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Deep Research Bench: Evaluating AI Web Research Agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 318d24b2-8816-4ef6-a277-f0cb637581eb · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Enabling large language models to generate text with citations
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation da12364b-f7d3-4649-a6eb-a79592244c2b · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Gemma 3 Technical Report
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 04325a0e-fa2d-4036-89a5-bc55b54321e7 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Gemini 2.0 is now available to everyone
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bcc17993-8406-43cf-b13b-d84b23233eeb · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Gemini 2.5 flash is now in preview
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 194b3c1d-b2f0-4459-8932-45a987bb5a55 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Gemini 3.1 pro model card
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5383023c-2a4f-4f3b-83b2-1000cbb0c944 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? The Llama 3 Herd of Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dca70804-cc33-4b7b-8515-941b3337bb5d · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 81b5bc79-ac7d-486d-8f14-d30d92cc335f · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Step-DeepResearch technical report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3898ed0f-35ae-4cc8-b26b-8a8d5e8b9f09 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? MetaTool benchmark for large language models: Deciding whether to use tools and which to use
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0f680140-27ca-4a5d-a314-fc12344d72fe · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Hwang, Varsha Kishore, Amanpreet Singh, Dany Haddad, Aakanksha Naik, Malachi Hamada, Jonathan Bragg, Mike D’Arcy, Daniel S
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6abafd53-aa75-4eb9-8c8a-fe8315b24a6a · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Toolscan: A benchmark for characterizing errors in tool-use LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 922621ac-c323-491b-aef0-811d3aa9a7ba · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Deepwidesearch: Benchmarking depth and width in agentic information seeking
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d087d9cb-616d-4dd8-aeae-6c3e5e20ec50 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Retrieval-augmented generation for knowledge-intensive nlp tasks
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 63874281-696b-42b9-ba4b-30048d32597d · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3e09ee99-d4ad-46ba-9b71-fb408140785a · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 14b5fa7a-51fe-4f1d-94dc-e7de63b735fd · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? G- Eval: NLG evaluation using GPT-4 with better human alignment
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a7836ef6-5a2f-4b99-92ea-09925b0210ce · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ReIFE: Re-evaluating instruction-following evaluation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b4e25498-b6a4-464b-9e3f-20a859819fe4 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? In: Zong, C., Xia, F., Li, W., Navigli, R
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 649f98e2-0ca0-4ada-ab91-a782de3a5971 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ToolSandbox: A stateful, conversational, interactive evaluation benchmark for LLM tool use capabilities
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aa31df63-7925-4dd9-a33e-9601a3395dde · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Agentrewardbench: Evaluating automatic evaluations of web agent trajectories
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1ba4f070-f8b4-4409-98f8-0a29e4918be6 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Smith, Hannaneh Hajishirzi, and Nathan Lambert
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 97044d86-e57e-4aa4-b6f0-68159b8ed96d · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? An expert schema for evaluating large language model errors in scholarly question-answering systems
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3322ce45-6ae3-4139-ad25-df4ee6620668 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 09522d14-4ec3-40e2-97c4-c19e784f5f1c · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? the moon is made of marshmallows
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fdaaacf6-2ddf-46c2-b030-58a7d2d6c1ec · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? WebGPT: Browser-assisted question-answering with human feedback
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ab0962c0-a14f-41e7-aaa3-c9c0ccb05c99 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? RAGTruth: A hallucination corpus for developing trustworthy retrieval- augmented language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 969d2b44-2c8d-48aa-979c-faa1880548ec · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? GPT-5 mini
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 92eb6bda-046e-4194-8ef5-d72d36a0c282 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Introducing gpt-5.3-codex
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 68e52a06-4b6c-47a3-8ff6-3ad435c5f18a · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Gpt-5.4 model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 912bad08-d1f3-4ca0-ae3a-ed7cf5b280f3 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Gonzalez
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5721c5c7-9218-4214-8127-a7b59ac5d1e8 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Direct preference optimization: Your language model is secretly a reward model
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c212035f-dee0-404a-bf18-d96c7ba46e55 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 56b562ed-bfb9-4e69-82cd-05316a156854 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ARES: An automated evaluation framework for retrieval-augmented generation systems
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4eff1a9d-e06e-4578-b371-4aaf7ec3b895 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Localizing and miti- gating errors in long-form question answering
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 47714c6d-5f04-41cf-ac81-9d31a29eba7d · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Toolformer: Language models can teach themselves to use tools
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bf251382-38d7-4561-85dd-ea40cf98d7cc · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? arXiv preprint arXiv:2509.22391 , year=
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3ac7d08b-3892-454b-ac3f-1decb6b4093f · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f269b83c-f0bb-48b5-877f-e617acc62f6e · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ResearchRubrics : A benchmark of prompts and rubrics for evaluating deep research agents
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 212c6f5e-7fe0-4b81-bab2-c7ece8e43cf7 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Judgebench: A benchmark for evaluating LLM-based judges
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation da382f13-bc90-461b-b3d5-58b74139c138 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Tongyi DeepResearch Technical Report
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 427f3014-7067-40f6-ae4b-62e6b35c2a06 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? DeepResearchEval: An automated framework for deep research task construction and agentic evaluation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4e4aad11-32d8-4b65-a560-297f5ee5b0d3 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Long-form factuality in large language models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6d507140-ac43-4703-9b41-7327e287f8c5 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Qwen3 Technical Report
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c5d36093-ef6e-449b-a5c2-01b396888bb9 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? React: Synergizing reasoning and acting in language models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d7327b84-da56-4402-b4b6-b38b382311ad · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Yao, Y ., Wang, Y ., Zhang, Y ., Lu, Y ., Gu, T., Li, L., Zhao, D., Wu, K., Wang, H., Nie, P., Teng, Y ., and Wang, Y
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2a731a01-b628-4c2a-a7d0-b4ba7f035e81 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? MiroEval: Benchmarking multimodal deep research agents in process and outcome
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e252c4b4-7021-41c4-84a1-9964e21c1868 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Yifei, Allen Chang, Chaitanya Malaviya, and Mark Yatskar
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f560a9ef-d05c-4437-b517-584a48670f8e · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Yifei, Allen Chang, Chaitanya Malaviya, and Mark Yatskar
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4c37ec1c-6c41-47c5-bb82-651c89a329da · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Automatic evaluation of attribution by large language models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b73c0e94-e94b-4087-ba06-fe4aa8320fe3 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Evaluating Large Language Models at Evaluating Instruction Following
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8f500d94-9707-444d-a5ef-0b0d76f3b977 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2dc813d6-d429-420a-8d3b-35eba1a23db2 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? SRR-Judge: Step-level rating and refinement for enhancing search-integrated reasoning in search agents
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ef89c5f7-10bd-4e41-9d91-679a9d9b9e6e · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8d90eb97-92ab-4e9d-9027-44b79ad5dd7b · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Chaining the evidence: Robust reinforcement learning for deep search agents with citation-aware rubric rewards
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f0575d26-c00e-42ab-afd2-5e1ef1aeb063 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ToolBeHonest: A multi-level hallucination diagnostic benchmark for tool-augmented large language models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8c846851-0a1e-454c-8505-e648e08f14d8 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Judging LLM-as-a-judge with MT-Bench and chatbot arena
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 30e59f45-05f2-4ec3-9ba0-76165526b25e · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5a1d8425-0a16-4cfa-b107-8bc06dd9f7ff · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Evaluating judges as evaluators: The JETTS benchmark of LLM-as-judges as test-time scaling evaluators
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2b9aa100-1956-4a75-b1db-1596db2c9d91 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? verbose database queries correlate with null results
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e7158c4f-48d6-4b03-9c90-ba437f82154e · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? coherence
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 82bcd1fe-3c08-46b5-a2c4-a6e7b8c8b29e · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4f120e13-9101-4a27-8dba-9969294aad0b · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1ef56faa-67bd-47fa-97a7-89579e9d75d5 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3d956fd0-8ce5-468d-adae-e7a766f5efc9 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation aa61c76c-9b30-4974-9dd6-8eae63d94220 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 968d2f76-2d18-46d1-920b-3a525e7bdb4c · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fe88269b-071e-40c9-bd60-de959bae05aa · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 934cfd80-e1cc-4e0b-974c-4a79709d4bf3 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f6a8d7db-9e8f-4de9-ba69-c0dd5ceac4dc · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0bace29a-5f93-4538-b101-8e04483d3f72 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7c6deaad-030e-42f9-8074-78d589c73673 · outbound
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? appropriate human control
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 77399758-2e3c-4b3a-bd63-8061acd33a9e · inbound
Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e182a7d9-0c82-410c-96fe-4efa8f5a5b55 · inbound
Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55fbe9c5-8aad-44a1-9f4b-1665da83e741 · inbound
Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.