Pith. sign in

Paper Citation Record · LEDGER

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

As of 23 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 3 inbound Pith citation observations for arXiv:2605.19196.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19196 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T10:15:37.280308Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:10:34.573818Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T13:10:34.802140Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact27
  • verified fuzzy38
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 356886a0-2e85-4089-9810-848988859a24 · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? gpt-oss-120b & gpt-oss-20b Model Card

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:18:11.838464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:5d35c381b4f415f8fee217e3e549cbf58f1d4102f5881fab09f37f9bbad51cca

Observation d591a02f-3dd0-42a1-8bd3-8eff9f8892d4 · outbound

This paper cites Introducing claude haiku 4.5.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Introducing claude haiku 4.5

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.358989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:c134a781c715e1e653f3e6eb1a80956d5ef3affcee38721ec16f5b7a83d5108b

Observation 1e2b9691-4bd0-49f3-ae10-d59cab5c1d98 · outbound

This paper cites Introducing claude opus 4.7.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Introducing claude opus 4.7

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.370094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:9f0af54c63edec1290541e25a192e89dc22a33bbc65220722c7f280469d090fc

Observation b9fccac4-766c-4e27-b441-3c6d35673f0f · outbound

This paper cites Benchmarking large language mod- els in retrieval-augmented generation.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Benchmarking large language mod- els in retrieval-augmented generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.346091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:cd1e026303266df506b79f50393decf9493a2cadcb7eb31a398f7a1e8ef1fd6f

Observation 969b4227-fb9d-4029-990a-5fb841b7973d · outbound

This paper cites Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.785323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:9c6d6630cea329e27072a4483590d40774e57272d94e4076a408e38d5f6763e9

Observation a097571f-6501-418a-b96b-aebcfc860649 · outbound

This paper cites DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:18:11.788133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:33d38a9bcae30b58773c4634dd92784bb8acfa43df3e42b1d966983ffc378f9f

Observation b01fbb24-2c84-4cfb-afd5-eed3f595cc59 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.349987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:92fc4b042346af2b32f4cf1a874d67e23685f9f36e80485262ac070d472b9b91

Observation c3584117-ebb2-4b22-911c-1f129064559a · outbound

This paper cites RAGAS: Automated evaluation of retrieval augmented generation.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? RAGAS: Automated evaluation of retrieval augmented generation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.344258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:b68ae456552c6a115e11653d39bcbfa4c07dda57cae3164d4bef9a34cff24d25

Observation 6233498b-3c8d-4fac-9938-7f12b5a55f94 · outbound

This paper cites Are we on the right way to assessing LLM-as-a-judge?.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Are we on the right way to assessing LLM-as-a-judge?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.778593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:22944a8c669167c0722fb5524f4d539cf5f3d5058cea6795e2cd750162af628c

Observation 0caa09fe-b545-48a3-9d4b-96c4ce779b96 · outbound

This paper cites Deep Research Bench: Evaluating AI Web Research Agents.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Deep Research Bench: Evaluating AI Web Research Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.859311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:b323a52fd9f0ab0e570b76dd354ab46ac6665224e7e3da0c26ac28bf30b823e8

Observation 318d24b2-8816-4ef6-a277-f0cb637581eb · outbound

This paper cites Enabling large language models to generate text with citations.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Enabling large language models to generate text with citations

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.342503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:907ac85581890d3bed8f6fc9a5a9f38d1665439524750af446237275d8b2244f

Observation da12364b-f7d3-4649-a6eb-a79592244c2b · outbound

This paper cites Gemma 3 Technical Report.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Gemma 3 Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:18:11.849298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:eac2a434093639cd7b438449499d5987b4ef707642ced014c1e5f24ee047d254

Observation 04325a0e-fa2d-4036-89a5-bc55b54321e7 · outbound

This paper cites Gemini 2.0 is now available to everyone.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Gemini 2.0 is now available to everyone

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.338763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:4ceafc770bb9968f80398249f7f1b749a868b644050745e6881825a8ef7d5722

Observation bcc17993-8406-43cf-b13b-d84b23233eeb · outbound

This paper cites Gemini 2.5 flash is now in preview.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Gemini 2.5 flash is now in preview

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.340710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:83342552078f127eaf117cb44c4180721e93c5fdd31194dcdffd7e52b562bc70

Observation 194b3c1d-b2f0-4459-8932-45a987bb5a55 · outbound

This paper cites Gemini 3.1 pro model card.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Gemini 3.1 pro model card

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.337057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:d30911714b65de9741a297d8c4aa5e2d81b9af88b8ebf2603f8a864dec01dd10

Observation 5383023c-2a4f-4f3b-83b2-1000cbb0c944 · outbound

This paper cites The Llama 3 Herd of Models.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? The Llama 3 Herd of Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:18:11.855294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:d0fe330b38b635d66c7dc19dab80e0496a81d843d7bda397354c866dc778f5f2

Observation dca70804-cc33-4b7b-8515-941b3337bb5d · outbound

This paper cites DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-14T02:21:26.356765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:23dab90bbfa25dbdfdc4dc88142c5ae012fc30da66dfb225dcef5dc4aebabc07

Observation 81b5bc79-ac7d-486d-8f14-d30d92cc335f · outbound

This paper cites Step-DeepResearch technical report.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Step-DeepResearch technical report

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.852269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:87d44e639fb5c2811fa51923af9291306e39cfea1dfdd4509a04b690e5df1031

Observation 3898ed0f-35ae-4cc8-b26b-8a8d5e8b9f09 · outbound

This paper cites MetaTool benchmark for large language models: Deciding whether to use tools and which to use.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? MetaTool benchmark for large language models: Deciding whether to use tools and which to use

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.335363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:154a75d40accb761d73e44fd1f59fb90f047f18f251791a5fb4c9f0000de39f6

Observation 0f680140-27ca-4a5d-a314-fc12344d72fe · outbound

This paper cites Hwang, Varsha Kishore, Amanpreet Singh, Dany Haddad, Aakanksha Naik, Malachi Hamada, Jonathan Bragg, Mike D’Arcy, Daniel S.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Hwang, Varsha Kishore, Amanpreet Singh, Dany Haddad, Aakanksha Naik, Malachi Hamada, Jonathan Bragg, Mike D’Arcy, Daniel S

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.769175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:c57a8bf5c47ce2a45b662caf6b3f4aef78a12a25f99bf880f5dad039a5a1c276

Observation 6abafd53-aa75-4eb9-8c8a-fe8315b24a6a · outbound

This paper cites Toolscan: A benchmark for characterizing errors in tool-use LLMs.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Toolscan: A benchmark for characterizing errors in tool-use LLMs

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.333548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:1767c15982faefe230e1a0998a53351d09d2d5fc5f916e08eb85e164dcc9dc5b

Observation 922621ac-c323-491b-aef0-811d3aa9a7ba · outbound

This paper cites Deepwidesearch: Benchmarking depth and width in agentic information seeking.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Deepwidesearch: Benchmarking depth and width in agentic information seeking

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.830405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:569a29df37e63cf6ff05d3c8d0c2047025db65580f9fede5060162f54f2b089d

Observation d087d9cb-616d-4dd8-aeae-6c3e5e20ec50 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.436930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:464f4e8362a20e42fd4949a11430a0e4aee3c0490372a263f200a9e13c4de5a8

Observation 63874281-696b-42b9-ba4b-30048d32597d · outbound

This paper cites ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.803436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:067fa2bf3a22fb1b76bd6bcfdc0cb01924d7869c43966809f7355803bdda313d

Observation 3e09ee99-d4ad-46ba-9b71-fb408140785a · outbound

This paper cites VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.800374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:7dc3f0e2f74aa2a1c9e96233f31fd0e61af968b71fb39665f863bbb474da7b57

Observation 14b5fa7a-51fe-4f1d-94dc-e7de63b735fd · outbound

This paper cites G- Eval: NLG evaluation using GPT-4 with better human alignment.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? G- Eval: NLG evaluation using GPT-4 with better human alignment

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.434253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:a09cac3a64b04d89c8ffb401e18e959a5b41b2ded8d650c4ffe2b7b15eaf6b42

Observation a7836ef6-5a2f-4b99-92ea-09925b0210ce · outbound

This paper cites ReIFE: Re-evaluating instruction-following evaluation.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ReIFE: Re-evaluating instruction-following evaluation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.439050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:1003863e4fdd0b8f8b2a0007ef9cb3d144d0dcea2466acda791cab8bdfac1c6f

Observation b4e25498-b6a4-464b-9e3f-20a859819fe4 · outbound

This paper cites In: Zong, C., Xia, F., Li, W., Navigli, R.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? In: Zong, C., Xia, F., Li, W., Navigli, R

Reference 28

Resolution
malformed identifier
doi_truncated, observed 2026-05-20T10:18:11.401397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:ab8aa59acba1b68d8d90a994acb6967a00ff20b3d79a3ae95e911ff5f95ac162

Observation 649f98e2-0ca0-4ada-ab91-a782de3a5971 · outbound

This paper cites ToolSandbox: A stateful, conversational, interactive evaluation benchmark for LLM tool use capabilities.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ToolSandbox: A stateful, conversational, interactive evaluation benchmark for LLM tool use capabilities

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.393784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:2baa9bbaa06fcd051aee1aae0bdaaf8b6f8c59ecea4b0f36b40237198ce29430

Observation aa31df63-7925-4dd9-a33e-9601a3395dde · outbound

This paper cites Agentrewardbench: Evaluating automatic evaluations of web agent trajectories.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Agentrewardbench: Evaluating automatic evaluations of web agent trajectories

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.832940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:090bdf1321e3f928e682d89726434736d74a046ddb3a6aa2a27668bb530f9182

Observation 1ba4f070-f8b4-4409-98f8-0a29e4918be6 · outbound

This paper cites Smith, Hannaneh Hajishirzi, and Nathan Lambert.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Smith, Hannaneh Hajishirzi, and Nathan Lambert

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.395994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:7da5f4f18e3fae6d6743d0733df3c22ad3c5a182aa6e7de7598c4cef2273a621

Observation 97044d86-e57e-4aa4-b6f0-68159b8ed96d · outbound

This paper cites An expert schema for evaluating large language model errors in scholarly question-answering systems.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? An expert schema for evaluating large language model errors in scholarly question-answering systems

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.399602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:aee1dc05ab167767ccd05dabbee65a793abf0f1dadd274e082da83e33d6f121f

Observation 3322ce45-6ae3-4139-ad25-df4ee6620668 · outbound

This paper cites FActScore: Fine-grained atomic evaluation of factual precision in long form text generation.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? FActScore: Fine-grained atomic evaluation of factual precision in long form text generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.391379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:73ba404eac9dba63f5010f27c5cc1d5436566de666948af3727166e778867561

Observation 09522d14-4ec3-40e2-97c4-c19e784f5f1c · outbound

This paper cites the moon is made of marshmallows.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? the moon is made of marshmallows

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.387570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:e28c6f0e7989e150d4787f827103a6253860c8c22ec2df21d7064442a185c576

Observation fdaaacf6-2ddf-46c2-b030-58a7d2d6c1ec · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? WebGPT: Browser-assisted question-answering with human feedback

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:18:11.815874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:235da1831cf7f936f5b3dc717253a3f8b8775f59accdd08e41e31c01aa118f16

Observation ab0962c0-a14f-41e7-aaa3-c9c0ccb05c99 · outbound

This paper cites RAGTruth: A hallucination corpus for developing trustworthy retrieval- augmented language models.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? RAGTruth: A hallucination corpus for developing trustworthy retrieval- augmented language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.409584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:d2c7b91a22f52dde0217fe300a8351bebeef5ec95df7befaff41632f05f1daa7

Observation 969d2b44-2c8d-48aa-979c-faa1880548ec · outbound

This paper cites GPT-5 mini.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? GPT-5 mini

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.432131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:213fdb31488ebf59cdd288ff2de035bc1ff99f056f4c6eeefe70b2eb419fcc2c

Observation 92eb6bda-046e-4194-8ef5-d72d36a0c282 · outbound

This paper cites Introducing gpt-5.3-codex.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Introducing gpt-5.3-codex

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.403453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:73d9362afc119164dd2ceace1289558387af620ce125f3a3b0803401c1597693

Observation 68e52a06-4b6c-47a3-8ff6-3ad435c5f18a · outbound

This paper cites Gpt-5.4 model.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Gpt-5.4 model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.407373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:e350984f1554b9d33fff519356a50bfd39ad6297b42142b43d3f5758b23d2800

Observation 912bad08-d1f3-4ca0-ae3a-ed7cf5b280f3 · outbound

This paper cites Gonzalez.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Gonzalez

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.405338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:9c277cf3e6ff601fcf2aa08754e92c3916d2943e8177fae22104e321d92a813f

Observation 5721c5c7-9218-4214-8127-a7b59ac5d1e8 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Direct preference optimization: Your language model is secretly a reward model

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.377510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:eb2534be211bc1d2e9d9acc48daff9324a63c87621113679e94a1785ce41baba

Observation c212035f-dee0-404a-bf18-d96c7ba46e55 · outbound

This paper cites an unresolved cited work.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:18:12.444985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:548be993e24dea492a8f730b91f36ddb3218b12eadfd6957eaa58450897449ce

Observation 56b562ed-bfb9-4e69-82cd-05316a156854 · outbound

This paper cites ARES: An automated evaluation framework for retrieval-augmented generation systems.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ARES: An automated evaluation framework for retrieval-augmented generation systems

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.451531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:1d1642885ab03ca2943d4373d36b71eb18f36b1fcd54317e5e153c8bf176d688

Observation 4eff1a9d-e06e-4578-b371-4aaf7ec3b895 · outbound

This paper cites Localizing and miti- gating errors in long-form question answering.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Localizing and miti- gating errors in long-form question answering

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.415492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:e41de6fcbc7d07f638339ff0301fcfa978ca13afd44ed008dc7a15706218fd11

Observation 47714c6d-5f04-41cf-ac81-9d31a29eba7d · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Toolformer: Language models can teach themselves to use tools

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.418108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:3aeda716e8d13a0374ad8ac666babbfb7132b02992c9d654031267802dc1adfd

Observation bf251382-38d7-4561-85dd-ea40cf98d7cc · outbound

This paper cites arXiv preprint arXiv:2509.22391 , year=.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? arXiv preprint arXiv:2509.22391 , year=

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:18:11.821552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:e82087f9083e7560376526ad5bd2a6e39326d3aa27e06e8c169bc457e72c96ad

Observation 3ac7d08b-3892-454b-ac3f-1decb6b4093f · outbound

This paper cites an unresolved cited work.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:18:12.366323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:281cb1d1d596b4c6ba8e017a55b430e85cd7075da991bfb82f5c27e90abc0aba

Observation f269b83c-f0bb-48b5-877f-e617acc62f6e · outbound

This paper cites ResearchRubrics : A benchmark of prompts and rubrics for evaluating deep research agents.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ResearchRubrics : A benchmark of prompts and rubrics for evaluating deep research agents

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.844903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:01ae60aeaad70dc511e43cd18a51f81a5cae17d5210682008e19230141ed0297

Observation 212c6f5e-7fe0-4b81-bab2-c7ece8e43cf7 · outbound

This paper cites Judgebench: A benchmark for evaluating LLM-based judges.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Judgebench: A benchmark for evaluating LLM-based judges

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.371973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:8fb6b3fe59a953412b144e0a1b051b6153b1374b6ba109776f610a9b73e22cf8

Observation da382f13-bc90-461b-b3d5-58b74139c138 · outbound

This paper cites Tongyi DeepResearch Technical Report.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Tongyi DeepResearch Technical Report

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:18:11.794452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:2ea1f9c74dbbef7f5b98f495747c90bb17e79e20bc2dd2c10d059ff0ebd98af0

Observation 427f3014-7067-40f6-ae4b-62e6b35c2a06 · outbound

This paper cites DeepResearchEval: An automated framework for deep research task construction and agentic evaluation.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? DeepResearchEval: An automated framework for deep research task construction and agentic evaluation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.835926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:bdc83e9bc9ead3fec925d7a4fcb44e8958603f8e2f060aefedaea6cb15c720c7

Observation 4e4aad11-32d8-4b65-a560-297f5ee5b0d3 · outbound

This paper cites Long-form factuality in large language models.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Long-form factuality in large language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.375517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:f6817a46e50818d79e893b0baf6f6451ef99ed7df9ebbf0ae0d437da538e42a5

Observation 6d507140-ac43-4703-9b41-7327e287f8c5 · outbound

This paper cites Qwen3 Technical Report.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Qwen3 Technical Report

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:18:11.841322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:dbc574f1615dd02ae4d659fe1855a222af3976516184c56b3c07c041209ac91c

Observation c5d36093-ef6e-449b-a5c2-01b396888bb9 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? React: Synergizing reasoning and acting in language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.381046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:7c3de55e6c1de1f6120d3bb7f57e9285f16e431a164b119c2430503232a7f201

Observation d7327b84-da56-4402-b4b6-b38b382311ad · outbound

This paper cites Yao, Y ., Wang, Y ., Zhang, Y ., Lu, Y ., Gu, T., Li, L., Zhao, D., Wu, K., Wang, H., Nie, P., Teng, Y ., and Wang, Y.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Yao, Y ., Wang, Y ., Zhang, Y ., Lu, Y ., Gu, T., Li, L., Zhao, D., Wu, K., Wang, H., Nie, P., Teng, Y ., and Wang, Y

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:18:11.806131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:949a651cbb4ba10af2db329b54f52dbff9311ba7a3ac6dd9f32aa4b50db7bdd1

Observation 2a731a01-b628-4c2a-a7d0-b4ba7f035e81 · outbound

This paper cites MiroEval: Benchmarking multimodal deep research agents in process and outcome.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? MiroEval: Benchmarking multimodal deep research agents in process and outcome

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.812625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:5b98b898a23d3ae15c4c4c4bf2c885e297a4b053742d5669330bc39f595bae0a

Observation e252c4b4-7021-41c4-84a1-9964e21c1868 · outbound

This paper cites Yifei, Allen Chang, Chaitanya Malaviya, and Mark Yatskar.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Yifei, Allen Chang, Chaitanya Malaviya, and Mark Yatskar

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.384073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:f50ce61aece906f5e84cdc00c478019654f996e087e371988e23bd59c618bd54

Observation f560a9ef-d05c-4437-b517-584a48670f8e · outbound

This paper cites Yifei, Allen Chang, Chaitanya Malaviya, and Mark Yatskar.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Yifei, Allen Chang, Chaitanya Malaviya, and Mark Yatskar

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.791303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:9fdfc44917ff4f1c1048b487338f34083b16c75f4b9dc8e59b4127fad66e17fd

Observation 4c37ec1c-6c41-47c5-bb82-651c89a329da · outbound

This paper cites Automatic evaluation of attribution by large language models.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Automatic evaluation of attribution by large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.401587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:0771b842ed049bd421d7ed3c6b131fec9ad942a01d020c9657ea6081b7c7889f

Observation b73c0e94-e94b-4087-ba06-fe4aa8320fe3 · outbound

This paper cites Evaluating Large Language Models at Evaluating Instruction Following.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Evaluating Large Language Models at Evaluating Instruction Following

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.809600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:4bdb4902dce9d15e07d0a8bb81031f010db4f53c104df4ae5e2898184d2e4aca

Observation 8f500d94-9707-444d-a5ef-0b0d76f3b977 · outbound

This paper cites Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-26T02:03:08.823370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:11d65d73ec4033e4d7ca3a20db3d07d9c40a313b7a6a9350d2bbea0f5f190bf5

Observation 2dc813d6-d429-420a-8d3b-35eba1a23db2 · outbound

This paper cites SRR-Judge: Step-level rating and refinement for enhancing search-integrated reasoning in search agents.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? SRR-Judge: Step-level rating and refinement for enhancing search-integrated reasoning in search agents

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.797516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:bdfbc214dd25e26fdc389514d3f250362fe5cb307271a96ba7d52c3ff6ca7eb4

Observation ef89c5f7-10bd-4e41-9d91-679a9d9b9e6e · outbound

This paper cites LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.772163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:039f292c47b9b4b96bfaaf7a5608ddace601649c6262223be6a38c0ffea938c2

Observation 8d90eb97-92ab-4e9d-9027-44b79ad5dd7b · outbound

This paper cites Chaining the evidence: Robust reinforcement learning for deep search agents with citation-aware rubric rewards.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Chaining the evidence: Robust reinforcement learning for deep search agents with citation-aware rubric rewards

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.775536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:837253f9cc5e2cfa5d67827dad8458ef09495700057706722f22109cea6e3166

Observation f0575d26-c00e-42ab-afd2-5e1ef1aeb063 · outbound

This paper cites ToolBeHonest: A multi-level hallucination diagnostic benchmark for tool-augmented large language models.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ToolBeHonest: A multi-level hallucination diagnostic benchmark for tool-augmented large language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.347961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:5c70ec8e1adc3cd89cf82b7000a34f12a6368ee8cddf651f67e6f2d5f4bb7b08

Observation 8c846851-0a1e-454c-8505-e648e08f14d8 · outbound

This paper cites Judging LLM-as-a-judge with MT-Bench and chatbot arena.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Judging LLM-as-a-judge with MT-Bench and chatbot arena

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.364556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:c4b38902f8c957448fc0d648bfd5537bfca8b56f4a74f35457bfb80b79bb3c06

Observation 30e59f45-05f2-4ec3-9ba0-76165526b25e · outbound

This paper cites ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.827795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:079f936ffe507b56264b5988be515057e263ccc19b92d17d7d8badfedec31584

Observation 5a1d8425-0a16-4cfa-b107-8bc06dd9f7ff · outbound

This paper cites Evaluating judges as evaluators: The JETTS benchmark of LLM-as-judges as test-time scaling evaluators.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Evaluating judges as evaluators: The JETTS benchmark of LLM-as-judges as test-time scaling evaluators

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.351898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:0b9bf1bca27f0ee6d3bf9eed7915c0dfdf555675d182c1855dda400af0a7cde9

Observation 2b9aa100-1956-4a75-b1db-1596db2c9d91 · outbound

This paper cites verbose database queries correlate with null results.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? verbose database queries correlate with null results

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:18:11.824969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:ebfe5a9bd08e37d2feb2357a581bbc18c5b74e14d85e5cecc22131161270996c

Observation e7158c4f-48d6-4b03-9c90-ba437f82154e · outbound

This paper cites coherence.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? coherence

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.456288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:8bc9500750a1edb13c85cb07bc3ff5668fe1a4a374aaa56655e196cb1e336bab

Observation 82bcd1fe-3c08-46b5-a2c4-a6e7b8c8b29e · outbound

This paper cites an unresolved cited work.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:18:12.353489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:3fff7127d8f9741b83effd26f8b914d1a85727ac56daa206c578b1f5aee6df78

Observation 4f120e13-9101-4a27-8dba-9969294aad0b · outbound

This paper cites an unresolved cited work.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:18:12.355185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:f98a5cb4c0001e34b31e83c1e74eeb465f82ac14dedc3cb82d2d46f93f809b29

Observation 1ef56faa-67bd-47fa-97a7-89579e9d75d5 · outbound

This paper cites an unresolved cited work.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:18:12.356819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:1e354cbf9c4f7c3ab8160e7fe667ed41b70f24fbf29de42b44f6c6a864c5dc44

Observation 3d956fd0-8ce5-468d-adae-e7a766f5efc9 · outbound

This paper cites an unresolved cited work.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:18:12.368066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:4f3829f695b30385194274dc743b3b6443ae6871371d85ee81c194668f85c928

Observation aa61c76c-9b30-4974-9dd6-8eae63d94220 · outbound

This paper cites an unresolved cited work.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:18:12.373685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:0344441f1a147bc25a29e892e54ba26c5e49f7b5ba521a4c902e4cd9176c5e74

Observation 968d2f76-2d18-46d1-920b-3a525e7bdb4c · outbound

This paper cites an unresolved cited work.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:18:12.427282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:8e3726d5a1336f03ecd9f60478a7596c4be931b69b165b47c8b99a5b0ea0e8a6

Observation fe88269b-071e-40c9-bd60-de959bae05aa · outbound

This paper cites an unresolved cited work.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:18:12.446816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:cd43abf5e89db858d69e146ca60f110a54e89d938e39d05f064079e779160d63

Observation 934cfd80-e1cc-4e0b-974c-4a79709d4bf3 · outbound

This paper cites an unresolved cited work.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:18:12.448611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:dbf329bfa431464ff4b7c2a4758e4d5718eb7adec08524668ba0526d0393f3d0

Observation f6a8d7db-9e8f-4de9-ba69-c0dd5ceac4dc · outbound

This paper cites an unresolved cited work.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:18:12.453579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:02cdd74a86df1f03845d3e56048435318be328e03ffddf175c857703c29d871d

Observation 0bace29a-5f93-4538-b101-8e04483d3f72 · outbound

This paper cites an unresolved cited work.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-05-20T10:18:12.443094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:c7d0150de3c4ebbd4355adfa44b21512e7031ba0d36a0d299ef300839c956249

Observation 7c6deaad-030e-42f9-8074-78d589c73673 · outbound

This paper cites appropriate human control.

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents? appropriate human control

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T10:18:12.441002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:15:37.280308Z digest=sha256:bb98381049809e5da1233de39c530dc575f33879717c7dac94268260bce11d41

Pith citing papers

Observation 77399758-2e3c-4b3a-bd63-8061acd33a9e · inbound

Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG cites this paper.

Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T10:22:55.558755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:22:55.558755Z digest=sha256:67f9c1a73ff46eb182a026b573d957198331e3f721174c79ea36b4fd579e0f3c

Observation e182a7d9-0c82-410c-96fe-4efa8f5a5b55 · inbound

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents cites this paper.

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T15:12:55.635905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:12:55.635905Z digest=sha256:5cf75b021cc08f88096abc4351b0077be5f1eb8692b30550b556ae3b9349faeb

Observation 55fbe9c5-8aad-44a1-9f4b-1665da83e741 · inbound

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models cites this paper.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:10:34.811346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T13:10:34.573818Z digest=sha256:41984257bf9aeb688c344b4c4cdbeb2f71225e8dcd68b445adecf6178f418939