Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2504.15253.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:56:01.605280Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:49:38.244210Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 9718b30f-8129-4090-b209-9e0b43a91073 · inbound
Can You Trick the Grader? Adversarial Persuasion of LLM Judges Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa498442-b8c9-46d0-9379-96303932e058 · inbound
Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3531d6c-a718-48ad-b4dd-4bfeb6ed5f0a · inbound
Towards Agents That Know When They Don't Know: Uncertainty as a Control Signal for Structured Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2052e7dc-8e34-4643-960f-36bb1c65a4ba · inbound
On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 142f02d3-7d4c-4ff0-abbd-df389318e56e · inbound
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2d5591c-75b5-4c20-a51b-0e8e55a72838 · inbound
Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df095008-aa51-4e87-9f08-e3532340a4ae · inbound
Counsel: A Meta-Evaluation Dataset for Agentic Tasks Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d44d8c1f-a5a3-4d80-88a5-19b8ac73ceec · inbound
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf7f5c99-4b5e-4db2-9b3a-f92db4a927ac · inbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.