Pith. sign in

Paper Citation Record · LEDGER

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

As of 11 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2608.07437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07437 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:39:03.692877Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved39
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02ee2bb2-7fe8-4c06-b2ee-ce4641c13f46 · outbound

This paper cites Stability analysis of fluid flows using Lagrangian Perturbation Theory (LPT): application to the plane Couette flow.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Stability analysis of fluid flows using Lagrangian Perturbation Theory (LPT): application to the plane Couette flow

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.400189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.400189Z digest=sha256:fcd6b0aedfea9873f3cce0ba331a99cc834b600e98066693edb09781252ed7a7

Observation f153cbbe-5435-4fe6-bc70-a6848ac07d3b · outbound

This paper cites Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.405648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.405648Z digest=sha256:1899f504627f51838f24d91bd93bd547859ed12bdbd975eb009b318eb457723b

Observation 0d367f4e-3feb-4514-9909-6a82962ea663 · outbound

This paper cites Deepanalyze: Agentic large language models for autonomous data science.arXiv preprint arXiv:2510.16872, 2025.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Deepanalyze: Agentic large language models for autonomous data science.arXiv preprint arXiv:2510.16872, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.410767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.410767Z digest=sha256:3ab894dc911b946ab9fe7803ead04efda622affae744dca9be2f701943b0d312

Observation bb3088c0-b3fa-4f21-b3ac-6010955f896f · outbound

This paper cites Data interpreter: An llm agent for data science.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Data interpreter: An llm agent for data science

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.752551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.415655Z digest=sha256:f946a053e5143f5aca91a05f47f0118d59ff84debffeb45df54bb8243d48b859

Observation 3688e365-d990-471c-8239-cf1a285070ff · outbound

This paper cites Infiagent-dabench: Evaluating agents on data analysis tasks,.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Infiagent-dabench: Evaluating agents on data analysis tasks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.736770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.420634Z digest=sha256:f27f6340308f53dcdc9425c2873acfcdc599ab2e121525c78455dde3b76b9d1f

Observation 644faab7-ad14-4c37-b046-38a7b75ab743 · outbound

This paper cites DABstep: Data Agent Benchmark for Multi-step Reasoning.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DABstep: Data Agent Benchmark for Multi-step Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.430946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.430946Z digest=sha256:09483acb95f5824e6ad874efef146c68a2b56157f26bffe4536e5c64d2fc5d3e

Observation d847d5b0-9d18-44da-930c-7c95e024f0b1 · outbound

This paper cites DA-code: Agent data science code gen- eration benchmark for large language models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DA-code: Agent data science code gen- eration benchmark for large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.436223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.436223Z digest=sha256:76b6025f4b450d9397161e39452ecde7f267156cfa93980fe3f0c37e397930a0

Observation 0a10aadc-bd13-43fb-ab13-3168d2574294 · outbound

This paper cites Are large language models good statisticians?Advances in Neural Information Processing Systems, 37:62697–62731, 2024.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Are large language models good statisticians?Advances in Neural Information Processing Systems, 37:62697–62731, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.720970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.441812Z digest=sha256:fbc18ef54ebe8864679b4576236471520caba93471c3cd8294e37bfedf857dd8

Observation 9ff204ad-749c-48ef-a97b-da696b13364a · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ReAct: Synergizing Reasoning and Acting in Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.446928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.446928Z digest=sha256:8b13ca8aa3dee297e116b3ef3f1a9eb6d27f80221821d1561a1fab295b2e6d7d

Observation 4ebded3c-7ce2-497e-a8ad-eeff6bf9b079 · outbound

This paper cites Ds-1000: A natural and reliable benchmark for 12 data science code generation.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Ds-1000: A natural and reliable benchmark for 12 data science code generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.704281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.452234Z digest=sha256:572ac144b0a40fe960d07e1eeeeb94208f570ba7de2932aac65c35d697360d25

Observation c9f45c4b-e365-44ce-94b6-d9d21dcce73d · outbound

This paper cites InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.457120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.457120Z digest=sha256:e65dd36ed8d414e05f1d7aafdf653cd7690ab8023e9414d8a891be8c834ae1c6

Observation b7b4fe64-9111-4aac-8e71-5a507dd9ac22 · outbound

This paper cites DataSciBench: An LLM Agent Benchmark for Data Science.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DataSciBench: An LLM Agent Benchmark for Data Science

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.461634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.461634Z digest=sha256:bdcf7dc0a8cc122bf95d1224435031e83e8df1d42d1947cce5f3923aabf67b29

Observation c1557f8b-cbf6-4f32-a05c-31576141a4ab · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.466801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.466801Z digest=sha256:36d91293252a1d988686b7b33344145f8df7015bbb773ffd5309a0a8bd298ccb

Observation 9a176489-9d48-4993-bd21-04ab319deee1 · outbound

This paper cites Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.471639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.471639Z digest=sha256:7eca616cefd2cb7fd474edeaa2356dd8fa6d1e3f7517df0244fd56b324d86563

Observation 8837324d-564e-4c35-8c3a-42f6175dc1af · outbound

This paper cites IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.476630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.476630Z digest=sha256:be2dde1b69cf8e6da8467fab4a568ac8de090b948610defc9a453e98daffcfec

Observation 30637093-fa4e-4b87-9771-5515245c4045 · outbound

This paper cites Fact or fiction: Verifying scientific claims.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Fact or fiction: Verifying scientific claims

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.481558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.481558Z digest=sha256:c193ebdba3c45992c0f7dcff4fafb6a5959156f894d62302132407abe4d193f6

Observation d8348684-9385-4de9-a54d-6444071f1631 · outbound

This paper cites Sciclaimhunt: A large dataset for evidence-based scientific claim verification.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Sciclaimhunt: A large dataset for evidence-based scientific claim verification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.678309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.486856Z digest=sha256:9412878ade7757d1b20a53070bcbfe8e470499579d0630beb3f8e12478da9bcd

Observation 967c7018-a6b1-49aa-9c25-6282fef5ebaf · outbound

This paper cites Musciclaims: Multimodal scientific claim verifi- cation.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Musciclaims: Multimodal scientific claim verifi- cation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.663704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.491749Z digest=sha256:745a11b9394d046e342941dbaa56658235b24ff26670a920812c627f21ba7d7e

Observation 75554491-30c5-42cb-bed5-39805595e3d7 · outbound

This paper cites Investi- gating the reproducibility of the social and behavioural sciences.Nature, 652(8108):126–134, 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Investi- gating the reproducibility of the social and behavioural sciences.Nature, 652(8108):126–134, 2026

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.648515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.496445Z digest=sha256:f76b195705c27d98abf3f606397462d6f7a1ef92e2b0b731be4f4ea7c009b60c

Observation 2e9418d9-cd96-49f7-b8f7-bbdd9163863e · outbound

This paper cites Towards end-to-end automation of ai research.Nature, 651(8107):914–919, 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Towards end-to-end automation of ai research.Nature, 651(8107):914–919, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.500807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.500807Z digest=sha256:1ff683ba1dcf075cdd909a211e0ce2317be9026e98680ea5ffe1b0cf079b624f

Observation bb557469-8e70-4498-822e-5276a4c060ee · outbound

This paper cites AI-Researcher: Autonomous Scientific Innovation.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing AI-Researcher: Autonomous Scientific Innovation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.505323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.505323Z digest=sha256:aa76793ef09255d810bbb0fbd551d813f1ef0a3654cbea90bc4db5257d48c103

Observation 2e96c553-6fcf-4e65-aa12-39d281943fde · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.509944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.509944Z digest=sha256:f96474913a447b14c8aabc2ea57b91ce841a57ea70218ebf6137cf607b184199

Observation 7097c805-7d63-451c-8122-a5b98befa13a · outbound

This paper cites Development economics field experiments (dfeep).

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Development economics field experiments (dfeep)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.620868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.514880Z digest=sha256:1e62aea7927899e62ffe53ed238d8d6cc749a3c1818226c2c49e15bf72d60737

Observation 7af7ec38-d793-45e9-ade6-72a95d998e44 · outbound

This paper cites The cbio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data.Cancer discovery, 2(5):401–404, 2012.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing The cbio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data.Cancer discovery, 2(5):401–404, 2012

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.605005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.519472Z digest=sha256:eacd78c449244f6e24478ccecfa96f4946ca827ca34e043db0db8fd253155113

Observation ba341f71-2d39-4448-8fd2-9d69c49fc373 · outbound

This paper cites Integrative analysis of complex cancer genomics and clinical profiles using the cbioportal.Science signal- ing, 6(269):pl1–pl1, 2013.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Integrative analysis of complex cancer genomics and clinical profiles using the cbioportal.Science signal- ing, 6(269):pl1–pl1, 2013

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.588856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.524397Z digest=sha256:f7546d2919a53eef307ff2880b7ebf4e22e7a82b15414c7bde488a1c738292f6

Observation 5b5ffc14-a414-42a7-b78d-11cf301e6476 · outbound

This paper cites BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.528813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.528813Z digest=sha256:315872b138f1dc1726d16de38a367c3bdff23cb48b78160d21aa8ff667837294

Observation 06e28d36-989a-43de-885b-022acebdf441 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.573011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.533331Z digest=sha256:eb3f6eec2ac2bf8aa0f0655333683f4139b91466ff51be535c43b4a4ad86b09b

Observation 3509a304-b336-450a-a2e1-85062991a107 · outbound

This paper cites Introducing claude sonnet 4.6.https://www.anthropic.com/news/ claude-sonnet-4-6, February 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Introducing claude sonnet 4.6.https://www.anthropic.com/news/ claude-sonnet-4-6, February 2026

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.556765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.537708Z digest=sha256:2c49e1202421798104bd9823b6854321ce102645b364b16af9764fd39cc577a9

Observation d71be0a8-203a-4c60-a02d-d6a67a1c3ce9 · outbound

This paper cites Executable code actions elicit better llm agents.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Executable code actions elicit better llm agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.542160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.542160Z digest=sha256:efd4f5a4c12bdfb39e40d7e6a424d9c915e9a50071b6729bd0a0064d1fc21c78

Observation 50d4aaaa-401c-49ae-be49-6e26115c50f2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.546532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.546532Z digest=sha256:2e5559ed9f337af0af307738abcfc0529cf3cfe1bed0dde898397cfdb3d146de

Observation da19215c-905d-4a00-b296-589b5f48acdf · outbound

This paper cites Qwen2.5-Coder Technical Report.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Qwen2.5-Coder Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.550965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.550965Z digest=sha256:56be9a28e7738e422e7a4db0bc6df0579aa938a24bb017779eb2d85ab818aa96

Observation 26ec706c-2bf8-4b48-8181-2fbbf363f3cb · outbound

This paper cites Introducing gpt-5.4.https://openai.com/index/introducing-gpt-5-4/, March.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Introducing gpt-5.4.https://openai.com/index/introducing-gpt-5-4/, March

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.555768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.555768Z digest=sha256:08fdf9079c5c2f015bc11ec287ff7869dcd8215961b8d314aac3e8e902a21578

Observation 71e73ab9-e116-4aed-86a6-0400ae6b77c6 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Deepseek-v4: Towards highly efficient million-token context intelligence

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.512901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.564690Z digest=sha256:90fc0e3c687936310b46e17b3bacd2459c6e984334d03835b7ebcc745cbac953

Observation 8c37b627-0242-4fa6-a6dc-129a9b438cbd · outbound

This paper cites Introducing gpt-oss.https://openai.com/index/introducing-gpt-oss/, 2025.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Introducing gpt-oss.https://openai.com/index/introducing-gpt-oss/, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.496607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.569100Z digest=sha256:46eeb3b493604b1611a6d6bc3aa5972a5b7957d9a5d8c779c6db0302a59d811b

Observation 6402fc39-cc60-48fe-9bb9-27fbc343b7b8 · outbound

This paper cites Qwen3-coder-30b-a3b-instruct.https://huggingface.co/Qwen/ Qwen3-Coder-30B-A3B-Instruct, 2025.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Qwen3-coder-30b-a3b-instruct.https://huggingface.co/Qwen/ Qwen3-Coder-30B-A3B-Instruct, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.480535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.573317Z digest=sha256:2efb64301be137c5f60ae1738368219ca7b7a217ae8190442df2bbb96d774e8f

Observation df0caddc-0ea6-49f8-8c9c-80b35a7704a1 · outbound

This paper cites Qwen3 Technical Report.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Qwen3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.577579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.577579Z digest=sha256:7136eb7d1eeccaac4fddb811ae804849715a0902cdef3248e407ee86ecd3a588

Observation 38312006-f686-462e-ae24-0e77b00ab9de · outbound

This paper cites Scaling generalist data- analytic agents, 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Scaling generalist data- analytic agents, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.582193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.582193Z digest=sha256:404a40975d6b5bd8d54f20e6a1fbf37591f1544db56790cc69c6403d3fa6cd3a

Observation c8b03bf1-d1ee-4461-a21f-d5d7b661b266 · outbound

This paper cites Accessed: 2026-05-06.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Accessed: 2026-05-06

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.465596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.586686Z digest=sha256:f3889bf0ee7a37c226d17b7e3dc7217bc9b445bb2a38c8e83375b0681a7b7b41

Observation e7de4fcd-e169-406b-91ea-3254483e4b57 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Self-Refine: Iterative Refinement with Self-Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.591340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.591340Z digest=sha256:a7343463aa0fb1bf8a0aed35b8a0b314e7e6876b1332e17406d99e9f0d87987a

Observation c44e2879-0050-436a-96b8-0081c40c9409 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in neural informa- tion processing systems, 36:8634–8652, 2023.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Reflexion: Language agents with verbal reinforcement learning.Advances in neural informa- tion processing systems, 36:8634–8652, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.450651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.596337Z digest=sha256:4ff38f13dbbbb73e9516a254f2dee5962f1fa05cb345a1bf05188afad661e814

Observation d7ab10fe-dee9-4e4b-be02-276bb43b150b · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.600566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.600566Z digest=sha256:fed2643d6a8a6bfbcf8fdb21c5c71a60e1ca1272bbda314d24574d9914b1a4bc

Observation c65ef131-db0a-42f3-ae75-2e46db8a5d69 · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated soft- ware engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Swe-agent: Agent-computer interfaces enable automated soft- ware engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.434898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.605328Z digest=sha256:40f1b599a7d81894d0bc94e33aa1820310d9f9a43ccfdf0d52dd2ba54795eae4

Observation 000e14c4-7303-4d23-bcd9-1596e22aecd7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Proximal Policy Optimization Algorithms

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.609725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.609725Z digest=sha256:1e72add681462106bf5ac1289dcce829ef4ba2e7311f8026a9ff3e2d0198f161

Observation c488ede0-f512-4ae6-9050-0a36aa648941 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.614383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.614383Z digest=sha256:b3e0c6d28de1988b477a26b5d1aa15c62fef4bf32b3efcdceb930b44d699f367

Observation aaa26ca6-40be-4162-acf0-908c80fce4d3 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.618872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.618872Z digest=sha256:87c901855dbde48e49cc0341c67cf2fda674fa24f3ce0da56bd95a37cc109e98

Observation ba795ca9-c6bf-41fb-af74-c838cfb26782 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.623408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.623408Z digest=sha256:574b6b41504c6fd9e3d8806502dbd83371e32ef483c5eb583aab3ceb11e7139e

Observation 53f52013-4298-4d13-9832-06e4463ba865 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.628209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.628209Z digest=sha256:7794c5fa56c2ca9da25a876c5abd9cf8c5dcaffd58992cf2e0d3f224a2435061

Observation 3c61b79c-030c-42a0-88ff-cbac19d0b4fb · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ToRL: Scaling Tool-Integrated RL

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.633077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.633077Z digest=sha256:2db82c1802cebce5a6b386218da8f2fd2a0150d701e83a04a6fb851101c4c315

Observation e6ce7ab3-ca53-44a0-960d-951f42109cd4 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.637815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.637815Z digest=sha256:a086a2e5871c6ce7dcff79b109472f6b855b070b56c1f5cdede01c5303800cd6

Observation 3a6b4ba9-845b-4056-be59-28eb4a5da243 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ToolRL: Reward is All Tool Learning Needs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.642514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.642514Z digest=sha256:dcd3212bec99e055f627bafc0d358d7178d76fa3996f094ed9c3b29e47f13417

Observation ca4ee6f6-5d92-472b-95e6-2c5810b604f1 · outbound

This paper cites Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.647525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.647525Z digest=sha256:314fdcb796236988a85d305c70bf1d1b54f877aa73df0de43ef650ce7f103ee2

Observation 65e2fbaa-1e39-42b7-8721-101cae37630c · outbound

This paper cites Effects of cognitive behavioral therapy and cash transfers on older persons living alone in india: a randomized trial.Annals of internal medicine, 176(5):632–641, 2023.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Effects of cognitive behavioral therapy and cash transfers on older persons living alone in india: a randomized trial.Annals of internal medicine, 176(5):632–641, 2023

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.399794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.652822Z digest=sha256:3cec3da4ec67991c692d165fc4630f6a8ce1727a5f44d65cd0fa270d03fa5ef4

Observation 5b30023b-5725-4082-80ba-ac74832d7b4f · outbound

This paper cites Genomic characterization of metastatic patterns from prospective clinical sequenc- ing of 25,000 patients.Cell, 185(3):563–575, 2022.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Genomic characterization of metastatic patterns from prospective clinical sequenc- ing of 25,000 patients.Cell, 185(3):563–575, 2022

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.384842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.657534Z digest=sha256:87bfd1c2279eeda2cf2d226b10234c06796eafcb3e136297b0deb6de85d502a7

Observation 5254a881-e8a1-488e-9cd5-2826d602f635 · outbound

This paper cites The support prognostic model: Objective estimates of survival for seriously ill hospitalized adults.Annals of internal medicine, 122(3):191–203, 1995.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing The support prognostic model: Objective estimates of survival for seriously ill hospitalized adults.Annals of internal medicine, 122(3):191–203, 1995

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.369442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.662098Z digest=sha256:fbc9ebdd0d6f72a80c185ab72ce5849ec1491af697fe16918b6fa9e3b92586da

Observation f15cf1cd-b341-43da-b393-5e0ccc2186a6 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.667845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.667845Z digest=sha256:9f86d1c7418c8b49551d5df6c4980ebaab564518edf47f6d007b27603b1ba41e

Observation 8c18f6e8-df20-41d5-b0d1-6c30635501c6 · outbound

This paper cites claim”: “drug improves patient outcome.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing claim”: “drug improves patient outcome

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.353710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.672908Z digest=sha256:e6be6ed54c8ae42c382be63966cd6d0ba9678ad9beb0d42eff677d33502b7182

Observation adbe0b92-7e51-427f-8525-b05ec970620f · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.336736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.678589Z digest=sha256:e53a732a1abdc34bb6696c793ecef2a0bcddc2d6baa989c871125ff947145025

Observation ddf3d3cd-3bf0-42c1-b65d-4b12483816bc · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.322049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.683802Z digest=sha256:1c8cf2beeaf5d0d41dfd65a1a098f6f7bed76c33abed66f216c8dfd73b82ae8e

Observation 93bd0286-2b0f-47e0-8e6b-13789a89d806 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.306832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.688424Z digest=sha256:a8628d3425de2cadc914113ec80e49f5e92d71cc812fc2e88e117a1c2bbc2df7

Observation 507ce3fc-5327-40f6-b266-2abf02d2f7e1 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.290784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.692877Z digest=sha256:c00cfdc53a715394663cdee70fe5f1ad963e8049cc7d7342a565e408b4a75044

Observation 163b0ec8-d481-4081-8989-a5e184270d91 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 2026

Resolution
parse uncertain
no resolver link, observed 2026-08-10T04:39:03.560376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.560376Z digest=sha256:b7315fb59e28237cc0d81a6087a1af1959f9601abcbadfaabc514175ae20b3a1

Pith citing papers

No inbound Pith citation observations are available.