Pith. sign in

Paper Citation Record · LEDGER

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

As of 11 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2608.07437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07437 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:39:03.692877Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved39
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02ee2bb2-7fe8-4c06-b2ee-ce4641c13f46 · outbound

This paper cites Stability analysis of fluid flows using Lagrangian Perturbation Theory (LPT): application to the plane Couette flow.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Stability analysis of fluid flows using Lagrangian Perturbation Theory (LPT): application to the plane Couette flow

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.400189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.400189Z digest=sha256:26e7becd204feb2d1af2cf10eb53e06217d927ea7a1112ad3e795c8d55f07848

Observation f153cbbe-5435-4fe6-bc70-a6848ac07d3b · outbound

This paper cites Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Data-Copilot: Bridging Billions of Data and Humans with Autonomous Workflow

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.405648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.405648Z digest=sha256:1fff17d19ec04e96607ae5c0c854fbc63180287a55e2690b97ec8a17595e494d

Observation 0d367f4e-3feb-4514-9909-6a82962ea663 · outbound

This paper cites Deepanalyze: Agentic large language models for autonomous data science.arXiv preprint arXiv:2510.16872, 2025.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Deepanalyze: Agentic large language models for autonomous data science.arXiv preprint arXiv:2510.16872, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.410767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.410767Z digest=sha256:bd8abe16c1ece67c37b63e1b7823f5cb5f0fcecaf6320418f6b22da85c0e8fa1

Observation bb3088c0-b3fa-4f21-b3ac-6010955f896f · outbound

This paper cites Data interpreter: An llm agent for data science.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Data interpreter: An llm agent for data science

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.752551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.415655Z digest=sha256:4203e903f253aacbb8b3242f14d21a9fdaf7b2ee54db08a57a108f1803fa3fde

Observation 3688e365-d990-471c-8239-cf1a285070ff · outbound

This paper cites Infiagent-dabench: Evaluating agents on data analysis tasks,.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Infiagent-dabench: Evaluating agents on data analysis tasks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.736770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.420634Z digest=sha256:e7aa185dc2879e4b9532876f44ae11fd3958894d11516c418305140c22be0751

Observation 644faab7-ad14-4c37-b046-38a7b75ab743 · outbound

This paper cites DABstep: Data Agent Benchmark for Multi-step Reasoning.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DABstep: Data Agent Benchmark for Multi-step Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.430946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.430946Z digest=sha256:ddfbb5ded697137f10d7602ddb27c40ff44eae9a43831491442ae5f8f1039dcf

Observation d847d5b0-9d18-44da-930c-7c95e024f0b1 · outbound

This paper cites DA-code: Agent data science code gen- eration benchmark for large language models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DA-code: Agent data science code gen- eration benchmark for large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.436223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.436223Z digest=sha256:09489e2127957445b5a463bea257fad66b887dbe6d97ac4d6aab443f1020afa9

Observation 0a10aadc-bd13-43fb-ab13-3168d2574294 · outbound

This paper cites Are large language models good statisticians?Advances in Neural Information Processing Systems, 37:62697–62731, 2024.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Are large language models good statisticians?Advances in Neural Information Processing Systems, 37:62697–62731, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.720970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.441812Z digest=sha256:ff4db66ac5f651dde0c3505b30b6507ad1535e39ac804e3c27da5da51962bc82

Observation 9ff204ad-749c-48ef-a97b-da696b13364a · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ReAct: Synergizing Reasoning and Acting in Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.446928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.446928Z digest=sha256:feb9635af9b9ec799798f7afb729b297597330b5d08baf76c6efbf9cb14a3a70

Observation 4ebded3c-7ce2-497e-a8ad-eeff6bf9b079 · outbound

This paper cites Ds-1000: A natural and reliable benchmark for 12 data science code generation.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Ds-1000: A natural and reliable benchmark for 12 data science code generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.704281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.452234Z digest=sha256:f355734f04d8d93c02ffc3c8c016211d97f8bce1c3ddb3dbd8bee9e532914724

Observation c9f45c4b-e365-44ce-94b6-d9d21dcce73d · outbound

This paper cites InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.457120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.457120Z digest=sha256:981902b4b29a8980f2efefc6c9bed808bc2a2541a29ade89d6c4ca011aa9336b

Observation b7b4fe64-9111-4aac-8e71-5a507dd9ac22 · outbound

This paper cites DataSciBench: An LLM Agent Benchmark for Data Science.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DataSciBench: An LLM Agent Benchmark for Data Science

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.461634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.461634Z digest=sha256:f27bea4ea11852e5af5676010f0404cbccce053f84aeecbd59aad7c1b4c908a4

Observation c1557f8b-cbf6-4f32-a05c-31576141a4ab · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.466801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.466801Z digest=sha256:c06039d94a8fac6a352f93567b9a3873a64da8ada4137bf5e97ab563305cf25a

Observation 9a176489-9d48-4993-bd21-04ab319deee1 · outbound

This paper cites Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.471639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.471639Z digest=sha256:5845b253baa5ddb9558114cb4cd532f178b477208ea59706560c9cfc2cc9aaba

Observation 8837324d-564e-4c35-8c3a-42f6175dc1af · outbound

This paper cites IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.476630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.476630Z digest=sha256:d9dd2b91234a2fdea1bcccbc1c785ac131291c336ef663607e4c3ad441c8a638

Observation 30637093-fa4e-4b87-9771-5515245c4045 · outbound

This paper cites Fact or fiction: Verifying scientific claims.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Fact or fiction: Verifying scientific claims

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.481558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.481558Z digest=sha256:fdd524946fbdacc67f8fde6b0deed57562a62116ecf85c8893a6f78dd95953cc

Observation d8348684-9385-4de9-a54d-6444071f1631 · outbound

This paper cites Sciclaimhunt: A large dataset for evidence-based scientific claim verification.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Sciclaimhunt: A large dataset for evidence-based scientific claim verification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.678309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.486856Z digest=sha256:c8c6bf3e171b94890bbe4dbe441343a2d293c0f114a76ab2f25598bb93c03b19

Observation 967c7018-a6b1-49aa-9c25-6282fef5ebaf · outbound

This paper cites Musciclaims: Multimodal scientific claim verifi- cation.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Musciclaims: Multimodal scientific claim verifi- cation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.663704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.491749Z digest=sha256:c4621de7dc9ddc9a25abf0f3ab44abd30b6cba9ed5737106e36e30ac818d551a

Observation 75554491-30c5-42cb-bed5-39805595e3d7 · outbound

This paper cites Investi- gating the reproducibility of the social and behavioural sciences.Nature, 652(8108):126–134, 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Investi- gating the reproducibility of the social and behavioural sciences.Nature, 652(8108):126–134, 2026

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.648515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.496445Z digest=sha256:4828568e4defc8a62416e82628eccec04e6d3b269cc4836a5ba9a8e62add7b37

Observation 2e9418d9-cd96-49f7-b8f7-bbdd9163863e · outbound

This paper cites Towards end-to-end automation of ai research.Nature, 651(8107):914–919, 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Towards end-to-end automation of ai research.Nature, 651(8107):914–919, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.500807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.500807Z digest=sha256:00a757af71a1eeea9975390cf64eb6d317f8dc6b4315b08011458fdfdd077c47

Observation bb557469-8e70-4498-822e-5276a4c060ee · outbound

This paper cites AI-Researcher: Autonomous Scientific Innovation.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing AI-Researcher: Autonomous Scientific Innovation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.505323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.505323Z digest=sha256:5aea5efb59a3b3ffb8b58872bc9b10cea1bfe0e6738daab015a7ba8c712b98e2

Observation 2e96c553-6fcf-4e65-aa12-39d281943fde · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.509944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.509944Z digest=sha256:6f797b0cf98ee99a7c83f51ddd06a57c8b94a7d5aaad74ec786f950590cfad21

Observation 7097c805-7d63-451c-8122-a5b98befa13a · outbound

This paper cites Development economics field experiments (dfeep).

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Development economics field experiments (dfeep)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.620868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.514880Z digest=sha256:f5d1590879e8b627aef8bd22df2104016ec58ed666f017006a67f9622ec44e6a

Observation 7af7ec38-d793-45e9-ade6-72a95d998e44 · outbound

This paper cites The cbio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data.Cancer discovery, 2(5):401–404, 2012.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing The cbio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data.Cancer discovery, 2(5):401–404, 2012

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.605005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.519472Z digest=sha256:83aff38ffc4c523184e416cb50f2922077816bdb6587b4c4f2b48c1688e2f18e

Observation ba341f71-2d39-4448-8fd2-9d69c49fc373 · outbound

This paper cites Integrative analysis of complex cancer genomics and clinical profiles using the cbioportal.Science signal- ing, 6(269):pl1–pl1, 2013.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Integrative analysis of complex cancer genomics and clinical profiles using the cbioportal.Science signal- ing, 6(269):pl1–pl1, 2013

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.588856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.524397Z digest=sha256:870a6295e7ce944008a28cdfc14adebc9ca234ff38613b27f9d94fdcd6349c5e

Observation 5b5ffc14-a414-42a7-b78d-11cf301e6476 · outbound

This paper cites BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.528813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.528813Z digest=sha256:012b9979ac67a08b6e607196a6258441aeb2dbccc8c979f005be6f92f82d66a7

Observation 06e28d36-989a-43de-885b-022acebdf441 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.573011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.533331Z digest=sha256:74a94a47e166d2e4a2cb422032f464f2c0cd95f50eb5667defcb4fbd22fe8801

Observation 3509a304-b336-450a-a2e1-85062991a107 · outbound

This paper cites Introducing claude sonnet 4.6.https://www.anthropic.com/news/ claude-sonnet-4-6, February 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Introducing claude sonnet 4.6.https://www.anthropic.com/news/ claude-sonnet-4-6, February 2026

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.556765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.537708Z digest=sha256:b6d5cd3e2148b69eb5d1873da4cb82d1f676d0efa8c3dc6e5d674958e16a0fe9

Observation d71be0a8-203a-4c60-a02d-d6a67a1c3ce9 · outbound

This paper cites Executable code actions elicit better llm agents.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Executable code actions elicit better llm agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.542160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.542160Z digest=sha256:c43c20f78d6de1bb3276281e4026ca6f02e0336a2aada81e8fe4809d4ff0ba42

Observation 50d4aaaa-401c-49ae-be49-6e26115c50f2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.546532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.546532Z digest=sha256:0cae4887ab2da2447075b6926b07aee6a810a4b3217616b869d9a62eb859d6d9

Observation da19215c-905d-4a00-b296-589b5f48acdf · outbound

This paper cites Qwen2.5-Coder Technical Report.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Qwen2.5-Coder Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.550965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.550965Z digest=sha256:1b691447dc721f4cc3e9783082b6563a65a31325a817ae6a3321f859b4c49597

Observation 26ec706c-2bf8-4b48-8181-2fbbf363f3cb · outbound

This paper cites Introducing gpt-5.4.https://openai.com/index/introducing-gpt-5-4/, March.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Introducing gpt-5.4.https://openai.com/index/introducing-gpt-5-4/, March

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.555768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.555768Z digest=sha256:7c481d8eea238f5468cce8588d2add67f4f0d517169471bcd3766e5ab4938cf3

Observation 71e73ab9-e116-4aed-86a6-0400ae6b77c6 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Deepseek-v4: Towards highly efficient million-token context intelligence

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.512901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.564690Z digest=sha256:911c8b44b03f80be44630f7ae066dad9841a8c366048279724f453b3fd52054e

Observation 8c37b627-0242-4fa6-a6dc-129a9b438cbd · outbound

This paper cites Introducing gpt-oss.https://openai.com/index/introducing-gpt-oss/, 2025.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Introducing gpt-oss.https://openai.com/index/introducing-gpt-oss/, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.496607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.569100Z digest=sha256:ffabfc39b7eb1efec3613e9bc7a8eb21ad0d512a435fdd6fea5bf3f5a45f035b

Observation 6402fc39-cc60-48fe-9bb9-27fbc343b7b8 · outbound

This paper cites Qwen3-coder-30b-a3b-instruct.https://huggingface.co/Qwen/ Qwen3-Coder-30B-A3B-Instruct, 2025.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Qwen3-coder-30b-a3b-instruct.https://huggingface.co/Qwen/ Qwen3-Coder-30B-A3B-Instruct, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.480535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.573317Z digest=sha256:085a532faf2a853e2384c7650d42699e3b240fa21bba360463832db4f6f8c4ac

Observation df0caddc-0ea6-49f8-8c9c-80b35a7704a1 · outbound

This paper cites Qwen3 Technical Report.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Qwen3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.577579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.577579Z digest=sha256:a2560ea28a4267fe6319c29e500a8404649f071e31a37e529ba7c90b933569e9

Observation 38312006-f686-462e-ae24-0e77b00ab9de · outbound

This paper cites Scaling generalist data- analytic agents, 2026.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Scaling generalist data- analytic agents, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.582193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.582193Z digest=sha256:3c4d212b6591a51c7072d129815f3a504d42fba30292f31f1f6a98381abb9d70

Observation c8b03bf1-d1ee-4461-a21f-d5d7b661b266 · outbound

This paper cites Accessed: 2026-05-06.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Accessed: 2026-05-06

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.465596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.586686Z digest=sha256:9fed2f899ba91b00e13e1953500d4bbb1718f893b4b514288f46f1cfcb8392ec

Observation e7de4fcd-e169-406b-91ea-3254483e4b57 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Self-Refine: Iterative Refinement with Self-Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.591340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.591340Z digest=sha256:5398d039d99ae38c4bdceec7e628543b463bc31a0bee25c4f798ed04af888d33

Observation c44e2879-0050-436a-96b8-0081c40c9409 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in neural informa- tion processing systems, 36:8634–8652, 2023.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Reflexion: Language agents with verbal reinforcement learning.Advances in neural informa- tion processing systems, 36:8634–8652, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.450651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.596337Z digest=sha256:4f76cfe6eed3ced20d48e5d106f3123b9ced00bb8d2bb77a4986fc8eb300d959

Observation d7ab10fe-dee9-4e4b-be02-276bb43b150b · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.600566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.600566Z digest=sha256:5043ff4df6b0af3ca3ac7a6c8f62ccd9e26d70675923e8758043e27997f5f466

Observation c65ef131-db0a-42f3-ae75-2e46db8a5d69 · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated soft- ware engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Swe-agent: Agent-computer interfaces enable automated soft- ware engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.434898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.605328Z digest=sha256:d5ce2a0c4c5568bc30dbd3a1927ef2ccd4f6e14b3a280b1e9325b066743c8ec1

Observation 000e14c4-7303-4d23-bcd9-1596e22aecd7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Proximal Policy Optimization Algorithms

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.609725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.609725Z digest=sha256:941c2ed8a4e0ce1e697ac600c4650c99cf92d197bf699372bcaa666330b30198

Observation c488ede0-f512-4ae6-9050-0a36aa648941 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.614383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.614383Z digest=sha256:849ca2b658b9e5c3a81bdc27d5dbb76ef91abd69dfba7a5cc9011fe4bf5308c3

Observation aaa26ca6-40be-4162-acf0-908c80fce4d3 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.618872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.618872Z digest=sha256:ebca5f7612c7595d9d9f7daa3297a0ee606dc94d67ca6ce7251f01556df0acdf

Observation ba795ca9-c6bf-41fb-af74-c838cfb26782 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.623408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.623408Z digest=sha256:d0815c2f805357933d77b3853ecb53376959d73eabc002fdd290c453ad19c1c5

Observation 53f52013-4298-4d13-9832-06e4463ba865 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.628209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.628209Z digest=sha256:af81da07e13315f08168dcfeba72d1c11eed9c37d73dd2b15f43af5201224c9a

Observation 3c61b79c-030c-42a0-88ff-cbac19d0b4fb · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ToRL: Scaling Tool-Integrated RL

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.633077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.633077Z digest=sha256:8e44a5d7a667e158e2215c42355ab121da5c61a96afbb3c391282123a5a2061b

Observation e6ce7ab3-ca53-44a0-960d-951f42109cd4 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.637815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.637815Z digest=sha256:184e301ec6c008923497485892086a3df151b3a406bd5b37841c7d6559d23cf4

Observation 3a6b4ba9-845b-4056-be59-28eb4a5da243 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ToolRL: Reward is All Tool Learning Needs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.642514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.642514Z digest=sha256:e817b37cbb669574f92d32ba50cf176c76c90d1ae9b2c5136b53874d44a2324a

Observation ca4ee6f6-5d92-472b-95e6-2c5810b604f1 · outbound

This paper cites Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.647525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.647525Z digest=sha256:5aa72615415c4a4fad0ec54110f041985bd9eab3fca2a66db342fdcf3f95f929

Observation 65e2fbaa-1e39-42b7-8721-101cae37630c · outbound

This paper cites Effects of cognitive behavioral therapy and cash transfers on older persons living alone in india: a randomized trial.Annals of internal medicine, 176(5):632–641, 2023.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Effects of cognitive behavioral therapy and cash transfers on older persons living alone in india: a randomized trial.Annals of internal medicine, 176(5):632–641, 2023

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.399794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.652822Z digest=sha256:56d384290aec23fbd19396355aeaed430474712ee7bccba0e2510a2327b5b47b

Observation 5b30023b-5725-4082-80ba-ac74832d7b4f · outbound

This paper cites Genomic characterization of metastatic patterns from prospective clinical sequenc- ing of 25,000 patients.Cell, 185(3):563–575, 2022.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Genomic characterization of metastatic patterns from prospective clinical sequenc- ing of 25,000 patients.Cell, 185(3):563–575, 2022

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.384842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.657534Z digest=sha256:075393de29b72821a952b43c25b4b83b66c4da19529d85716640d6ae3bc7a578

Observation 5254a881-e8a1-488e-9cd5-2826d602f635 · outbound

This paper cites The support prognostic model: Objective estimates of survival for seriously ill hospitalized adults.Annals of internal medicine, 122(3):191–203, 1995.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing The support prognostic model: Objective estimates of survival for seriously ill hospitalized adults.Annals of internal medicine, 122(3):191–203, 1995

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.369442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.662098Z digest=sha256:61240d2a5af8ab4de8e384d87416dbb9d56fa7b4f18066f3d901821d60fd4163

Observation f15cf1cd-b341-43da-b393-5e0ccc2186a6 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.667845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.667845Z digest=sha256:7a0d33d66f1905e226fa3c6eb943a869c222da5bfc68198bbba7fe8975d03e15

Observation 8c18f6e8-df20-41d5-b0d1-6c30635501c6 · outbound

This paper cites claim”: “drug improves patient outcome.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing claim”: “drug improves patient outcome

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:39:04.353710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.672908Z digest=sha256:d84627e2c62b01c69b567bbdeefd2d214c5a8115daedb9c9d140ec5410225f0f

Observation adbe0b92-7e51-427f-8525-b05ec970620f · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.336736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.678589Z digest=sha256:ef1cca0ffe2ef144c8908b01c48f73463bee5cb4d619a2e41b3a37c496e0207e

Observation ddf3d3cd-3bf0-42c1-b65d-4b12483816bc · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.322049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.683802Z digest=sha256:951fc0619d354cdb5da150e9f176bfedc82d8d9724f11b34f8f7798b0f9848a1

Observation 93bd0286-2b0f-47e0-8e6b-13789a89d806 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.306832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.688424Z digest=sha256:652949f55a37eec61160dc864ca5defd5bac376a061dcbb1d49357adf5363196

Observation 507ce3fc-5327-40f6-b266-2abf02d2f7e1 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-10T04:39:04.290784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:39:03.692877Z digest=sha256:57dd1f36ef50915e525798be4b270699fb9419f15ab7aa6128d724c230763201

Observation 163b0ec8-d481-4081-8989-a5e184270d91 · outbound

This paper cites an unresolved cited work.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Unresolved cited work

Reference 2026

Resolution
parse uncertain
no resolver link, observed 2026-08-10T04:39:03.560376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.560376Z digest=sha256:f59dad11a407d9b965d02b6a85d60f55fa1ac3a3e89ae9e8494e838460eaf49b

Pith citing papers

No inbound Pith citation observations are available.