Pith. sign in

Paper Citation Record · LEDGER

The Hallucination Tax of Reinforcement Finetuning

As of 8 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 9 inbound Pith citation observations for arXiv:2505.13988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13988 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:44:36.534742Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:52:50.299350Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T18:57:16.598004Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee8a17e6-6a68-4f1f-a790-a77f4d267511 · outbound

This paper cites online" 'onlinestring :=.

The Hallucination Tax of Reinforcement Finetuning online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:29.905006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:29.905006Z digest=sha256:a4305ce90a39429c9c3d63b5f8305cc7b245e087bb121d2664de475e31fb70ab

Observation 1bd65997-4777-4fe7-9e09-495e03bcfbc3 · outbound

This paper cites write newline.

The Hallucination Tax of Reinforcement Finetuning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.081158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.081158Z digest=sha256:34bb911e02219d8ca290894a861a814ae22320282257e54c6df0655dd442b1b0

Observation 995a819f-5fbc-4bb5-90c7-4cf4c8a0ac33 · outbound

This paper cites Phi-4-reasoning Technical Report.

The Hallucination Tax of Reinforcement Finetuning Phi-4-reasoning Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.194934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.194934Z digest=sha256:e5080779fcdcd50b73c60fd68aa979b06dbef01eafef1065728b4ad90f4e2d47

Observation 6c6d2e74-fa0b-4a8b-bb16-c419cb0eefa0 · outbound

This paper cites HalluLens: LLM Hallucination Benchmark.

The Hallucination Tax of Reinforcement Finetuning HalluLens: LLM Hallucination Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.304843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.304843Z digest=sha256:e2ebe2d470d22c2acace691a2b7685e42c7c322072b8bf00511d19ce04fe0fda

Observation 24326b6b-640a-4b9a-abe6-962ebb645b7e · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.416483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.416483Z digest=sha256:ad332bd9f430293ed0e17e36e67d31fe41a386684b9a64766e42c98824e9b9a5

Observation 5c1703b3-86d1-46ea-9da9-93516346b420 · outbound

This paper cites Mitigating Open-Vocabulary Caption Hallucinations.

The Hallucination Tax of Reinforcement Finetuning Mitigating Open-Vocabulary Caption Hallucinations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.575619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.575619Z digest=sha256:8976594e5b35ddf5483751e4534a7cc4ce8a2a07fc470d56b3d92afc9d15d833

Observation e0c669cd-5eb0-4592-82c1-95bfa02b08bf · outbound

This paper cites Evaluating Hallucinations in Chinese Large Language Models.

The Hallucination Tax of Reinforcement Finetuning Evaluating Hallucinations in Chinese Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.675099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.675099Z digest=sha256:f6e22c2a84ece40a0b852c7ca493759ede83e791333d7e2b7a44dcc45589b27f

Observation 41939ca4-a260-4759-91f5-d40a2cfe6b20 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

The Hallucination Tax of Reinforcement Finetuning Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.756069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.756069Z digest=sha256:23470c2bd960574ac763625af0c65c89bd3adaf93ea6a42f68bab4c2edade24c

Observation 3a56fac1-4f09-493f-ad1f-d88b1ad02e51 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

The Hallucination Tax of Reinforcement Finetuning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.864841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.864841Z digest=sha256:93896609caa9751e78b6e1964a66e73252aceda9f85c09e64caed2a7ef5002c7

Observation 71483c38-9e7a-49ef-8260-ff5294e1f711 · outbound

This paper cites Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?.

The Hallucination Tax of Reinforcement Finetuning Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:30.995190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:30.995190Z digest=sha256:1365297fab2607ad057189829aa3217ae223b6f4548915f65109e02a2cade268

Observation 854c0351-2a20-4b30-92cb-d939b26418a0 · outbound

This paper cites The Llama 3 Herd of Models.

The Hallucination Tax of Reinforcement Finetuning The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.106531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.106531Z digest=sha256:518d80919b92b161028a611c8c1fd782618cf306557b8b006b98854a1a0246ff

Observation ce49eb88-0866-4122-8767-c1fdbd6c4c21 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

The Hallucination Tax of Reinforcement Finetuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.222493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.222493Z digest=sha256:5a1097230b0955b69e0ab185ed57d11eda0457821a67506832fb8180fa263f83

Observation 819102f4-c5fd-4cb0-b98d-60b315192ed4 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

The Hallucination Tax of Reinforcement Finetuning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.355068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.355068Z digest=sha256:8ba451259b3ecd08ba516bf561525b2d57c3746b25233a4001b31e471acf79b9

Observation f5f2b607-eb5a-428d-ac1d-fb36a23b78d3 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.447643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.447643Z digest=sha256:ebae256ecde11dda6b3d0c8901fc6a415cf1dc83ea80b532f1eb09cda919b13a

Observation 8f7481b6-dca4-4af9-aaed-449f5323eaa3 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

The Hallucination Tax of Reinforcement Finetuning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.518305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.518305Z digest=sha256:d5cad27324c0160a70db9ae7efa17eb854cc6cb31826bb2a5732db57ae30f69d

Observation 84052823-48f1-4594-8298-fe0688dfcfe8 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

The Hallucination Tax of Reinforcement Finetuning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:31.690376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:31.690376Z digest=sha256:c39b0ace4a98198f82157ba3816b02e90f00966b58b7247ab47ee42355f37ee9

Observation 2b37d6ff-c0f1-4c99-a109-81665178032f · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:43.373694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:44:31.804578Z digest=sha256:c8d14ead296115530eff584459a50bf646703b80528821acca1d36afb2886984

Observation 25972b8f-464b-49e8-94e8-f8f5994bcaff · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

The Hallucination Tax of Reinforcement Finetuning O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.005515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.005515Z digest=sha256:408b6416b982ef3c80dfb12ab3f2791be632773aa9d3863b2b08accfe17df28b

Observation 5f1ffac1-a515-450a-8696-fd8c930e9a5e · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.124833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.124833Z digest=sha256:a5b304680aa7dd04729599a88c1d95e38cfb2abf26746319995614659a69444d

Observation 911ec246-e3ca-4ef7-a54b-253350311b84 · outbound

This paper cites The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language Models.

The Hallucination Tax of Reinforcement Finetuning The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.201925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.201925Z digest=sha256:9f53a93ee956aa6c3cc2bbdee2f7d2988f633487bf05726e9c01cf4ac03195a0

Observation 27b411fc-cbc0-4d0d-9643-2eeee438d1af · outbound

This paper cites Treble Counterfactual VLMs: A Causal Approach to Hallucination.

The Hallucination Tax of Reinforcement Finetuning Treble Counterfactual VLMs: A Causal Approach to Hallucination

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:44:37.888022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:44:32.260799Z digest=sha256:6a4c5eb4599f8c20ec35efa61d2e139f4c708ffe54ed60a895332f509403ba0b

Observation 9418ddc7-8284-4590-81cc-d8bc205438cb · outbound

This paper cites LIMR: Less is More for RL Scaling.

The Hallucination Tax of Reinforcement Finetuning LIMR: Less is More for RL Scaling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.385679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.385679Z digest=sha256:6321d56c28ac4c928e8bef5cd3cdf5a4f1d74aee8fbf69654bbcfeb4b0951d0c

Observation bc76dd61-de76-4e95-b56a-b6224cd025a2 · outbound

This paper cites Let's Verify Step by Step.

The Hallucination Tax of Reinforcement Finetuning Let's Verify Step by Step

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.534874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.534874Z digest=sha256:311d09a9397066b7acfae1d863eb6f7fc97a94a3b54fbb3cc2bdcc26e81f5c4e

Observation 711946dc-cdd6-42db-a1c6-d3ed63c0ed2f · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.681966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.681966Z digest=sha256:221a889e49466774465365557836f6c48d4874141a0f26271be8c97517d2cc73

Observation 152c4b32-31b7-4b31-b8a7-ce1516fce7a9 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

The Hallucination Tax of Reinforcement Finetuning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:44:43.042890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:44:32.794896Z digest=sha256:e8b926ee45cded3d5738abda89c5571badad86026f7b185d01decdd3ec3ca3c9

Observation 655dcc39-5a30-44d3-ab7d-6c4fd97b7cc5 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:32.954751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:32.954751Z digest=sha256:a21627baee94c765b6dd0dfbc9f2cab074aa36f462fcfcb25c02e083307c93f9

Observation 10c551ec-e18d-4b78-987f-de8e7b0ceff8 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:42.634345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:44:33.126414Z digest=sha256:b0c06cb4596aea3787232fb7ba0a353668ac782060421581455dd7a5fc82896e

Observation 2ee8666a-b3cf-42e7-b1d0-e4cff21c07ea · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:42.366454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:44:33.324958Z digest=sha256:7534d36a9749527320007e85234b0b00ba4d104a4fe75bb2eb07c1cf09e4c93c

Observation e825bd42-c4fd-4479-9e05-a2f9c9a63638 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:33.545500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:33.545500Z digest=sha256:900655bad2430a3e16fee748637fcfbbf17151dc8d07fbf19a5418759416a494

Observation 4e638485-7ee6-4076-8934-97bbd59cd17a · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:41.964745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:44:33.628322Z digest=sha256:804e910ce0896db3e71722e6bdcb409577a24553696f09606bb4e50d59ff1377

Observation be302502-a318-4dc2-87ee-44644080a929 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Hallucination Tax of Reinforcement Finetuning Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:33.782472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:33.782472Z digest=sha256:148921d1578564aeaa4f674d25abd9f9deee99cbf56e3f67bc3e5527bdf4c620

Observation 7c226e53-6050-4b90-b6c2-edf442f7a5d1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

The Hallucination Tax of Reinforcement Finetuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:33.894983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:33.894983Z digest=sha256:fefb16d0993b2b6f46415b5352f9b9466084eea795a6c09dd05fccf5bf6ded73

Observation 16abeb0e-4769-4dee-bee0-62ab852816d5 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.007922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.007922Z digest=sha256:7ddbad9c4c688c6a920268d893b9b7c7a5db76b8865996d0e7c1a851637f607d

Observation b6932d08-7faf-4c4a-9d61-9ad113926933 · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

The Hallucination Tax of Reinforcement Finetuning Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.170709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.170709Z digest=sha256:882a22768ed65fb8cdd1d6cfe24a506766db801a588527a7a97ab31e0421d35e

Observation d8527608-d5de-4952-b276-8553a0ab47df · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.272336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.272336Z digest=sha256:f6f11aad1d472b3348444f4fe4618531da87f507ad15e8c0c4d6e43099c9831f

Observation ed215d06-63d0-4e1a-96ed-001d416d4174 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:41.595034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:44:34.448645Z digest=sha256:61abf86712bf75b20911946b3279f2a0abb26510e8ff67a46ae9c9de8ba6d2e9

Observation 1e62e25e-fcd5-4aa5-9bbe-597e3bacba13 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

The Hallucination Tax of Reinforcement Finetuning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.610778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.610778Z digest=sha256:d5b17c0efc9fb654d6ee0815a6614de675a8cda3ec0b9e092610959b56ec2eb1

Observation 1e57963e-b949-4603-86b2-f160291ca989 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:41.284869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:44:34.775008Z digest=sha256:07c9dc7b1e14244dec2cec5a3dda59676052ec2c2a0328501360f98de3627fed

Observation fdd7360a-d90f-4de7-905a-48d9044a6cfb · outbound

This paper cites A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models.

The Hallucination Tax of Reinforcement Finetuning A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:34.912984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:34.912984Z digest=sha256:9c2d15b15af4d238aafc495f908213289c8c3f655beee5e07f9bd6d121af0896

Observation 1e86e968-b904-4aa8-af69-2222fd6b1be5 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

The Hallucination Tax of Reinforcement Finetuning AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.037495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.037495Z digest=sha256:53ac7053c72b463aca0546c651c99b820bc24ae975de09ba984f07038c97b18d

Observation 9e7e68fb-7fad-46bd-a097-8001bb4f84c5 · outbound

This paper cites Tina: Tiny Reasoning Models via LoRA.

The Hallucination Tax of Reinforcement Finetuning Tina: Tiny Reasoning Models via LoRA

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.167445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.167445Z digest=sha256:e2b7e72b8ab7796b1a44a30e33581daa579710172a654c6f7351d3d5f24183cf

Observation ec9d1ffe-a12c-470d-9379-96f40be417bc · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

The Hallucination Tax of Reinforcement Finetuning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.310902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.310902Z digest=sha256:a7d13fd986e93f0ec8aba88a4421e4a3a948f5cfb695ce4219c3a08de9d6260d

Observation 176734b8-7600-48fd-a725-72d93dc4c722 · outbound

This paper cites Simple synthetic data reduces sycophancy in large language models.

The Hallucination Tax of Reinforcement Finetuning Simple synthetic data reduces sycophancy in large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.430765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.430765Z digest=sha256:f018a2eabd99418b0fcb4def61ff66dbdfaa77333d9eb9f09aa8a6a331eb182b

Observation 9aebf6e4-c24a-47e1-9f0a-9503c9f6beb1 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:40.994636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:44:35.515063Z digest=sha256:b7f86ca40241bf296694bcc1ddf6560330b70f867b69a885fc5818a28e9ef434

Observation 007b1964-85c9-4886-b1c9-f9788c4ad213 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

The Hallucination Tax of Reinforcement Finetuning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.621729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.621729Z digest=sha256:bbd1942330522c275cac3cb8282d8625bba190b5100f6ff984589ea894c388bf

Observation 7be9aab1-c81d-42c3-9abd-b00308dcd272 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:40.628475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:44:35.719317Z digest=sha256:500c63e0a246ae5068914a1e9b412e8c538064eecb2e5df7fed15d87f1c51287

Observation 16848e90-3df9-41b2-bccf-5254ad9f571f · outbound

This paper cites Do Large Language Models Know What They Don't Know?.

The Hallucination Tax of Reinforcement Finetuning Do Large Language Models Know What They Don't Know?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.850344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.850344Z digest=sha256:890616bba0121ae3217ed5e71cd05c1852bee734c81c5b40f31ef1cb057a900e

Observation b778e018-1636-4f23-891f-dbaaaa36b509 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

The Hallucination Tax of Reinforcement Finetuning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.961914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.961914Z digest=sha256:77eadd5d563d3df9e516250293ef9309f2b81ca1b9d379d309b2641407f463d1

Observation 9a8df657-c924-449c-b6ab-22dd12a69bd1 · outbound

This paper cites an unresolved cited work.

The Hallucination Tax of Reinforcement Finetuning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:44:40.383837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:44:36.094973Z digest=sha256:dd7fd5528355df0249361e28f6ae0592afa911f386692186b39dbe41b55ef3c2

Observation 9268f459-e4da-4b2a-901c-79f32205026e · outbound

This paper cites How Language Model Hallucinations Can Snowball.

The Hallucination Tax of Reinforcement Finetuning How Language Model Hallucinations Can Snowball

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:36.213918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:36.213918Z digest=sha256:92018f103546fe8c03cab4814650212648a5d37b6b6d4e019ebae50fd7edbfe2

Observation f32edaa0-bd51-4cb7-81a7-533956c446c1 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

The Hallucination Tax of Reinforcement Finetuning Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:36.354899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:36.354899Z digest=sha256:b63c11dc02df7d8bcec61213e55c60473b8623a9b424414ab1b37539518709df

Observation 13000c1c-0526-4a67-bb6d-f74ffe031b72 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

The Hallucination Tax of Reinforcement Finetuning Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:36.534742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:36.534742Z digest=sha256:2f49292fdd2a6f4e983207eda885f1868e7aa771d551eee41d677f9356a42354

Pith citing papers

Observation f96d3bbd-0892-49ad-9c12-ce012dde8f0d · inbound

Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration cites this paper.

Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI Collaboration The Hallucination Tax of Reinforcement Finetuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T10:52:50.299350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:52:50.299350Z digest=sha256:5d23a771f4d5b5f137dec807c806e57df66c952637121350986e5318d4888ea1

Observation d1ed60c3-31ad-4d33-8b55-7c3267e55276 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:03.113610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:03.113610Z digest=sha256:d85487d03643bf34548a464005171b448da4169808a707bf12615886dc1d6251

Observation 1d18fcaa-1364-4c4f-9724-4d6f3d7debac · inbound

MoCo: A One-Stop Shop for Model Collaboration Research cites this paper.

MoCo: A One-Stop Shop for Model Collaboration Research The Hallucination Tax of Reinforcement Finetuning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:43.795988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:17:37.129753Z digest=sha256:d3479aa1d81858e9918342343aa0151f92f8f021e9f75fe9db9ba0629e2bf916

Observation 7a4542a9-ef81-404b-97c1-25a1b7813634 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:29.704183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:e99081336c5e5053652125f33d0f7255006ef3c5b6fc202d8666673d74856c3b

Observation 92e46cd5-8b6e-4d71-83a9-2633117e1cc6 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:16.560735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:b9419e42302a1fbd1444f2ef60ac60d336129f2ab6c6e724d29de052bd2654b3

Observation 1686a16f-1bbe-407e-b0bd-10c6fe6c2dc3 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models The Hallucination Tax of Reinforcement Finetuning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:02:40.737194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:531e7c96bd2c1a13fc349dcc0db20a0307ad9468ae786757d70f6b19347d763d

Observation d9620be2-8cc6-45fd-88ce-ff858571b6e4 · inbound

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information cites this paper.

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information The Hallucination Tax of Reinforcement Finetuning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.821929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T12:37:16.138816Z digest=sha256:5e5136f656116744462c34470864d0665a67aaa544097807f85b69645b3cfecf

Observation 74983e86-3310-4efc-aab7-7357ede6e45f · inbound

Scaling Participation in Modular AI Systems cites this paper.

Scaling Participation in Modular AI Systems The Hallucination Tax of Reinforcement Finetuning

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-02T18:57:16.599301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:49:27.042616Z digest=sha256:ed3fdbc86df317817c702381b78a0609f64ba047905d527bdbf5588f1ecb17c1

Observation a0a77b3d-b3f9-4a9d-8d75-9ebe41e5c81a · inbound

Mechanistic Attention Guidance for Agent Memory Refinement cites this paper.

Mechanistic Attention Guidance for Agent Memory Refinement The Hallucination Tax of Reinforcement Finetuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T17:30:40.711089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:30:40.711089Z digest=sha256:7bf82e6c86906c7fa5300fa275836baafdb4218f00d5c2c8b1e6fee7c5e16746