Pith. sign in

Paper Citation Record · LEDGER

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2605.20061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.20061 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T05:35:45.084011Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T22:51:52.859785Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact27
  • verified fuzzy16
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 012ea2ec-3a8e-4547-aaf9-aac28ec6acaf · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.623982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:598ec0b2bb430a359d4aabecc704c4e5330f5ee4a98b4236a6a7158ada7c154a

Observation dd1f9a8d-ad26-4961-a91f-ad0d90bc6ffc · outbound

This paper cites Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.595551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:7b31f04fb1c4287baacb432dca7bb1f2b9089ce9613d1c3bf375837a9ae2a29d

Observation 8312bc98-f50d-4c52-9068-4cd3ade9f260 · outbound

This paper cites Exploration by Random Network Distillation.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Exploration by Random Network Distillation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.576948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:c9677e3d21de08c1addbdf797976ee1096e2e02e2d5f09ca8c2d4ddbf055ba1a

Observation 96854874-9cc5-4dd8-8337-aaa084eb7e79 · outbound

This paper cites arXiv preprint arXiv:2511.16108(2025).

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents arXiv preprint arXiv:2511.16108(2025)

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.629513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:2d39b1bde4fee121948e2a679ac969b4dde398c2531e80157f8e2045a3cfc3ca

Observation 7f9a707d-2739-4f66-a25d-cfedcc0afaa5 · outbound

This paper cites Sparse2dense: A keypoint-driven generative framework for human video compression and vertex prediction.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Sparse2dense: A keypoint-driven generative framework for human video compression and vertex prediction

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.637540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:003269e88894eec2ef751036878c579a972ac309346df148ad73233981895966

Observation fe5798ea-0ed8-40ee-bd67-1950eebb2d83 · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.569721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:5650c8ae5bab8f7c8dce28ff53ed4a4fbd47951378a1f61a2c892e79383b09b8

Observation fea07252-93ec-4980-b5e9-f72545c22899 · outbound

This paper cites In-Place Feedback: Reliable Refinement for Multi-Turn Expert-LLM Collaboration.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents In-Place Feedback: Reliable Refinement for Multi-Turn Expert-LLM Collaboration

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:56.900644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:921be2129dbb28dc07262fe4578177f03100bfd10a0a6795f57fb48fabd57fc5

Observation 3c5d387c-8159-4439-8748-2ae3e97f5b7b · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Process Reinforcement through Implicit Rewards

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.597087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:73b1cb8007aa6f0b7cc19115b253e182043f0bcd1129e12d793d1c41ca7b5e53

Observation 31fa0f1d-f338-4d9d-99d8-9a2c52ff711a · outbound

This paper cites Causal-Guided Active Learning for Debiasing Large Language Models.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Causal-Guided Active Learning for Debiasing Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.609691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:f80703e1dc282cd35f37a4eb0b9fd8f212417cdcfe8892ef8df0a929b00cff7d

Observation d6ee478c-0d33-46d2-b74f-66c82d0dfd3a · outbound

This paper cites UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.556220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:4e721370c0a1c7607cef5f0af24359a0e59dcc23f783789d1fa262386d0c2e0f

Observation 37cda70b-f966-4fed-82e0-ff2e5503d418 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.561954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:55ab4b3bc8ac32768cffa56cc62c7ec318e2b9e948bcd62288924d24aa51161f

Observation 9b88eeb2-e5d9-444f-82b7-04f9cc4a0ca4 · outbound

This paper cites Reward shaping to mitigate reward hacking in rlhf.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Reward shaping to mitigate reward hacking in rlhf

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.615287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:eb184e4ffbcb91225dd483da8ae0fe7972ff84cf76d54212913694a85a803926

Observation eb0fd076-7e1f-4a4f-b105-2640a98e096c · outbound

This paper cites A survey on llm-as-a-judge.The Innovation.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents A survey on llm-as-a-judge.The Innovation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.592049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:cd5096bbde25f83e546c89ea16614b2154012c17c71ae4cf383f418acab7d10e

Observation ccad7092-8913-47ea-95d8-154beb4e1c4b · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.580201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:3370907629d6ff08af7cfe0199c30b59d0d3539f34925ce420bc8f6e9e4e6edc

Observation 536ea10e-a0fe-474f-afa9-d6eb8a971c87 · outbound

This paper cites Retrieval augmented language model pre-training.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Retrieval augmented language model pre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.588522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:125901bd421acd1821cae3d6f3550c2029f1d5f4d78b0ccb867530a7360f2978

Observation 2eef5082-7dca-4653-9404-39ef37fe7458 · outbound

This paper cites Reason- ing with language model is planning with world model.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Reason- ing with language model is planning with world model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.556080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:af85fcebd3e9c4a18440ab2a60ed4cffd6619fbcb451a3a11ac07a3c14730ff9

Observation c5a900b7-0a5f-4dde-b407-57e7a7d7707e · outbound

This paper cites Sample-Efficient Multi-Round Generative Data Augmentation for Long-Tail Instance Segmentation.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Sample-Efficient Multi-Round Generative Data Augmentation for Long-Tail Instance Segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.574964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:005dd5af6e8c58829479193d4c9ffdb5250c4f414d0fd47258fddba93ae9070f

Observation 5dc5fd74-4f94-469e-99f6-62e8a1d63375 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:38:05.656716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:1306007afd1277d373babb68499165444e1b189491a65db66ae83ee3afe84d70

Observation 72c95792-12f3-42aa-ac25-33d7645719b5 · outbound

This paper cites ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-05T02:15:32.426400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:aa754a83976ce0114997247ebb388f7e700d64f6c87113a0650155f21f048031

Observation c57c37ad-9882-4127-b0e7-4c4e70b8d627 · outbound

This paper cites Let’s verify step by step.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Let’s verify step by step

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.559369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:d1fd4f1d5b440e4a008a415b057c7d18593b05dc17b65fd79a1b7f3b72577ea6

Observation dbdc0305-8223-4178-8ff1-cade817a5fc2 · outbound

This paper cites Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.476364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:fcc394a013397465e933d58bdd1d1a1aebc6b00b4f797d3ff4c248e8d248bf19

Observation 1d16f286-4699-4d31-9965-58a52a70115c · outbound

This paper cites Agentic reinforcement learning with implicit step rewards.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Agentic reinforcement learning with implicit step rewards

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.616554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:6297ef06a382c81d6f45ff157370e1d3ff28b82aa5226e6c63fe54261c0e6ebf

Observation dde12716-7245-4f66-8666-2fff40799569 · outbound

This paper cites Symbolic and subsymbolic geoai: Geospatial knowledge graphs and spatially explicit machine learning.Trans.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Symbolic and subsymbolic geoai: Geospatial knowledge graphs and spatially explicit machine learning.Trans

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.584526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:7e14e8e5935b9e8fd6ffdedf7060db913f75c3bb436eb487193621baf98411ac

Observation 66ff8920-b637-422e-9234-891c994c2a83 · outbound

This paper cites Augmented Language Models: a Survey.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Augmented Language Models: a Survey

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.663138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:fb36da606fc52ea9204b2fabb01d583aeb82f2fa6ea7144522f4f8cef0c741a7

Observation 607163d8-842d-4eca-8978-8dd92539bc64 · outbound

This paper cites Uncertainty quantification in llm agents: Foundations, emerging challenges, and opportunities.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Uncertainty quantification in llm agents: Foundations, emerging challenges, and opportunities

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.619793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:1a82e2ab626ce447db80f10ed85f697ac4b70dd45847e92a4edb7d4b10d5cee3

Observation dead8b95-81ca-4233-9490-4607066f35e9 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.628346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:8071c8cba04b27db2d8716c991ae46e577499517fc82b6b7ba13b166c3d2be5c

Observation 0f1eefb2-5760-4d62-8312-4efae2694fa6 · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:42:06.579174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:64bf63b0cd4f128917aacd9c13608d70adc96a9afb3f0ca8830d621fe39833d2

Observation b069b3c8-cb86-418f-94b9-e3117a4981e0 · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.549697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:c5432222f8ec07d22e8f4266236f4a934be1a028f24eba9142e0be5c0511ce4c

Observation 74f3a646-af73-4093-9624-7a67d59c5b8a · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents ToolRL: Reward is All Tool Learning Needs

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.643313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:5385ca702b754202ade02e6d0d23c39632a8381f2d6d37235ccb788226554c6b

Observation cad7fdda-c1b9-451c-b078-33c44222fcaf · outbound

This paper cites A possibility for implementing curiosity and boredom in model-building neural controllers.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents A possibility for implementing curiosity and boredom in model-building neural controllers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.611174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:28e214bebd681d5ced29eecc1ba1969ea5354fa5856152256427fd4a26bd778b

Observation 30ef472a-92d5-4cc3-97a4-902a8226c098 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.589891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:9c9a5078553b35b2316214aae48c5acfc07c34835c0d60840c06ab3839cabf7c

Observation 7a71902c-632d-4363-9a7e-e9bd2d163928 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.543201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:5965e103d60a6509febe6ac9b817d03d5e5b2902eff7f6abb623e942e1991f42

Observation c2ca204c-6241-40e0-afc0-1f11f80191cb · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.525894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:b2469f45bb9dcedfc1cbf6f07b4a32c4a8070c724748a35274ff4a2b8c44ab44

Observation d180fe40-4558-4732-a493-9eefd30a594c · outbound

This paper cites Monte carlo tree search: A review of recent modifications and applications.Artificial Intelligence Review, 56(3):2497–2562.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Monte carlo tree search: A review of recent modifications and applications.Artificial Intelligence Review, 56(3):2497–2562

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.607710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:cdc378870fd354f53031da14efa58ce4c80a62eff18d147585d2dd96b23da90f

Observation e890ac0e-3c93-4f75-a7f8-2a656467d5f8 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Solving math word problems with process- and outcome-based feedback

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.649948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:c444b20e5e42b50aef25a169c69e1f4a5f750857b132b6821919affe588bfc2a

Observation 71b714c2-ce8d-4902-b461-f6355dafef08 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.583384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:91a44d2589745269f698b1f2b012e24eef46d149dde275b085ee3fa8721c7b34

Observation 8f6170ba-ec94-4679-9880-94e7e6a78c23 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.463079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:0f0a83b881eecb8120db96027463c52aa1535f0c99676bafd245799822982c6f

Observation b59c1e80-b311-4087-a119-9dbbe090a57a · outbound

This paper cites AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.470067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:9395884f234107517e4dd1eaa6f61d1e378d546cd25f2cc479eedc5182ac7a37

Observation 17c7d3ff-d50a-45d5-868b-bfe81a49f6fa · outbound

This paper cites Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.519281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:61e4d543f81bc962ff2a31acfc9b354e7e234e5dbc0a518ed304c114e3123235

Observation 62165ce6-d375-4406-99a3-6bcd1478482b · outbound

This paper cites Qwen2.5 Technical Report.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Qwen2.5 Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.603208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:9fc32a39debe8ac90e650a019e2929c79026480a881f115a222aa48e9a64a1c4

Observation fc52fa50-7a6d-4d03-9159-74fa85313a7b · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.603719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:1a20659779a8ce7247fdfe58ef9a746812e2565a2fe0a8a2a5efe81eb6e50c58

Observation 87b15eba-f30b-4a0f-aa93-ef773b398c8e · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:24.599605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:3bed26d1725774c175b8435362edb023328692c13a6542860359b03c083aaae2

Observation d1c10b34-de1c-4fc3-874a-c8745195a556 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.537641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:7284c69f98c361ef5bd964336c1c5bccce10c8198ecba51227c416a7a3308702

Observation 4ac2a3c3-7399-4a83-aacd-b7eedd783c24 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.483436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:ed6ac5633e49ec3c635d49b335a4b048fd1cd681ad862c574db4f803a5b60e49

Observation efe056c7-d2ce-416f-b3d3-3d9bc68915a3 · outbound

This paper cites Enhancing agentic rl with progressive reward shaping and value-based sampling policy optimization.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Enhancing agentic rl with progressive reward shaping and value-based sampling policy optimization

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:38:05.496454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:1fbd065ee276693d1ffcdc9fb02d0258c22dcd56eddd6b70a81367047ba4acb3

Pith citing papers

Observation be776092-685d-4b9a-8c28-90ea9b07680a · inbound

From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents cites this paper.

From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-31T22:51:52.859785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:51:52.859785Z digest=sha256:e3aaa534e608df1df03f0b1ea2aa143d770fce07d9c7921823ec00dd9d36bce6