Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T05:35:45.084011Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2605.20061.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T05:35:45.084011Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T22:51:52.859785Z
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 012ea2ec-3a8e-4547-aaf9-aac28ec6acaf · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dd1f9a8d-ad26-4961-a91f-ad0d90bc6ffc · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8312bc98-f50d-4c52-9068-4cd3ade9f260 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Exploration by Random Network Distillation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 96854874-9cc5-4dd8-8337-aaa084eb7e79 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents arXiv preprint arXiv:2511.16108(2025)
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7f9a707d-2739-4f66-a25d-cfedcc0afaa5 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Sparse2dense: A keypoint-driven generative framework for human video compression and vertex prediction
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fe5798ea-0ed8-40ee-bd67-1950eebb2d83 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fea07252-93ec-4980-b5e9-f72545c22899 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents In-Place Feedback: Reliable Refinement for Multi-Turn Expert-LLM Collaboration
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3c5d387c-8159-4439-8748-2ae3e97f5b7b · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Process Reinforcement through Implicit Rewards
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 31fa0f1d-f338-4d9d-99d8-9a2c52ff711a · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Causal-Guided Active Learning for Debiasing Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d6ee478c-0d33-46d2-b74f-66c82d0dfd3a · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 37cda70b-f966-4fed-82e0-ff2e5503d418 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Group-in-Group Policy Optimization for LLM Agent Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9b88eeb2-e5d9-444f-82b7-04f9cc4a0ca4 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Reward shaping to mitigate reward hacking in rlhf
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eb0fd076-7e1f-4a4f-b105-2640a98e096c · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents A survey on llm-as-a-judge.The Innovation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ccad7092-8913-47ea-95d8-154beb4e1c4b · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 536ea10e-a0fe-474f-afa9-d6eb8a971c87 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Retrieval augmented language model pre-training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2eef5082-7dca-4653-9404-39ef37fe7458 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Reason- ing with language model is planning with world model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c5a900b7-0a5f-4dde-b407-57e7a7d7707e · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Sample-Efficient Multi-Round Generative Data Augmentation for Long-Tail Instance Segmentation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5dc5fd74-4f94-469e-99f6-62e8a1d63375 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72c95792-12f3-42aa-ac25-33d7645719b5 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c57c37ad-9882-4127-b0e7-4c4e70b8d627 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Let’s verify step by step
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dbdc0305-8223-4178-8ff1-cade817a5fc2 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1d16f286-4699-4d31-9965-58a52a70115c · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Agentic reinforcement learning with implicit step rewards
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dde12716-7245-4f66-8666-2fff40799569 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Symbolic and subsymbolic geoai: Geospatial knowledge graphs and spatially explicit machine learning.Trans
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 66ff8920-b637-422e-9234-891c994c2a83 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Augmented Language Models: a Survey
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 607163d8-842d-4eca-8978-8dd92539bc64 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Uncertainty quantification in llm agents: Foundations, emerging challenges, and opportunities
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dead8b95-81ca-4233-9490-4607066f35e9 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0f1eefb2-5760-4d62-8312-4efae2694fa6 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b069b3c8-cb86-418f-94b9-e3117a4981e0 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 74f3a646-af73-4093-9624-7a67d59c5b8a · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents ToolRL: Reward is All Tool Learning Needs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cad7fdda-c1b9-451c-b078-33c44222fcaf · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents A possibility for implementing curiosity and boredom in model-building neural controllers
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 30ef472a-92d5-4cc3-97a4-902a8226c098 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7a71902c-632d-4363-9a7e-e9bd2d163928 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c2ca204c-6241-40e0-afc0-1f11f80191cb · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d180fe40-4558-4732-a493-9eefd30a594c · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Monte carlo tree search: A review of recent modifications and applications.Artificial Intelligence Review, 56(3):2497–2562
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e890ac0e-3c93-4f75-a7f8-2a656467d5f8 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Solving math word problems with process- and outcome-based feedback
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 71b714c2-ce8d-4902-b461-f6355dafef08 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8f6170ba-ec94-4679-9880-94e7e6a78c23 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b59c1e80-b311-4087-a119-9dbbe090a57a · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 17c7d3ff-d50a-45d5-868b-bfe81a49f6fa · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 62165ce6-d375-4406-99a3-6bcd1478482b · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Qwen2.5 Technical Report
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fc52fa50-7a6d-4d03-9159-74fa85313a7b · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 87b15eba-f30b-4a0f-aa93-ef773b398c8e · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d1c10b34-de1c-4fc3-874a-c8745195a556 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4ac2a3c3-7399-4a83-aacd-b7eedd783c24 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation efe056c7-d2ce-416f-b3d3-3d9bc68915a3 · outbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Enhancing agentic rl with progressive reward shaping and value-based sampling policy optimization
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation be776092-685d-4b9a-8c28-90ea9b07680a · inbound
From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.