Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:08:13.270416Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2411.19547.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:08:13.270416Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b0164537-341b-4af9-9018-a9ad87aa6898 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Dota 2 with Large Scale Deep Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3911caca-315b-48dc-ab0f-151a45476f58 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models FireAct: Toward Language Agent Fine-tuning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef04eb93-7dd7-4bca-bd77-6e1efd7c23fd · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f793bbd-7632-4028-b1b5-75ba93f4de6d · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28587567-42b0-4924-987c-a60391b68760 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 718e2d56-b39d-4f36-be58-647422554ef5 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db37de30-47ce-433c-a044-113fc4eb932f · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Large Language Models Can Self-Improve
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d8aa783-376a-4892-b90d-fe4e86b880c8 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models SelfEvolve: A Code Evolution Framework via Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fd38620-f1e9-4e65-b8e7-b22ecbfcf6a4 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Self-training language models in arithmetic reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8b844dc2-c8b8-4dc0-9edb-81bc36c5c626 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 28d9802b-9719-4342-a050-5e0f7222eac1 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d6a8eb0-7b47-468f-a64e-8bce11b4c6ed · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8711026-5546-4613-a603-0378ef86e67a · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models WebGPT: Browser-assisted question-answering with human feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22770d27-8490-46bf-b7e4-ddb966e6a351 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc839b6f-707d-4c4d-b329-ff278a5b0229 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Gorilla: Large Language Model Connected with Massive APIs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8342b52c-e514-4a8a-9361-2ac05e32dcba · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 464f3854-a11b-4c8a-9b4f-92286fd17e73 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e5d3a35-5ba9-467f-afd8-0d04ad5bd38a · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93d9b338-0bdf-4cd7-ae47-22bb95f431e7 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca4bf90c-b752-47e8-ab71-18f14afc5c18 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30f9efe6-4235-41ab-80fe-fae95410b525 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 961843a0-f1d2-4a48-a538-4833fad57d09 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10a9cc14-2510-4922-9bd9-38ad76b28956 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4bca0c1-da40-4e88-a251-c4d224b5cad8 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9ee9c83f-1709-4c35-82f4-122a2c3981aa · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66349724-564e-4ced-804f-1e789f21684b · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models The Rise and Potential of Large Language Model Based Agents: A Survey
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c40f4bc0-521e-4a6c-aaa8-3033dc969e04 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Lemur: Harmonizing Natural Language and Code for Language Agents
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 56c74e4a-5805-45d7-af8c-af24eb062a9e · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models ReAct: Synergizing Reasoning and Acting in Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2894738e-c67c-4bff-8d49-7d942bb8df60 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Yi: Open Foundation Models by 01.AI
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8030f6ea-7901-4745-9615-cce2475dd0d9 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models AgentTuning: Enabling Generalized Agent Abilities for LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af16f172-6a01-4b14-89ea-2613ee381586 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2395cfca-8a4e-477b-8f87-0948d7c215ed · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models online" 'onlinestring :=
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 159d06ae-b93f-4975-ad0e-a008cba38128 · outbound
Training Agents with Weakly Supervised Feedback from Large Language Models write newline
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.