Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:47:23.367773Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2507.07441.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:47:23.367773Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d1c58432-c1bf-4e9d-9d6d-d27380de17db · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1345c0c-97ec-4bdb-ba76-79784be37139 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f951891-0a63-4b3b-98d3-0325132dc1e0 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation GPT-4 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ece5b4b-0d92-46bc-a299-ae61de169c11 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation FireAct: Toward Language Agent Fine-tuning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b399fd-71c6-495a-a48b-3b945872a3c1 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation febc73ab-b7d9-42e1-9fb6-6c1cc19d2ae6 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation ATLaS: Agent Tuning via Learning Critical Steps
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c1495a5-144f-40fd-bcbb-48e8e818efe4 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation The Llama 3 Herd of Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e922bb64-c738-4675-aed0-39776862d27a · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d479af49-068f-4c8a-b185-056e521f23de · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83ca68c7-130b-41d1-95ca-d456ecddac34 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c83b9cf-8803-4826-b029-1e45f8a4ea5d · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Internal Consistency and Self-Feedback in Large Language Models: A Survey
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcac6f08-a539-4ce4-a1a4-233b120afa61 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfeec77f-50a9-43e2-98f4-73c923cc7f89 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a1dce29-d0dc-4b0b-ad32-2a4042ff241d · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation WebGPT: Browser-assisted question-answering with human feedback
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e540441-45df-4b22-9e46-e1c6f479d971 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9d1d53f-8349-4011-adc4-2a88be6ebd52 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1f0d83a-7e41-47e9-bcfb-b47baaca3f27 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8dd96be-fad7-46d0-b042-04279887dcd1 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad2e3530-7d07-41c3-a562-a53f5142c921 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d8b0230-658b-4639-bf03-1e40356300d7 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 165adf27-7a16-4ede-b22a-c2a605c72513 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1cad929e-0d5f-4796-9417-4ee8127b40f1 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation ScienceWorld: Is your Agent Smarter than a 5th Grader?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6362ff3-168c-41a8-a62c-e50eb85720ec · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f015df02-002d-4ea6-832d-a0b9e3e87cd4 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0bc5883-362f-47e2-b5bd-e51e120a841e · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation AgentRM: Enhancing Agent Generalization with Reward Modeling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 099e01cc-e3d8-45c5-af70-45b54d5a7a0c · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a51431e-3524-489a-8c88-44714305f71a · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation MPO: Boosting LLM Agents with Meta Plan Optimization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 192b2273-cd1a-48aa-ac22-ec79c06ddf31 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb61d8e1-ed56-40a8-80e1-784696faf346 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Qwen3 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad3bb6f5-ff79-4a3f-b4b1-70497cacb9c7 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6659a5b5-b057-4067-bf44-5df7e27399e9 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfaa7b3a-f669-4c3d-badf-bb1c07864b51 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e784283-0055-4fc5-9ae7-06df5910fdd3 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bd3cfd6-4049-470e-af4d-e5c63b7094da · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4efc397-8e73-4806-ab5e-4ab405074120 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55f2cdf5-e890-4e21-bc6b-14ac32574dba · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3b23f49-3da8-4173-9fa1-fe4e5894c6aa · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bff288c3-8dd9-4cdd-a884-46bef3babba4 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation A Survey on Multi-Turn Interaction Capabilities of Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7da4d6c5-15f8-4e26-8e69-46c479e6fc78 · outbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.