Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T00:13:30.397515Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2602.11351.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T00:13:30.397515Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4cf34756-4688-4a94-a3f3-eb8df2a296f7 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Consistently simulating human personas with multi-turn reinforcement learning.arXiv preprint arXiv:2511.00222,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 925a5ce1-d2d7-4178-bfe5-2fb732e95e7c · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d47239c-cc22-4a5b-98c7-0fd887ce84b7 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f6680bd-5aac-44f1-a237-74399e3fa718 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de66711d-3d1d-4196-a7cb-7a50aa09c23d · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Contextual Markov Decision Processes
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3cd077b-4b50-4363-b7c5-5ba09127493d · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization R., He, J., Yu, H., et al
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed6abef3-9eaa-4136-af20-fc317702da7d · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Quagmires in sft-rl post- training: When high sft scores mislead and what to use instead.arXiv preprint arXiv:2510.01624,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdafd0cb-8359-4589-80fa-a54d0b1c4e9f · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dbddb93-2803-4ca7-88e5-63fe4e100703 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd3124c4-1797-4d42-9b63-7bc0125cb811 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f75ca2ed-f3e2-4825-8c23-754567efa4a3 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21967e0c-e5f6-42a9-a2a2-d93f2d683900 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c097a627-c7e7-408b-87bf-6b323c113e63 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab94e218-f79c-409f-8062-ecaef6c54f67 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization HierTOD: A Task-Oriented Dialogue System Driven by Hierarchical Goals
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8a8e0be-0c63-4936-a4fb-ca7f1444b804 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97c2fc39-6466-4cf3-a527-32042fbe3344 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 670737dc-3c3a-42ab-a041-0bf224cc8170 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f86f967-51b6-442c-a056-e36475a4f012 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization ToolRL: Reward is All Tool Learning Needs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df72b7ae-2db8-4bda-985b-47de9f1a2166 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2465e9b8-0ca7-4897-9838-026eacf3cd3d · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbaf1417-b5da-400e-9a9f-40f7dd5c0742 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Language Model Personalization via Reward Factorization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc728a6f-6756-4464-bffe-ece9910977f7 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d746fc4d-08f7-4f89-804f-0866a12d889f · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Training proactive and per- sonalized llm agents.arXiv preprint arXiv:2511.02208,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b738638-e145-4a6d-a552-07f26960a010 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Gemini: A Family of Highly Capable Multimodal Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10028b63-61cd-4504-8a77-e75ff9eb10a6 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization En- hancing personalized multi-turn dialogue with curiosity reward.arXiv preprint arXiv:2504.03206,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9391ca7b-429b-4994-a220-a8f2c6f0f8b9 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b279e7df-c54c-4fe5-a995-6298419c2e5b · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Boad: Discovering hierarchi- cal software engineering agents via bandit optimization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43f0cd01-802e-484c-b23a-15c30fe2faac · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Qwen3 Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bd0f250-4516-44ba-9c1a-1f3650c12246 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Demystifying Long Chain-of-Thought Reasoning in LLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23557dd4-11bf-41ca-905a-c1c1bfb8f343 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Demysti- fying reinforcement learning in agentic reasoning.arXiv preprint arXiv:2510.11701,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f78ddb-4cbf-4659-9b2a-0d815ea1ca68 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10def8fe-8f92-4e75-af87-e71342188605 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Teaching language models to evolve with users: Dynamic profile modeling for personalized alignment.arXiv preprint arXiv:2505.15456,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3738c32-c27f-4543-aa82-ea31e4711187 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4cb23e-93d5-4e7a-b7a3-f78585c61635 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization E., and Zhou, W
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f994dda-d6a8-42fd-964a-a8595da18906 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Yes”, “No
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a0a8997-cc09-4a0b-a013-6621dda5a016 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization GPT-4o System Card
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11359bac-b2ad-4c72-8e1a-64e34cd6f14f · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ce43951-02aa-4c32-9da1-fd317ea5175d · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b8c592-c102-4c97-a2f1-41c7335db57c · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization OpenAI o1 System Card
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f09582dc-a9eb-4fe9-9b35-23ec975f43e3 · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Behavior injection: Preparing language models for reinforcement learning.arXiv preprint arXiv:2505.18917,
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a279625d-8e07-44d8-a331-6aec4580110f · outbound
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.