Pith. sign in

Paper Citation Record · LEDGER

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization

As of 6 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2602.11351.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.11351 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:13:30.397515Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4cf34756-4688-4a94-a3f3-eb8df2a296f7 · outbound

This paper cites Consistently simulating human personas with multi-turn reinforcement learning.arXiv preprint arXiv:2511.00222,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Consistently simulating human personas with multi-turn reinforcement learning.arXiv preprint arXiv:2511.00222,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:24.996133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:24.996133Z digest=sha256:93f46d089bd5615e119d90c54c5ffb5bb2b961961d217e496f980f468712fb1e

Observation 925a5ce1-d2d7-4178-bfe5-2fb732e95e7c · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.232880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.232880Z digest=sha256:ebd0d4c361b67e175faa01f898f61274cdb5a849d387900ce277b89480f850e2

Observation 5d47239c-cc22-4a5b-98c7-0fd887ce84b7 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.386643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.386643Z digest=sha256:3c791b8ffdb958081586d2efb60d931c88957787dbdaf7f24a2ec53d8127ed36

Observation 2f6680bd-5aac-44f1-a237-74399e3fa718 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.542401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.542401Z digest=sha256:0eeb0ed7f4bc9c47e51f99f95ecf6a22112dd1f4d31ac50a9131946a2bdc0a7b

Observation de66711d-3d1d-4196-a7cb-7a50aa09c23d · outbound

This paper cites Contextual Markov Decision Processes.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Contextual Markov Decision Processes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.745024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.745024Z digest=sha256:31106d29d246fb7f36eabd473162e8a66a0fe414cf5af739167e3190f3808822

Observation d3cd077b-4b50-4363-b7c5-5ba09127493d · outbound

This paper cites R., He, J., Yu, H., et al.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization R., He, J., Yu, H., et al

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.242813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.242813Z digest=sha256:de1c3a7b1ebd86d0b4879ff08b2c21e46ef725c28c5972f810f7f3846c16c5a8

Observation ed6abef3-9eaa-4136-af20-fc317702da7d · outbound

This paper cites Quagmires in sft-rl post- training: When high sft scores mislead and what to use instead.arXiv preprint arXiv:2510.01624,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Quagmires in sft-rl post- training: When high sft scores mislead and what to use instead.arXiv preprint arXiv:2510.01624,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.329894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.329894Z digest=sha256:8f7b92c02a90f02c4cfb3e681b5c5964b6a3e359936cf65700989cebc17ccd62

Observation fdafd0cb-8359-4589-80fa-a54d0b1c4e9f · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.456107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.456107Z digest=sha256:867018b5c8934da5a0f15f27419aa881bf2a4ca5819a188c8e0d25c40b0175ef

Observation 1dbddb93-2803-4ca7-88e5-63fe4e100703 · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.575467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.575467Z digest=sha256:fb967a21b8411f4162bde8b076d700bf6d3c43bb8170a10224dba7fe5aa19eae

Observation bd3124c4-1797-4d42-9b63-7bc0125cb811 · outbound

This paper cites A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.701496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.701496Z digest=sha256:fb0d74fb707cb6a9ae77dc662a9b8add57dc25fa625e9e7d8c141742a96c6226

Observation f75ca2ed-f3e2-4825-8c23-754567efa4a3 · outbound

This paper cites A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.854490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.854490Z digest=sha256:30efd4ee6ddd66954fe38b980770057137601884b906228994f538db51851849

Observation 21967e0c-e5f6-42a9-a2a2-d93f2d683900 · outbound

This paper cites UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.015676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.015676Z digest=sha256:98a0701144124ecd1e8f9d162425f8d9769483130a60192ad4a6a74832348cf3

Observation c097a627-c7e7-408b-87bf-6b323c113e63 · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.189438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.189438Z digest=sha256:4e465d04afa64cb1fd9b2286a8710384035613620caee4bab5b7c07b4a234660

Observation ab94e218-f79c-409f-8062-ecaef6c54f67 · outbound

This paper cites HierTOD: A Task-Oriented Dialogue System Driven by Hierarchical Goals.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization HierTOD: A Task-Oriented Dialogue System Driven by Hierarchical Goals

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.347739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.347739Z digest=sha256:f7feed6c6e351a2d78a3dd0276feb7a534a0ab0a742b8547b0385cd59792c82f

Observation a8a8e0be-0c63-4936-a4fb-ca7f1444b804 · outbound

This paper cites Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.574670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.574670Z digest=sha256:16497bcad2a83a9606b66619b386963683473b8541d2adbc7acc88b1e31746f2

Observation 97c2fc39-6466-4cf3-a527-32042fbe3344 · outbound

This paper cites BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.793264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.793264Z digest=sha256:37bfad88fd69ab3b1841f521922c5169f072ceb6d0ee4e81fb3ae968a8a6425a

Observation 670737dc-3c3a-42ab-a041-0bf224cc8170 · outbound

This paper cites APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.863872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.863872Z digest=sha256:2f3b9542423e442bee72d54ae0fcdd6d08bfbc28aab0311ae32829073705046c

Observation 1f86f967-51b6-442c-a056-e36475a4f012 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization ToolRL: Reward is All Tool Learning Needs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.985409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.985409Z digest=sha256:872c37ed808e4afa44b11f04e46d88fb40659f7bcbcfa4cdfba764cee12afed6

Observation df72b7ae-2db8-4bda-985b-47de9f1a2166 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.112926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.112926Z digest=sha256:5ef5c5b23bea71e059d846f8f1abbf01a3d92323999df4a6b52cb2d0945fe27c

Observation 2465e9b8-0ca7-4897-9838-026eacf3cd3d · outbound

This paper cites Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.211988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.211988Z digest=sha256:d29a50d90e49939136cfc10309d922f06e21187bfcbd3ae808ef5b6b8ebaf750

Observation fbaf1417-b5da-400e-9a9f-40f7dd5c0742 · outbound

This paper cites Language Model Personalization via Reward Factorization.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Language Model Personalization via Reward Factorization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.352379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.352379Z digest=sha256:c48e9412b40bcf2b3d57dd103ee6fe91467fe7b973f8921bb58ee5f073ddcf62

Observation fc728a6f-6756-4464-bffe-ece9910977f7 · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.480694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.480694Z digest=sha256:81731a484db01df57ddf1f4f0fc892be3cd08ebb58c9efe2c84b2973e5e46938

Observation d746fc4d-08f7-4f89-804f-0866a12d889f · outbound

This paper cites Training proactive and per- sonalized llm agents.arXiv preprint arXiv:2511.02208,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Training proactive and per- sonalized llm agents.arXiv preprint arXiv:2511.02208,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.610154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.610154Z digest=sha256:721e67efb1e15798061379f04cfe514a7edabbe7af4e9df52ddafab21b8be4e1

Observation 0b738638-e145-4a6d-a552-07f26960a010 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.684575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.684575Z digest=sha256:2559fb80688380d03527b7dca7e76d35adddb20d3dbf754ec4273b25c34e3834

Observation 10028b63-61cd-4504-8a77-e75ff9eb10a6 · outbound

This paper cites En- hancing personalized multi-turn dialogue with curiosity reward.arXiv preprint arXiv:2504.03206,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization En- hancing personalized multi-turn dialogue with curiosity reward.arXiv preprint arXiv:2504.03206,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.984404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.984404Z digest=sha256:d4c37645909545a547c54a34a6bf09cefc15f2733380c27d498db214fad6b7a0

Observation 9391ca7b-429b-4994-a220-a8f2c6f0f8b9 · outbound

This paper cites OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.080919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.080919Z digest=sha256:139c924a7e7f4222ce925b659561ab542cc066d60aa9793c753249501db41c25

Observation b279e7df-c54c-4fe5-a995-6298419c2e5b · outbound

This paper cites Boad: Discovering hierarchi- cal software engineering agents via bandit optimization.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Boad: Discovering hierarchi- cal software engineering agents via bandit optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.386885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.386885Z digest=sha256:441d26d977103eca5a03b0f2423eea40ccc6d563f88130c43513700d1f1bfbdb

Observation 43f0cd01-802e-484c-b23a-15c30fe2faac · outbound

This paper cites Qwen3 Technical Report.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Qwen3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.423336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.423336Z digest=sha256:4900304562d8609558571a6d119e47e59504cf18587362882549c03fcb789691

Observation 3bd0f250-4516-44ba-9c1a-1f3650c12246 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.528540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.528540Z digest=sha256:676058784ed0af0033de81bfc5c56646068f9d570cdd7f5ddbbc2cc7e49f1227

Observation 23557dd4-11bf-41ca-905a-c1c1bfb8f343 · outbound

This paper cites Demysti- fying reinforcement learning in agentic reasoning.arXiv preprint arXiv:2510.11701,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Demysti- fying reinforcement learning in agentic reasoning.arXiv preprint arXiv:2510.11701,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.699949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.699949Z digest=sha256:252d4c629e66b12aa66a79e720c5bb70a9edec7774eaedba2dd2170422a473dd

Observation 91f78ddb-4cbf-4659-9b2a-0d815ea1ca68 · outbound

This paper cites Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.838143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.838143Z digest=sha256:9bc5cd89d454debdd25548582426dd62b7392bb6209649ef65a085411c36c47d

Observation 10def8fe-8f92-4e75-af87-e71342188605 · outbound

This paper cites Teaching language models to evolve with users: Dynamic profile modeling for personalized alignment.arXiv preprint arXiv:2505.15456,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Teaching language models to evolve with users: Dynamic profile modeling for personalized alignment.arXiv preprint arXiv:2505.15456,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:30.021924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:30.021924Z digest=sha256:c5b570a2f7b2bfad753eeb4a6a167260a501373a7e917dd1c5a6e9dcdc9a583d

Observation e3738c32-c27f-4543-aa82-ea31e4711187 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:30.206953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:30.206953Z digest=sha256:a0f95a1de78ec1d43f461795cfadb21bb69ed0791537d20ecde2b521479fa43a

Observation 9a4cb23e-93d5-4e7a-b7a3-f78585c61635 · outbound

This paper cites E., and Zhou, W.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization E., and Zhou, W

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:30.321333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:30.321333Z digest=sha256:89d96044a3ed8ec3259c32ad54d2681e55069e297af6ba12fc1176e26c665ae8

Observation 0f994dda-d6a8-42fd-964a-a8595da18906 · outbound

This paper cites Yes”, “No.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Yes”, “No

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:30.397515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:30.397515Z digest=sha256:894f23c6e75144b074187fee8228e4d6cc8e4f1d40b1f4bf9183645991b3a642

Observation 5a0a8997-cc09-4a0b-a013-6621dda5a016 · outbound

This paper cites GPT-4o System Card.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization GPT-4o System Card

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.873768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.873768Z digest=sha256:4259c846058d903d59aade345a34aa3cf88e17da19a8e93a9e2c3dfe7c3b0364

Observation 11359bac-b2ad-4c72-8e1a-64e34cd6f14f · outbound

This paper cites ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:27.661295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:27.661295Z digest=sha256:e1c3f9fd8bebc49b429382342aca0c1cc2b9ce4e29d1173d7c50bcb11715494f

Observation 8ce43951-02aa-4c32-9da1-fd317ea5175d · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:28.882418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:28.882418Z digest=sha256:aa69b7bcce4b2d96ac5563097b4f5437684d0478434d1eca7bce7a90624b87ef

Observation d3b8c592-c102-4c97-a2f1-41c7335db57c · outbound

This paper cites OpenAI o1 System Card.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization OpenAI o1 System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:26.072593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:26.072593Z digest=sha256:73cbc2e7da806994dc1cb8de630031c0e0701b521ce5dee3822f49aa5468d70b

Observation f09582dc-a9eb-4fe9-9b35-23ec975f43e3 · outbound

This paper cites Behavior injection: Preparing language models for reinforcement learning.arXiv preprint arXiv:2505.18917,.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization Behavior injection: Preparing language models for reinforcement learning.arXiv preprint arXiv:2505.18917,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:25.140784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:25.140784Z digest=sha256:767eb4dd4174407ad21cd9ed2c6e7011f312ae578c8c58ee82220c377ddaf65d

Observation a279625d-8e07-44d8-a331-6aec4580110f · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T00:13:29.208730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:13:29.208730Z digest=sha256:6c375a53a09749bffa1e4b41f23bafd2941e8c4087d8ab7de6489486678e0287

Pith citing papers

No inbound Pith citation observations are available.