Pith. sign in

Paper Citation Record · LEDGER

SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2503.15478.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.15478 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:56.255061Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:34:57.712764Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 14a9e272-3755-4f82-8ebb-4c5d0b603de9 · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:56.255061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:56.255061Z digest=sha256:26641ccd3e7c1fceef5accd04c9e877ba0d1fc14e251b4d9eaf77eeba31f16db

Observation b169f929-fb11-4c83-8be7-a284473d8032 · inbound

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay cites this paper.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.796768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.796768Z digest=sha256:7aca5f9dd3dd1e7d060c3fe1df6c393c71e13ba0a011d1e24db81670bdc3a4e2

Observation 73ce5631-804b-42e7-a0dc-7ac9d5f4e945 · inbound

Get Experience from Practice: LLM Agents with Record & Replay cites this paper.

Get Experience from Practice: LLM Agents with Record & Replay SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:31.351435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:31.351435Z digest=sha256:046a87f92837cfcf4ded3e6a5893a7a527b97ba7ffec85f2bff7db98b7e5bfb5

Observation b6dca655-ce83-4790-b963-216ffd7ff111 · inbound

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution cites this paper.

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:59.729856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:59.729856Z digest=sha256:9b8cf9dfc5c6f06cbf7ac0460db9ea9e74c5fae39543ca645875316f218afad6

Observation 1454204d-6f50-4207-9c68-4aac41fc2c92 · inbound

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation cites this paper.

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:35.042447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:35.042447Z digest=sha256:52a45acd584594d08ad41a68d0af7314605f91c0e2ba1f89844a91c20ff552e7

Observation 4f807b02-017d-4bed-a1be-c32a4eab9262 · inbound

TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence cites this paper.

TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:52.419144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:52.419144Z digest=sha256:e9d3ae7eb145908694a8e52332e96571e21ffa6a19d33dffd591073ca4384d37

Observation 125d7087-241d-4a98-9ab3-b3ac2c9313c8 · inbound

ARIA: Training Language Agents with Intention-Driven Reward Aggregation cites this paper.

ARIA: Training Language Agents with Intention-Driven Reward Aggregation SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:01.135410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:01.135410Z digest=sha256:baca21a23b80db70863397a5c62ae9b7c05c77b722ad67a866e125c7830895b0

Observation 17e33e8c-9b01-4cc9-a34c-5da8a28ac835 · inbound

Self-Challenging Language Model Agents cites this paper.

Self-Challenging Language Model Agents SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:58.657810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:58.657810Z digest=sha256:4be90fb8b2150c02f7882792150d8092eb2687a754f115051d0503ec2c7e5b13

Observation 3d37a5d8-7487-4031-94a2-c63c0b08b245 · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.314881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.314881Z digest=sha256:a66bb202385c383626980e6e56a9ae9d47298c5f5bf3dd9b6309bba34204cc66

Observation 7625a69a-a765-439d-8c35-df5ee4ddd1c3 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:39.160869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:39.160869Z digest=sha256:110ccdc81ef7a6b38d6926aca4878f98c745d9db215de98790d43edc91e21c02

Observation 1f0dc4f4-f5e2-4c7e-91bb-808ff5e823e3 · inbound

Conversation Forests: The Key to Fine Tuning Large Language Models for Multi-Turn Medical Conversations is Branching cites this paper.

Conversation Forests: The Key to Fine Tuning Large Language Models for Multi-Turn Medical Conversations is Branching SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:45.270627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:45.270627Z digest=sha256:693186d43e74d6334d5ea07a8933fe877b182374077afbc5cc77358c9cb48deb

Observation 633bb555-fc2c-4003-9188-8f313ed67ce6 · inbound

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs cites this paper.

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T18:13:00.021058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:13:00.021058Z digest=sha256:a7665ddbcf1ff63d995cbe042e02774aa66b0aa5e6b88a8c77fda0ce0f11f5bb

Observation dd58bca8-26d4-4d0c-b6ab-e0f0866f15b4 · inbound

General Modular Harness for LLM Agents in Multi-Turn Gaming Environments cites this paper.

General Modular Harness for LLM Agents in Multi-Turn Gaming Environments SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:11:22.785205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:11:22.785205Z digest=sha256:475dd1de3be2c9b736a6c75509dba694605fba1f85927a9f2e796516fed56578

Observation aa65f488-02e3-49d7-8feb-539207c4b1eb · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.728989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:682644850e8ff73f028916dc68962ab89f06aae11ab457a7762e66e782292ce1

Observation df998f68-4a2b-44f7-8386-e11527426948 · inbound

RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs cites this paper.

RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.226541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:29:46.136115Z digest=sha256:5ae821fe500345829b3b9e664be38cacd672e3b07e5bcbdb6f4d2a0f4057ee40

Observation 43c7aad6-49c0-49e6-906c-976e2360c91e · inbound

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse cites this paper.

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:05:39.034862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T02:04:46.559881Z digest=sha256:f90f20f3899cdbbe9a2bb84cd9dc9bcaac4ab4be61b47d3df30fe227758da95c

Observation 1a6179a0-be0e-4fec-a489-4edb3cfde94d · inbound

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction cites this paper.

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T14:34:51.341451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:34:51.341451Z digest=sha256:24200a7c3f729971f73b0793a0d5f176e36a0f28503bcbd1eea9503fc0bfdb09

Observation d137f4c8-ea46-414d-83e9-9e000525dd33 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 235

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.300750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:685b0f1eeb894c19b2db621c9ca3e43d2a5108c691797c004dc7aa1e8f3f2e9f

Observation 1a126558-2aa2-41fa-822f-bb3b714ebddb · inbound

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning cites this paper.

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:22:31.136212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:21:01.335414Z digest=sha256:d5d30707eb48896241cc5c6243a3f09e346481d41567da39d5801b977019a872

Observation 04ab4af9-3fd6-40cc-b505-75a93ffb9a05 · inbound

Multi-Turn Reinforcement Learning for Tool-Calling Agents with Iterative Reward Calibration cites this paper.

Multi-Turn Reinforcement Learning for Tool-Calling Agents with Iterative Reward Calibration SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:53:11.529146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T19:51:50.503945Z digest=sha256:7de7782df87d8c29057f1429debe03fa683a3d78edcc769201af90d51087890b

Observation f3d42dab-7dce-4859-ac9e-c1ed570faaeb · inbound

ActivityEditor: Learning to Synthesize Physically Valid Human Mobility cites this paper.

ActivityEditor: Learning to Synthesize Physically Valid Human Mobility SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:05:50.000162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:42:36.437661Z digest=sha256:6111eebbda95f4811fc309b21922a97aa5b80db7d9c58b26063a3202e4ba69c3

Observation 26c3d4ed-4455-4ae1-9410-e2bf0424da5c · inbound

PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent cites this paper.

PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:06:11.230632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:17:59.156121Z digest=sha256:9488cab41830859ec9ac88e8920d4b8ef149c8dcb041dd24793609517e723632

Observation a0fdffef-ee8e-439a-865e-3d2bcf9dfcf6 · inbound

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants cites this paper.

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:29.471183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T05:48:00.486572Z digest=sha256:10c695098f111f4c8684b1f5a1f0a596d4e33dacaec995f1388f43b01e05489e

Observation 11058383-d232-4920-a4e6-e116d4043f25 · inbound

CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators cites this paper.

CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:24.049413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T00:51:57.796883Z digest=sha256:ab93013136a5a43407d7f18d02ef1e2ead5021deef9f04e4493578821eb82c88

Observation 3390848c-dd81-46ca-98fc-9a18b501fb71 · inbound

CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators cites this paper.

CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T14:37:24.690196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:37:24.690196Z digest=sha256:bc80053cff04d4e2183cd04199b8885966990b25327d93e3d4f5e66f86b29c63

Observation b9ad27e3-ca85-4d65-96da-dcefaabc01cb · inbound

Step Rejection Fine-Tuning: A Practical Distillation Recipe cites this paper.

Step Rejection Fine-Tuning: A Practical Distillation Recipe SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:06:35.455960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T03:44:33.525966Z digest=sha256:5351a23b770ea40c2bf87c1517a4215fe91b2d0586be908a97c2966c9aacef70

Observation 754fb7ba-0268-41af-9679-6dd5bcb6f288 · inbound

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization cites this paper.

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:43:12.469797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:41:25.205368Z digest=sha256:caa3b809bdc2ec5eafa522a4977a4525d2c6d57e5e515b7acc20746a0a79e947

Observation 7e2d00a5-471e-438c-94c9-204a55df7b01 · inbound

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization cites this paper.

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T02:23:32.483381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:23:32.483381Z digest=sha256:7be6e3419be36b908fff50cf43605ad0f108f46bd181b0de71b05f84f10d9cf9

Observation f5193e78-721e-4d63-849f-01c0f30ae694 · inbound

Unlocking Proactivity in Task-Oriented Dialogue cites this paper.

Unlocking Proactivity in Task-Oriented Dialogue SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:40.521037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T05:31:37.171422Z digest=sha256:54c77105596b9050eeeb2dd39916abd6d5416aa3cd05154adc5139bbae1928f8

Observation 8d6e6666-2b71-4cb6-b8a8-6c70dc737ad5 · inbound

Unlocking Proactivity in Task-Oriented Dialogue cites this paper.

Unlocking Proactivity in Task-Oriented Dialogue SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.714454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:31:11.074231Z digest=sha256:89dba830c54e3282719207d3bbe86c5de203daf539bbe3558123b3b279e6d842

Observation 2e6c0734-d37c-44a7-ab34-e99e658384e5 · inbound

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use cites this paper.

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T07:36:55.011277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:36:55.011277Z digest=sha256:7df4d31cb8d6d7b1b034826e0975928f12f648f50256f6be0220d8e0cf346a12

Observation fb439d78-5096-4625-ab82-618defb3f50a · inbound

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning cites this paper.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:30.431790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:30.431790Z digest=sha256:8db98208038dfdf6897f49877150be3ce835ca184ea408cc51c187741a819f52

Observation 0188d5a1-8c76-42b0-9bc7-0a7d4b590917 · inbound

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning cites this paper.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.291734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.291734Z digest=sha256:1c6a512165eb20e5c4635703d2096d86446f8cc50ca5a85de3fff682649cf1e4

Observation 29aec9b2-098f-4024-ada1-5ec02895ecfb · inbound

TCPO: Turn-Level Credit Policy Optimization cites this paper.

TCPO: Turn-Level Credit Policy Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.690152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.690152Z digest=sha256:a9167d8cbe1866b65509c9728b4d103c02f4acbdc31a108b4cd5d0527b609a5a