Pith. sign in

Paper Citation Record · LEDGER

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning

As of 4 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2607.25369.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.25369 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:44:33.432754Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 58f3f583-70bc-4a8e-9769-6e73fadbce18 · outbound

This paper cites GPT-4 Technical Report.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.030089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.030089Z digest=sha256:3a857cc797e33383f512e293051324b2882c0ab95fcb63c4a12fb858d60fdfc9

Observation 15a77535-4fa7-4cce-8d30-7cc0e1a9d627 · outbound

This paper cites Qwen3-VL Technical Report.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.142384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.142384Z digest=sha256:f1782bb3fc37ad0c8b23a04c23e1fa123fae12d306b20112839e4dfa111f3255

Observation e6014760-aaad-43d8-afac-a648af427baa · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.254075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.254075Z digest=sha256:2a6a9da6e3054205ae52d445b85238199efcba58ce16ad4a488d36746aef0ffe

Observation 212faca4-28ac-4214-93c4-009490a9726b · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning FireAct: Toward Language Agent Fine-tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.355941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.355941Z digest=sha256:d4568cd802a016e9eaab60ddf0aa0bd4e4c575fe88c642cfc44cca08949ffed7

Observation 66d0f88d-5d33-4e42-bb02-4fd408dc819e · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.468395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.468395Z digest=sha256:d5923274c3941efb6942358aa67bbaac366eb092ef380cdd582e633b66e33fb1

Observation 2c8899e8-849c-475b-95ac-088f14c3c23b · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.578227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.578227Z digest=sha256:9b04813357fbbc1caee0cc914381aa1a4fad565997e9d8de284a9ebb659d42ae

Observation 0fe7d395-8db0-4109-96b0-e308c5d50ef1 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.653544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.653544Z digest=sha256:7791586e305cdac8d1776a639da517ef42100ab0d0be8454bee4b0c8d4cbe9cc

Observation c2808dcb-3851-44b4-bd24-dd3d3550767a · outbound

This paper cites From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.873290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.873290Z digest=sha256:86a214c30ee0f67620259716399933fa8d728fc8394f7cc4e1a71c144c0753ad

Observation ee2297b6-fc2e-4a24-a527-71498abb493e · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.974453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.974453Z digest=sha256:20a863b02b78201ee2e2ba22ada3842fb056e4da7c6210c58feefe55d636ec8b

Observation 40ce625e-4d6b-4884-b8bc-3a84d544f609 · outbound

This paper cites Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.092180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.092180Z digest=sha256:c349ba54fa6de949ba4e6ac9fca137f5635f2515a0fea1836e265ca11840ff1b

Observation 71d403d3-eda2-4fc8-a613-73c58194ca33 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.232639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.232639Z digest=sha256:34b45eea821de3d895670cb2ceaf14c38010732b43f0a9baaa90c6228962072a

Observation 880bd765-69b8-4ad1-b4f0-c7c87fd570b7 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.342917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.342917Z digest=sha256:31315b7ddcc5d8ddea7ffcd768353e428fb35ef8a8d3cddf4b45aaa4091be299

Observation 388dcb4c-d70e-42b3-9bd3-dececbbff4de · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.380852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.380852Z digest=sha256:2554efc7c9ca9e32ba7a0dddb170b50dfdd16adddee7297ba3641e24db9fca47

Observation ee12d701-c700-437b-83dd-c7063ba2cbfd · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.449127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.449127Z digest=sha256:3d509819e9f4ad27ea85295a3513eadc8a3ff7af5555e81bbc9217a31009f2ad

Observation 2c282ac1-3a35-4f29-a61f-aafcc247f28b · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.487990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.487990Z digest=sha256:a4c191fe0860df9f6ba478aef4cbff805127ace1a1661c6886a004b2bbceb521

Observation ec876918-15e4-4592-8e12-99592aa3cfcd · outbound

This paper cites From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.559900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.559900Z digest=sha256:e0bbe4d5b5eeb1e683888fca747ad987a32c65fe0a3d660b108b244ed7cbf565

Observation 6fa4f9bb-57a3-4eae-8f77-8758f82df654 · outbound

This paper cites Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.630565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.630565Z digest=sha256:f74722e2c5dba4e81e71c54465ee33781a211862a2466ab63d2671b1cb18c4d7

Observation eece9899-17b5-4219-ac0b-587ba4dfd51b · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.704787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.704787Z digest=sha256:a8952d61ffe83b1680690179a56f737b6fba4238c3de9d359a3ecfd54f862f18

Observation 9960dda1-b0ef-4a73-ad1e-49e17baeccdf · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.747447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.747447Z digest=sha256:61415259ccaa0d75492c8ddcab848d84ad8c4aada49e0e4b73e866cbe587c820

Observation 157abbdd-6fb9-4ddb-b113-ac938e7bf219 · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.818457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.818457Z digest=sha256:1f3c1ff135a8dc627af3a36f2283af9aec73f507fb97c6cb1507418a260adfbb

Observation 73398d19-4860-4aa9-84e5-9c476ca629a8 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.917786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.917786Z digest=sha256:d7b2cbc49276ba6ec2449325c1198c074ee4e70ee93c619bdcca2386d022a027

Observation f7d06e82-c850-44e9-81cf-6b54fed0c298 · outbound

This paper cites X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.959960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.959960Z digest=sha256:18eecfc9458cb367ccb4e67807179a3cce3b7308d9cee8e175bf651cd351db68

Observation 91497f2d-7e15-4fbb-b50f-cee25b30a192 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.029300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.029300Z digest=sha256:9a506589a0fa216ad5346d4ad1479e1430157978e5cef831a9a67a22cd65388c

Observation 4afd7b9c-59bf-44ab-a362-c5a946216c51 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.102170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.102170Z digest=sha256:dd58d5857b4d7cfb1e7fdf5991c27131f62e3b6c584023f91ea1496355bbb782

Observation 7b63ecdf-9291-4b39-92cf-eb54e93d250e · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.319484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.319484Z digest=sha256:51c66e5aa967002a49ee1e90510bb45977975a5403a5ff2dc85bb2c579769fcb

Observation 4b2c6bd0-cc90-4cb9-a844-6e61c5acfdb7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.427858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.427858Z digest=sha256:0646e698f814b809ce38881434a42299e8294b1204652533f1986b370db7d06e

Observation e864c572-b28a-4bf4-bf8f-900a67cd20d8 · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.519890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.519890Z digest=sha256:ec103db203bf2046b4e41b45209fd824a53be3018986d472531b28900d725d01

Observation a2a54a8e-37d7-4566-a743-38c3d5bdfa09 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.607597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.607597Z digest=sha256:7b012873dd0eb27931682a989fce5550f10e2772448bc968b40c52d07b710b2d

Observation 773a5ffb-195f-4335-af1b-349e45e0d3b1 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.741806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.741806Z digest=sha256:78ef071ebcc5038e42cb811cc70d4c9a11b5c67ec7137b33e8c745566dbc683e

Observation 064b171f-b589-4cda-9428-1e1d2bb06a69 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.932750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.932750Z digest=sha256:e2e707a9be10744bbffd5dbbb43a6cd21b1b8ebb8bd0fd3e8c605711bcc1e129

Observation cb2d267d-00e2-47f7-9aeb-5b1c3b4e103f · outbound

This paper cites OpenClaw-RL: Train Any Agent Simply by Talking.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning OpenClaw-RL: Train Any Agent Simply by Talking

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.111226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.111226Z digest=sha256:2de7e45298a1eeee21b4ab221ab927f1fc559bf14ae0ab2d384a22a4d304fff5

Observation cb3e8800-5cea-494b-ae26-5caf540394cf · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.308466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.308466Z digest=sha256:29e7e07dd69a1349a99fb1004da2f8413596fbcd4f4906ccd8d8c6fe36325327

Observation 945e01cb-44cc-438e-b99b-edd7d0b760b5 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.508057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.508057Z digest=sha256:f49a7d7b7efd5ee2b8ff41176f0950023d5f1bc6d5111bf464c12082055d4444

Observation 5ba483d8-be47-44fc-a746-f91a8801b4f8 · outbound

This paper cites Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.617211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.617211Z digest=sha256:40ee6fec1a31ac6360a274cc88f14e831b1f9c39ec2392bf599cc0cbc8e964e7

Observation 9e394e04-c7f0-45bc-8e30-e4f8de9faa4f · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.737062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.737062Z digest=sha256:134f05f4a72f2443fb642edf3b7dec704796f6db27123afb3d0483bc9f1bc6cf

Observation c871051c-eac8-464e-be0f-4efce5de4c16 · outbound

This paper cites Agentic Reasoning for Large Language Models.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Agentic Reasoning for Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.776319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.776319Z digest=sha256:a3a23c80805fac28912230ec3cdf76ba5b2771001e0a8b2239eb4f744cd5c95d

Observation bcf65ac5-2bd4-467b-b53c-074acb317f50 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.815219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.815219Z digest=sha256:51803bf0702f3f1c11ee005c6db4dc05609bc29d42e09a8e1525f505ebebc724

Observation 4c98dd14-f00e-461f-bebb-80a32febe6e1 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.853689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.853689Z digest=sha256:1bf54fb444d12a1d78622c776259a4a9ba9d79ca8db4f398204d6dce8153b52e

Observation c09ca760-3772-45a1-80b9-78f678767612 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning ReAct: Synergizing Reasoning and Acting in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.890559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.890559Z digest=sha256:7e5811627c443d3a2b8dc197b2557aec75b1122280f4613845f05d2a962a4da1

Observation 3297b684-4248-4c10-8cdf-c7210307d948 · outbound

This paper cites HomeRobot: Open-Vocabulary Mobile Manipulation.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning HomeRobot: Open-Vocabulary Mobile Manipulation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.954619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.954619Z digest=sha256:6f52cd5d168a8b59504a02be740d9038083850697b8c988b092473ad14af21c7

Observation b418a0bd-df24-4634-9464-13a4da7c0c8f · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.991522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.991522Z digest=sha256:f475975c984b609261ff5fa53484fd3939e7dceae9b522b8dbe05a641f061720

Observation 106d3082-ea8b-412e-b1ae-1ac231548a4e · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.028513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.028513Z digest=sha256:a8103e95f690900d4ea80bf79be2a9daa076ad20d5a2733f9ea802c842d12394

Observation be864231-b730-46a7-acdf-c92c11cabaae · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.060330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.060330Z digest=sha256:20df6f963bee119ff19e851514b95e6e5ad50125c6081b0e9ec93b0924871b27

Observation 284279d5-12bf-425d-b36c-799523dce73c · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.092825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.092825Z digest=sha256:48c46b134c42294856aecc717b8aad3210359652984ed737548f73e960a2c7d4

Observation 20078f02-6ba7-405c-9c0c-58e27ae032b4 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.162251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.162251Z digest=sha256:aee98679ba0986e81a482720287c4d6743c1b6ca112b72f0660547f935808cf7

Observation 479c00b9-a6ab-4484-b106-ecd04b4ee11d · outbound

This paper cites The doubly librating Plutinos.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning The doubly librating Plutinos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.192307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.192307Z digest=sha256:66450d3e8c3be7893ac3a3df205813fffb3115c1a7af7aff70d70726a77396f3

Observation cbdb7d06-b2f5-4131-94e2-7b3c64c7fe40 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.224055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.224055Z digest=sha256:4096aa237efca571ff60b5482c7d221ef4a28404b76c4b23cb816d2525b2e8b7

Observation 19789d3d-f449-4af7-a8a8-fa564f4626c9 · outbound

This paper cites A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.128004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.128004Z digest=sha256:7b42d84bca85d3280cb4fa4abba9d2d174152d323f74d6b5dc29aaf3a78c0eae

Observation 19ff2c2a-57ae-48cf-a597-819c5ef15596 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.307285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.307285Z digest=sha256:ed373ac9671a36c0ac1bc55baa946c02ad25de32c26da272c3b409dc2f20f0e1

Observation 422feada-5b73-4fdc-8002-63a85f8d2062 · outbound

This paper cites SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.316386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.316386Z digest=sha256:9c0fe90cc4e512370445177a480c509a0263f430bf19d6087035c6cbc206bbbc

Observation 0188d5a1-8c76-42b0-9bc7-0a7d4b590917 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.291734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.291734Z digest=sha256:5fd2e5702d1ddc56d3e1e71d1cc3e995752846fe6a3f2ec013213d69cacc60dc

Observation 8006beff-8efc-4d38-94de-7308fdd69321 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.338196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.338196Z digest=sha256:7bc8720ad29f17f236d1c476c45f7561bf43f74dbd6fee9708a13dd735234423

Observation ebeefab7-4b2c-4ed8-9225-d7d84263b6b1 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.362229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.362229Z digest=sha256:7d0ec92d76f743f15794ab298220c3426e21c9f3cae0cc3a046ff30fd141cc5e

Observation 15578491-cf05-434c-8e67-15d42301103b · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.384019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.384019Z digest=sha256:77f50416d113ba532441cc530995f0e426a9ba99b0b00989b8801eee3cd39bf4

Observation 29cdaf4a-0c77-48df-b4d1-b46a3e74bfe3 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.407391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.407391Z digest=sha256:096ed1dd27797cc9a5467e0fa459e8273e29f8f1794bac9cf056ec06d75e7771

Observation 00cdaf78-3c0b-40e8-88d5-b60f44f28883 · outbound

This paper cites •Do not output multiple<think>/<answer>blocks.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning •Do not output multiple<think>/<answer>blocks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.432754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.432754Z digest=sha256:9f156d47a02cebaf3009e36a496d9ba400094c74e87c03f4159d188b9d80ad04

Observation f80b8360-9a88-48f3-9601-f6a549bd68e6 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.212369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.212369Z digest=sha256:2129711b9683ce64a8a554d39f1afd45ab447569bc75b18ec3444dabd3e07448

Observation 003a6523-95fe-4860-8329-fc4f1bbafd50 · outbound

This paper cites Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.759049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.759049Z digest=sha256:4b64e46e6f5811a14f4104202a8e17e0753d6496d3c08c5ed5ca5f9823996d14

Observation 66ea9cd1-ea2b-4dbd-8c5e-76a897a75c15 · outbound

This paper cites InThe Fourteenth International Conference on Learning Representations.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning InThe Fourteenth International Conference on Learning Representations

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.665859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.665859Z digest=sha256:13d794d7e1c50bf97dee636396f1bab6c0571de89965d8839efadecc253643ae

Pith citing papers

No inbound Pith citation observations are available.