Pith. sign in

Paper Citation Record · LEDGER

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning

As of 19 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2607.25369.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.25369 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:44:33.432754Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 58f3f583-70bc-4a8e-9769-6e73fadbce18 · outbound

This paper cites GPT-4 Technical Report.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.030089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.030089Z digest=sha256:d94783e16789d957ded84d2fd8d7460411357eb3d6c14e15a92407c792cb825d

Observation 15a77535-4fa7-4cce-8d30-7cc0e1a9d627 · outbound

This paper cites Qwen3-VL Technical Report.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.142384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.142384Z digest=sha256:ece9f0c151d5b222f7e532d85cba2a174219f3a30b3c3587acdf79c58676b8f7

Observation e6014760-aaad-43d8-afac-a648af427baa · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.254075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.254075Z digest=sha256:a512530a67bd1d87b2d69c2941e03bc6617651bdeb6855eabd58db71bc54107d

Observation 212faca4-28ac-4214-93c4-009490a9726b · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning FireAct: Toward Language Agent Fine-tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.355941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.355941Z digest=sha256:7d524cca528728bc33a10360d70340338b1ba0e327a9e6dcab8a1d04a908e9a6

Observation 66d0f88d-5d33-4e42-bb02-4fd408dc819e · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.468395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.468395Z digest=sha256:48c3ebdae917369fc665a1e230db924fc74eb0892fea1a88fa411ba95c32e0ec

Observation 2c8899e8-849c-475b-95ac-088f14c3c23b · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.578227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.578227Z digest=sha256:0ab382c9cce0d8fa5b287f922770a2e439615101621830088c0e160657fa706a

Observation 0fe7d395-8db0-4109-96b0-e308c5d50ef1 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.653544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.653544Z digest=sha256:9db85dab61d8a0f9a4dcfed17c4818c53dbdb288244a200ec6b7de042b6da116

Observation c2808dcb-3851-44b4-bd24-dd3d3550767a · outbound

This paper cites From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.873290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.873290Z digest=sha256:d4307bab208207d5bdd98cbd4dd94fb5b0d419c7fc09d5ae600e4cca93689ce9

Observation ee2297b6-fc2e-4a24-a527-71498abb493e · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.974453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.974453Z digest=sha256:79d85739ba08cf3c0f7648154084a013afbb197cc7f81ee6166410d70ebefbf7

Observation 40ce625e-4d6b-4884-b8bc-3a84d544f609 · outbound

This paper cites Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.092180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.092180Z digest=sha256:94681b5b65380da1859f27b96746b5e0f87ca586911af610a661d7907a704a8f

Observation 71d403d3-eda2-4fc8-a613-73c58194ca33 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.232639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.232639Z digest=sha256:3302895eec255c0585bd8356b515fd02af8d4a11ab58a8bb0de61a4bb524f3c4

Observation 880bd765-69b8-4ad1-b4f0-c7c87fd570b7 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.342917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.342917Z digest=sha256:922cc774c7d1702524578acb09d31a35a134f94d6eec6b68b8665bab2db034d6

Observation 388dcb4c-d70e-42b3-9bd3-dececbbff4de · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.380852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.380852Z digest=sha256:6a7b18fb240cca6747bc470de91c86d3ee3cee9a00a0cca4d3b3be8c74d25be1

Observation ee12d701-c700-437b-83dd-c7063ba2cbfd · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.449127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.449127Z digest=sha256:679f6c4893e3a8f3e77607ab9263fddae194a9ec8e92a6d0cf523a60ae3d2118

Observation 2c282ac1-3a35-4f29-a61f-aafcc247f28b · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.487990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.487990Z digest=sha256:b11e66a81b3f17a9b20e5771113a1e9b6fe0648276eb4755b9f5cf76a79dd9d7

Observation ec876918-15e4-4592-8e12-99592aa3cfcd · outbound

This paper cites From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.559900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.559900Z digest=sha256:9ec088f19384bd71c03d3c5a86d75ff6105badf80e0b8c0040637a359868b687

Observation 6fa4f9bb-57a3-4eae-8f77-8758f82df654 · outbound

This paper cites Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.630565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.630565Z digest=sha256:7f7854e4122301d2e1e022292dfe2aa78e6ad53f8995940a98926cfc086b47af

Observation eece9899-17b5-4219-ac0b-587ba4dfd51b · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.704787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.704787Z digest=sha256:1eb70fd9867423ad1ab364c88c2be0833501a2db3651edf8e470a06a27cea212

Observation 9960dda1-b0ef-4a73-ad1e-49e17baeccdf · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.747447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.747447Z digest=sha256:99ebbd495b030ce992d2d9b00e4e60770ba5b6d41bd79f3e2c07e19f7f75d0ff

Observation 157abbdd-6fb9-4ddb-b113-ac938e7bf219 · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.818457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.818457Z digest=sha256:8a08b78ff3baf92ed277411b90d9f4cc9106501b931b2333f61805f9e011b966

Observation 73398d19-4860-4aa9-84e5-9c476ca629a8 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.917786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.917786Z digest=sha256:0e1b12f554041420e752d0fc02affe44772d67637798cf8ac10b99d065ab9069

Observation f7d06e82-c850-44e9-81cf-6b54fed0c298 · outbound

This paper cites X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:30.959960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:30.959960Z digest=sha256:a66f0fb8a8f565b19d699c656c723af046f001e4b034f3dc7a2741159e00f21c

Observation 91497f2d-7e15-4fbb-b50f-cee25b30a192 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.029300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.029300Z digest=sha256:852c649ccc6f8fb6b4cac5831c80d33799040da0997a40f0d07a3220930831ff

Observation 4afd7b9c-59bf-44ab-a362-c5a946216c51 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.102170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.102170Z digest=sha256:cd3fbe40b09fbe1b531c4082f44e25035b8bbcd702de574366588b65b6fcf864

Observation 7b63ecdf-9291-4b39-92cf-eb54e93d250e · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.319484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.319484Z digest=sha256:a4648654dd56b94b04ce14ca71f98e3bab03554e41d379915794adf348419e9d

Observation 4b2c6bd0-cc90-4cb9-a844-6e61c5acfdb7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.427858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.427858Z digest=sha256:6457bdd12a27de8aa43fbf902e6a585fdd1014e56080a2add65195c27867e326

Observation e864c572-b28a-4bf4-bf8f-900a67cd20d8 · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.519890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.519890Z digest=sha256:71a9235d931739c0c4daf47b6802f3295d77a05546c5ec259dcfcdca0bc21e19

Observation a2a54a8e-37d7-4566-a743-38c3d5bdfa09 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.607597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.607597Z digest=sha256:c0352f90201ec8de51fda4f4b0599742780e5663ec522f3dbc8727b73a4bd089

Observation 773a5ffb-195f-4335-af1b-349e45e0d3b1 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.741806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.741806Z digest=sha256:b49297c2f4b370ecd0d034924bc18faa1483c231a110d54b3ae6caa166c08cc8

Observation 064b171f-b589-4cda-9428-1e1d2bb06a69 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.932750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.932750Z digest=sha256:319c1490a9fd816f5367563e02feb7a2ec2c6e626460b865055283f071659f5b

Observation cb2d267d-00e2-47f7-9aeb-5b1c3b4e103f · outbound

This paper cites OpenClaw-RL: Train Any Agent Simply by Talking.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning OpenClaw-RL: Train Any Agent Simply by Talking

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.111226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.111226Z digest=sha256:a3e4ca71b598af74dacebcf894613286d665c9f0f048c71e6572918fc0e673c6

Observation cb3e8800-5cea-494b-ae26-5caf540394cf · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.308466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.308466Z digest=sha256:6135528d36f4c918d3606c59b1fb4416d959031491487ab10afac7a9360f331c

Observation 945e01cb-44cc-438e-b99b-edd7d0b760b5 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.508057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.508057Z digest=sha256:f7eec3275e2d14feef6065ae4ebf9685a694dc0d17528e3a34b77398dedc1ea7

Observation 5ba483d8-be47-44fc-a746-f91a8801b4f8 · outbound

This paper cites Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.617211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.617211Z digest=sha256:457896193b2ed4ef8e0e657a638a4a3b900c3bc124be0738d1479cbabceac289

Observation 9e394e04-c7f0-45bc-8e30-e4f8de9faa4f · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.737062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.737062Z digest=sha256:f5ece5e1c1a75964511e461a40194c0ccdd03917cc5574a1ba8768043171048e

Observation c871051c-eac8-464e-be0f-4efce5de4c16 · outbound

This paper cites Agentic Reasoning for Large Language Models.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Agentic Reasoning for Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.776319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.776319Z digest=sha256:ae48856613255deb63c0671084fe1e9b930df68d3822f9bbfb5a36a8343a8623

Observation bcf65ac5-2bd4-467b-b53c-074acb317f50 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.815219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.815219Z digest=sha256:97ed067c6863f10550f00077b4068ff931cbfe875f229b5ab85f3670e8a501cd

Observation 4c98dd14-f00e-461f-bebb-80a32febe6e1 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.853689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.853689Z digest=sha256:d56315f8e9a81895a3b6d4c902d2acf0cfadcda2f0857862ff7e2873896469d5

Observation c09ca760-3772-45a1-80b9-78f678767612 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning ReAct: Synergizing Reasoning and Acting in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.890559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.890559Z digest=sha256:69a7e090ba819ccf86f6b31ad7537a8999b57bceacf6ad70b47527042b6c7934

Observation 3297b684-4248-4c10-8cdf-c7210307d948 · outbound

This paper cites HomeRobot: Open-Vocabulary Mobile Manipulation.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning HomeRobot: Open-Vocabulary Mobile Manipulation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.954619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.954619Z digest=sha256:0e9053a79ead5725452509bd1a92eda755c8c7372f7e80bd82874f59a5eedf16

Observation b418a0bd-df24-4634-9464-13a4da7c0c8f · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:32.991522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:32.991522Z digest=sha256:22ca5b34e6fdcca9f9b3128f3c77bedf77c7073305e78ff364ffea53392253ea

Observation 106d3082-ea8b-412e-b1ae-1ac231548a4e · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.028513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.028513Z digest=sha256:0023f0fc3cc769a422af3517e09d728474b6792ae8ecc0fbf714fce8b2100806

Observation be864231-b730-46a7-acdf-c92c11cabaae · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.060330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.060330Z digest=sha256:0e9ce5c5523b4eaa3b1c3dfc96b9b80d6735a0394baf02703a7e0522954e732b

Observation 284279d5-12bf-425d-b36c-799523dce73c · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.092825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.092825Z digest=sha256:4e99939d00ff095f54572a31f6f6eee265650144fe207c1713b263c3b39126d8

Observation 20078f02-6ba7-405c-9c0c-58e27ae032b4 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.162251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.162251Z digest=sha256:6f79ea8c1cd8af80ea047f6c8ff21ca8648e52bdf8b0bcfc81f5fc4f1cacc838

Observation 479c00b9-a6ab-4484-b106-ecd04b4ee11d · outbound

This paper cites The doubly librating Plutinos.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning The doubly librating Plutinos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.192307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.192307Z digest=sha256:57dac7315c1f8cbd17c572705d9f41fd1e35bc864637ed949d25105182856b94

Observation cbdb7d06-b2f5-4131-94e2-7b3c64c7fe40 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.224055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.224055Z digest=sha256:abd062a6f95636457c8eb280b39d8b2c972528e3b20a16150a6e23d10ab34fc1

Observation 19789d3d-f449-4af7-a8a8-fa564f4626c9 · outbound

This paper cites A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning A Survey of Large Language Model Empowered Agents for Recommendation and Search: Towards Next-Generation Information Retrieval

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.128004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.128004Z digest=sha256:6567fcae6d477e1527d6e84820f4845247e57eaae646196db37c7f56603d693d

Observation 19ff2c2a-57ae-48cf-a597-819c5ef15596 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.307285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.307285Z digest=sha256:a095a00c019f1dba6813a4955acfed5b7fb2ba4767ffbfbf76f7933c237bbf8a

Observation 422feada-5b73-4fdc-8002-63a85f8d2062 · outbound

This paper cites SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.316386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.316386Z digest=sha256:a95e0e0d0e0dd2052e81f188e553eaeec202a84fa698c2d9679cea377f6ec834

Observation 0188d5a1-8c76-42b0-9bc7-0a7d4b590917 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.291734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.291734Z digest=sha256:bb2d06c7782ee2df440bec444255a2c37f8446531aff01f94281ff1368112b15

Observation 8006beff-8efc-4d38-94de-7308fdd69321 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.338196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.338196Z digest=sha256:79278f3d29759ca7f636193bab30f47b80d0fb4ca6ab08a07e99b6544975fe05

Observation ebeefab7-4b2c-4ed8-9225-d7d84263b6b1 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.362229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.362229Z digest=sha256:09adbd8e250ef15514f0db654709e38333c12d5cdd28fdb0e546858323a43edb

Observation 15578491-cf05-434c-8e67-15d42301103b · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.384019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.384019Z digest=sha256:b4c551e712576d69b6ba3ec61c7addecde41cc556e9e4671e73435f34a46235e

Observation 29cdaf4a-0c77-48df-b4d1-b46a3e74bfe3 · outbound

This paper cites an unresolved cited work.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.407391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.407391Z digest=sha256:39d88d700e3f1f1ba548ae13701ac64658e5e2373804988296263aa0a59b2045

Observation 00cdaf78-3c0b-40e8-88d5-b60f44f28883 · outbound

This paper cites •Do not output multiple<think>/<answer>blocks.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning •Do not output multiple<think>/<answer>blocks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:33.432754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:33.432754Z digest=sha256:147cf5acfc0d340d0ff7c4b2050714edcf20028736448cdfd104edea9aaf30ee

Observation f80b8360-9a88-48f3-9601-f6a549bd68e6 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.212369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.212369Z digest=sha256:218fa02e75335f38707fd8422a73ce1801dfde759c46ab2d8ab0e987675957b5

Observation 003a6523-95fe-4860-8329-fc4f1bbafd50 · outbound

This paper cites Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:29.759049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:29.759049Z digest=sha256:c599e6cde7336be19af131c5c2f0f40628a087a61754eede94bd2e749c18c951

Observation 66ea9cd1-ea2b-4dbd-8c5e-76a897a75c15 · outbound

This paper cites InThe Fourteenth International Conference on Learning Representations.

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning InThe Fourteenth International Conference on Learning Representations

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:31.665859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:31.665859Z digest=sha256:883875747d9b37488aff6822fbe1fbc99b7c4c27125f30afcabb5f958e52724c

Pith citing papers

No inbound Pith citation observations are available.