Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T11:59:18.223000Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2606.21925.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T11:59:18.223000Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 18eff5ca-7a2d-4808-834e-35318cf2d6cb · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control PhiBE: A PDE-based Bellman equation for continuous time policy evaluation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 04c12d0c-ce18-4373-bfe1-5e6fa0c88265 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control A deep learning-driven iterative scheme for high-dimensional HJB equations in portfolio selection with exogenous and endogenous costs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 934246e1-cb9d-45eb-84a2-3eee0650766b · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics) , volume=
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19e74262-7d46-4d61-9f3b-3f924dd8b227 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Machine learning , volume=
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20cdcf70-8ac3-493a-9c1c-b0e2849869f6 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Machine Learning Research , volume=
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ab243e7-2d46-42bf-8a0d-d1c89d485692 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Machine Learning Research , volume=
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1224526-dfcb-4d76-9ee3-2a8d49eec90c · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Machine Learning Research , volume=
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fbe1b95-bc16-432d-8202-4af28da38db8 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Machine Learning Research , volume=
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89c20810-68b5-41fa-9374-78d75d8fdd3b · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control nature , volume=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8778096-3981-4ca8-acda-f4cfc621209d · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control nature , volume=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bd96a63-0a59-4b12-a9b6-d48096270142 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Nature , volume =
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4550f723-c310-4007-93d9-8312435f2de7 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Proceedings of Thirty Third Conference on Learning Theory , pages =
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 750b08af-8a36-4a6a-bc19-a502a0cbc4cb · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Learning for Dynamics and Control , pages=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab6bb7f-f4bf-43df-9290-3949d92dd0ce · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2012 , publisher=
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c6a27e9-382c-4505-b43a-f8eebf07a442 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control On Bellman equations for continuous-time policy evaluation I: discretization and approximation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d9afcbe3-f083-4e61-88ed-de3303f542c4 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Advances in neural information processing systems , volume=
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6008699-7e41-4acf-bf61-0341000a0432 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0d74310-c4de-4f6e-ad4d-c97510bf9d10 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Proceedings of the 33rd International Conference on Machine Learning , year=
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ac63e5b-ca1e-4f7b-b4fa-71d0ef671faa · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control International Conference on Machine Learning , year=
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15ddb8c4-cefe-4f94-ae70-e530de688b26 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control International Conference on Machine Learning , year=
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19621f60-3ccd-415a-9074-9ebf79af2bcb · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Advances in Neural Information Processing Systems , year=
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 053fdde9-d793-47c2-a1db-ebc9126bc6d9 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control International Conference on Machine Learning , year=
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf1cc9bd-5d35-44ba-897b-676e9579f6ca · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control nature , volume=
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 367d7271-099a-4518-a20c-05bb9a19bd78 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Fine-Tuning Language Models from Human Preferences
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bb19c0c9-32fd-4e67-a652-f445688e1a6f · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Biomedical Informatics , volume=
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 855c7827-76bd-4b66-893d-37df84b4e8cd · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control IEEE reviews in biomedical engineering , volume=
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e36efa1-0d45-482c-bc05-fcd60a7accd4 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Advances in neural information processing systems , volume=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f44cb654-dfaf-4dc3-bdfe-6ab863782256 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Diabetes care , volume=
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49ebfd26-0602-4599-9382-2cdb3e85c80b · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2009 , PAGES =
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be9f6bc-7a30-4fbc-a4e1-1f40b10b3435 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 1999 , PAGES =
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf9aff3e-d1e9-4aeb-9f94-ef391324fb40 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2013 , publisher=
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08f09fbe-dbf9-4e3e-b8af-ef43f1c8c0b1 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2018 , publisher=
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b26388-1008-47df-86e6-f8ecf738da50 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Communications of the ACM , volume=
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b4a959f-98c3-4204-97ad-8d9828ae82f2 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Machine learning , volume=
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f58064-b45e-46ec-bc6b-7ec294a0024e · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Neural computation , volume=
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994c117e-a4a6-447b-8c56-40322671a9e5 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Machine Learning Research , volume=
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d865529-02cb-459e-bcad-e98ac21a55cf · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Mathematical Finance , volume=
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f8b2071-5aaf-4f28-b923-af70ebaf0bdb · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control SIAM Journal on Control and Optimization , volume=
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05eef019-2554-42e5-8703-382cf5b13121 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Machine Learning Research , volume=
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ff858d-e86e-46ef-bdcb-9c6803eaefde · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Exploration-exploitation trade-off for continuous-time episodic reinforcement learning with linear-convex models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 71230e14-7b31-4992-ae3b-8da3e7268cb6 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Proceedings of 1994 IEEE International Conference on Neural Networks (ICNN'94) , volume=
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06b65086-d6d3-4a33-a329-49e9f4da9ce0 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control International Conference on Machine Learning , pages=
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38dd8a99-14f8-4979-af37-682c35b998f5 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control arXiv preprint arXiv:2312.11797 , year=
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e6df4a6b-d4e8-4945-9440-a0f53987d181 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Accuracy of discretely sampled stochastic policies in continuous-time reinforcement learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a6779442-fbec-4060-a4f5-4003830d89e2 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control SIAM Journal on Control and Optimization , volume=
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbb84136-02e0-4a41-85f4-0d7c5d34707a · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bd6d39ca-0319-4470-b8c2-21cfa1cac476 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control arXiv preprint arXiv:2501.15910 , year=
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 03482a95-84e2-4a18-ae21-16ed0955cb42 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Scientific Computing , volume=
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb9bcc7d-18bf-4e12-a4ee-ebdfe57e34d7 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control siam REVIEW , volume=
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 530f72c9-89f1-40ed-bbda-ffbeb9089ca3 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e27d9c20-043e-4bae-b7d5-fa6858c9b9cc · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control GPT-4 Technical Report
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cd053fce-c077-4e74-b0cd-f5cb5fec808d · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c385a958-195c-47d4-ab89-c77c0ea35295 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Proceedings of the 38th International Conference on Machine Learning , pages =
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f7d6dff-f7bf-42f0-84a6-5afefbb7daa2 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control International Journal of Control , volume=
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe4782a-75b7-41fa-9ad9-429e017ec662 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Automatica , volume=
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a906685-70ed-4412-9a02-7adfaef95846 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control and Soner, H
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfd43f8e-7d29-409b-a7cc-eb86dd3fd42e · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Advances in Neural Information Processing Systems , volume=
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b18012b3-f462-4aed-8cd3-ef1127ebfc9e · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Optimal-PhiBE: A PDE-based Model-free framework for Continuous-time Reinforcement Learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 279677a0-3c9e-4781-aa95-284ab8c065cd · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Foundations and Trends
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 990d68b9-c7ee-4c6d-9d52-9ec1dcc31224 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control The Thirty Sixth Annual Conference on Learning Theory , pages=
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5ad3f98-74fc-4d20-86c6-6614bfe0eb44 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 1998 , publisher=
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c8c50dd-4ccb-4b62-90e5-f08ee8e9ca14 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Proceedings of The Web Conference 2020 , pages=
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92bee1c0-48ed-4617-b883-8a4cfed91be3 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2013 , publisher=
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feb815fa-22da-4f6f-b8db-a1a88838bbc2 · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2006 , publisher=
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 884ddb00-8c2a-446c-aabf-7dd0336dbc9e · outbound
PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2003 , publisher=
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.