Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:19:51.917733Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2502.01876.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:19:51.917733Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:21:33.975739Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-15T15:02:45.745385Z
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a9e2aed3-f7cb-4920-b02b-d9226a645ef8 · outbound
Reinforcement Learning with Segment Feedback write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeb90841-1345-4e41-b206-70885f44a7a9 · outbound
Reinforcement Learning with Segment Feedback Improved algorithms for linear stochastic bandits
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 48bb65b3-5523-46c3-8f38-e793d7315b3d · outbound
Reinforcement Learning with Segment Feedback Near-optimal discrete optimization for experimental design: A regret minimization approach
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 797df648-711a-4078-b2ba-0d730ef58bd0 · outbound
Reinforcement Learning with Segment Feedback Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 047d5f60-4c25-45bc-98ea-aad9dcd95919 · outbound
Reinforcement Learning with Segment Feedback G., Osband, I., and Munos, R
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f64a0e19-359e-4ce6-949e-26d36f838ad3 · outbound
Reinforcement Learning with Segment Feedback and Sundberg, C.-E
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9708ab51-edfd-4bf4-b011-f37a9bdfb6ae · outbound
Reinforcement Learning with Segment Feedback On the theory of reinforcement learning with once-per-episode feedback
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4254ac76-f449-49fa-b2f2-080697f08db9 · outbound
Reinforcement Learning with Segment Feedback Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e906830f-dac6-4f56-9ed3-365145e617dd · outbound
Reinforcement Learning with Segment Feedback Reinforcement learning with trajectory feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fb8f7803-616b-4706-98c9-9f1248569426 · outbound
Reinforcement Learning with Segment Feedback Improved optimistic algorithms for logistic bandits
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d7eccb-65ee-42ee-bcd1-d734a6f62da8 · outbound
Reinforcement Learning with Segment Feedback Parametric bandits: The generalized linear case
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a9985149-adc6-43e9-91ed-5dadc67c580c · outbound
Reinforcement Learning with Segment Feedback Harnessing causality in reinforcement learning with bagged decision times
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c2cb0f95-f8d3-4aef-9dcf-5b58cf0598e5 · outbound
Reinforcement Learning with Segment Feedback Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d422c481-9aa1-411d-8d41-79dc6dd9e9cd · outbound
Reinforcement Learning with Segment Feedback Near-optimal regret bounds for reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5b0bfe69-5799-4b8a-8fd0-7f183172296f · outbound
Reinforcement Learning with Segment Feedback Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b5dad29f-7bf6-47e8-92c7-26568c55f415 · outbound
Reinforcement Learning with Segment Feedback and Hutter, M
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9be24f8c-4c6a-45fe-b684-70e86e12cc16 · outbound
Reinforcement Learning with Segment Feedback and Massart, P
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 22347b28-8b4a-44a4-84d6-ac0ff433e6ce · outbound
Reinforcement Learning with Segment Feedback D., Jonsson, A., Kaufmann, E., Leurent, E., and Valko, M
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efd83852-8ee7-485f-a728-2325face270d · outbound
Reinforcement Learning with Segment Feedback and Moore, A
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 05dfcaf2-1f5b-4aa8-8b19-d3ae9b83aabb · outbound
Reinforcement Learning with Segment Feedback Optimal design of experiments
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f4451481-f3ee-475b-82b5-6f603bcb7e04 · outbound
Reinforcement Learning with Segment Feedback Self-concordant analysis of generalized linear bandits with forgetting
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 19dcb54f-00ab-46af-96f0-463ae0cb0835 · outbound
Reinforcement Learning with Segment Feedback Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed638af5-d215-4e19-996b-97ec7b8aa83e · outbound
Reinforcement Learning with Segment Feedback Reinforcement learning from bagged reward
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 84586d59-6cec-4806-a0e8-c061cd2a6dd2 · outbound
Reinforcement Learning with Segment Feedback Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cad5ac9-f212-46e7-903c-f6f62a6953da · outbound
Reinforcement Learning with Segment Feedback Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3bc721f4-5633-47d3-8fc7-93a12b879514 · outbound
Reinforcement Learning with Segment Feedback and Brunskill, E
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2f6a9377-259a-45e1-9173-21492ddfe58f · inbound
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Reinforcement Learning with Segment Feedback
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e5a35949-7b26-4bf3-b165-64b8ebb71333 · inbound
Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning Reinforcement Learning with Segment Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.