Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Segment Feedback

As of 16 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2502.01876.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01876 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:19:51.917733Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:21:33.975739Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T15:02:45.745385Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a9e2aed3-f7cb-4920-b02b-d9226a645ef8 · outbound

This paper cites write newline.

Reinforcement Learning with Segment Feedback write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.837280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.837280Z digest=sha256:328804f795ce38039309bd8b61b9881cc1ab96e694c98f8723e13c94aa3748ac

Observation eeb90841-1345-4e41-b206-70885f44a7a9 · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Reinforcement Learning with Segment Feedback Improved algorithms for linear stochastic bandits

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.176914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.842071Z digest=sha256:acaca13562ff872fe7a88b46af3230316a2bfe59138c2951395981898ad82639

Observation 48bb65b3-5523-46c3-8f38-e793d7315b3d · outbound

This paper cites Near-optimal discrete optimization for experimental design: A regret minimization approach.

Reinforcement Learning with Segment Feedback Near-optimal discrete optimization for experimental design: A regret minimization approach

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.166256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.845372Z digest=sha256:2b6f5022259bc1155cc6f82cd7aa744d803dc2cbd10239c08ea705ca146c1b86

Observation 797df648-711a-4078-b2ba-0d730ef58bd0 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:19:52.155293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.848793Z digest=sha256:77fc4639213c67784f13d08d5491b036c9d6aaf5afa2993e66f62beb1a35204c

Observation 047d5f60-4c25-45bc-98ea-aad9dcd95919 · outbound

This paper cites G., Osband, I., and Munos, R.

Reinforcement Learning with Segment Feedback G., Osband, I., and Munos, R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.142619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.852301Z digest=sha256:6ba0f4d9947daef2cd6264da174c920fac124ed39efad907bb523de44e125af1

Observation f64a0e19-359e-4ce6-949e-26d36f838ad3 · outbound

This paper cites and Sundberg, C.-E.

Reinforcement Learning with Segment Feedback and Sundberg, C.-E

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.131058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.855536Z digest=sha256:cdd2dcffc211affd341a54ee19fa06f90c4049eb4f3438e690db1cb82d4d12ef

Observation 9708ab51-edfd-4bf4-b011-f37a9bdfb6ae · outbound

This paper cites On the theory of reinforcement learning with once-per-episode feedback.

Reinforcement Learning with Segment Feedback On the theory of reinforcement learning with once-per-episode feedback

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.119563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.859125Z digest=sha256:af64e722ff153b7995b1fa67c3769be6c1b83220cd021ad216585a07feb32db9

Observation 4254ac76-f449-49fa-b2f2-080697f08db9 · outbound

This paper cites Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning.

Reinforcement Learning with Segment Feedback Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.107444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.862300Z digest=sha256:afcf4446731717bddfc4363ef7bd94c92b4ba340b7d8bb4f7c499c8f0eda02a6

Observation e906830f-dac6-4f56-9ed3-365145e617dd · outbound

This paper cites Reinforcement learning with trajectory feedback.

Reinforcement Learning with Segment Feedback Reinforcement learning with trajectory feedback

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.096493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.864992Z digest=sha256:bde24df127e53050bd0cfa4129b61bb48e7d61bf2aed706601734654b5bc8365

Observation fb8f7803-616b-4706-98c9-9f1248569426 · outbound

This paper cites Improved optimistic algorithms for logistic bandits.

Reinforcement Learning with Segment Feedback Improved optimistic algorithms for logistic bandits

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.868220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.868220Z digest=sha256:ddf920157d912793d80c9fcc99c1867095a87ee3665228972ab0f0b3ed8ed997

Observation d7d7eccb-65ee-42ee-bcd1-d734a6f62da8 · outbound

This paper cites Parametric bandits: The generalized linear case.

Reinforcement Learning with Segment Feedback Parametric bandits: The generalized linear case

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.080447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.871243Z digest=sha256:d5c457d2dac42e4e866f209ed7d8bbdda4875b5cd322d9cf03c3c09a98913463

Observation a9985149-adc6-43e9-91ed-5dadc67c580c · outbound

This paper cites Harnessing causality in reinforcement learning with bagged decision times.

Reinforcement Learning with Segment Feedback Harnessing causality in reinforcement learning with bagged decision times

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.070439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.874352Z digest=sha256:c2a949bf044432ca367929a215c692ddbab0703456e01044afd812ef581583ec

Observation c2cb0f95-f8d3-4aef-9dcf-5b58cf0598e5 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:19:52.058986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.877211Z digest=sha256:02a6ae9b5ec06400f626aa7a7042aa9cba969ba3d58e913692e29c1b329631d9

Observation d422c481-9aa1-411d-8d41-79dc6dd9e9cd · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Reinforcement Learning with Segment Feedback Near-optimal regret bounds for reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.048612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.879915Z digest=sha256:fcdccf53003ea1758c05c722e4f961ed6d8931292a6a03bf54a0d5ee31097d70

Observation 5b0bfe69-5799-4b8a-8fd0-7f183172296f · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:19:52.038642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.882770Z digest=sha256:b0bf5dac789c4e5c7bc0ac5f49635feb67858dc98ebe126f2cc149830e6c2b21

Observation b5dad29f-7bf6-47e8-92c7-26568c55f415 · outbound

This paper cites and Hutter, M.

Reinforcement Learning with Segment Feedback and Hutter, M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.028471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.886343Z digest=sha256:e8a27369d5da14b76aafb8216d795ab3e9cf8a1d876529e44b38cd1050d09143

Observation 9be24f8c-4c6a-45fe-b684-70e86e12cc16 · outbound

This paper cites and Massart, P.

Reinforcement Learning with Segment Feedback and Massart, P

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.019473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.889347Z digest=sha256:7ee32378e6a2755d8e31805a9c45a7c4d370dda47aa3410f890f87f508254283

Observation 22347b28-8b4a-44a4-84d6-ac0ff433e6ce · outbound

This paper cites D., Jonsson, A., Kaufmann, E., Leurent, E., and Valko, M.

Reinforcement Learning with Segment Feedback D., Jonsson, A., Kaufmann, E., Leurent, E., and Valko, M

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.892457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.892457Z digest=sha256:c708d3d522ff78c5b6629ff13e4f87e87f3482eae3d2a384beac25581353a72a

Observation efd83852-8ee7-485f-a728-2325face270d · outbound

This paper cites and Moore, A.

Reinforcement Learning with Segment Feedback and Moore, A

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.003308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.895685Z digest=sha256:ec5bf1ed42ec80dce0c068ce6d953cd8a99a9e572b5830147eb19687f39d38d3

Observation 05dfcaf2-1f5b-4aa8-8b19-d3ae9b83aabb · outbound

This paper cites Optimal design of experiments.

Reinforcement Learning with Segment Feedback Optimal design of experiments

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:51.995256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.898605Z digest=sha256:3b2f13d1ed7829909b65623ed095ce562a82159b7fb9f05e886134039dae3539

Observation f4451481-f3ee-475b-82b5-6f603bcb7e04 · outbound

This paper cites Self-concordant analysis of generalized linear bandits with forgetting.

Reinforcement Learning with Segment Feedback Self-concordant analysis of generalized linear bandits with forgetting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:51.986161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.901392Z digest=sha256:3621e0b3442ee64b49a8ff829e82de2bee9bea120b5ceb6d3bf271744436deaa

Observation 19dcb54f-00ab-46af-96f0-463ae0cb0835 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.904455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.904455Z digest=sha256:134404af719408062d586c2205d9a726115bbf6f929760f2eadfe89f7866cfcc

Observation ed638af5-d215-4e19-996b-97ec7b8aa83e · outbound

This paper cites Reinforcement learning from bagged reward.

Reinforcement Learning with Segment Feedback Reinforcement learning from bagged reward

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:51.970970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.908053Z digest=sha256:14b21ec7fc81150ce414e2b3faf7da66d42dbcd907f7bd97b1eb5c9c8f2d92e2

Observation 84586d59-6cec-4806-a0e8-c061cd2a6dd2 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.911511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.911511Z digest=sha256:43318329ab9debb6522557f5ab1081fe8624be8ad238187d200106517ad58c8c

Observation 4cad5ac9-f212-46e7-903c-f6f62a6953da · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:19:51.954438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.914853Z digest=sha256:c69a2e3ff6ef4f04770844b35916e06666f27fd64a3cc573bd9bf3d6ee1703d9

Observation 3bc721f4-5633-47d3-8fc7-93a12b879514 · outbound

This paper cites and Brunskill, E.

Reinforcement Learning with Segment Feedback and Brunskill, E

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:51.944380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.917733Z digest=sha256:0916cdead8f587eccddfcd684954357a0eb53c412124e2b87e7cf04d4e4d9a02

Pith citing papers

Observation 2f6a9377-259a-45e1-9173-21492ddfe58f · inbound

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling cites this paper.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Reinforcement Learning with Segment Feedback

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:02:45.752291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:02:45.617996Z digest=sha256:02925d558104059cf228bf6fc6611230cc0496ebce010969ae22ac721b7468bf

Observation e5a35949-7b26-4bf3-b165-64b8ebb71333 · inbound

Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning cites this paper.

Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning Reinforcement Learning with Segment Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:33.975739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:21:33.975739Z digest=sha256:9a126a5f60af99535fb01b8325eee2fa8c2d5da41828b3e0aa9f884ef93795f3