Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Segment Feedback

As of 16 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2502.01876.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01876 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:19:51.917733Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:21:33.975739Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T15:02:45.745385Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a9e2aed3-f7cb-4920-b02b-d9226a645ef8 · outbound

This paper cites write newline.

Reinforcement Learning with Segment Feedback write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.837280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.837280Z digest=sha256:0c8159ca5b079c4470122c492af97d74601805e1eefcf2fe93606a3801b25f5f

Observation eeb90841-1345-4e41-b206-70885f44a7a9 · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Reinforcement Learning with Segment Feedback Improved algorithms for linear stochastic bandits

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.176914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.842071Z digest=sha256:3eacdec267d82006124085f1f1adc66e5f9b5cc55cc3b2210d917c1030bc6a15

Observation 48bb65b3-5523-46c3-8f38-e793d7315b3d · outbound

This paper cites Near-optimal discrete optimization for experimental design: A regret minimization approach.

Reinforcement Learning with Segment Feedback Near-optimal discrete optimization for experimental design: A regret minimization approach

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.166256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.845372Z digest=sha256:210516fe9680fe6358d1c07f11aacf5628e8a2f64c825979f1c00e32765521ba

Observation 797df648-711a-4078-b2ba-0d730ef58bd0 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:19:52.155293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.848793Z digest=sha256:267ac47c812a1e5e06d7877c7cacf02b217c75879727a7d2dcf71006635f8093

Observation 047d5f60-4c25-45bc-98ea-aad9dcd95919 · outbound

This paper cites G., Osband, I., and Munos, R.

Reinforcement Learning with Segment Feedback G., Osband, I., and Munos, R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.142619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.852301Z digest=sha256:f1af833309d72c0499bed905a1d2b1c822e574d1e4f4ee7529e21894c4d1861c

Observation f64a0e19-359e-4ce6-949e-26d36f838ad3 · outbound

This paper cites and Sundberg, C.-E.

Reinforcement Learning with Segment Feedback and Sundberg, C.-E

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.131058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.855536Z digest=sha256:7971b7e71182222cf97b1f12e0fba1cb41b2f5a349a9647bb36e6bdbe367cb59

Observation 9708ab51-edfd-4bf4-b011-f37a9bdfb6ae · outbound

This paper cites On the theory of reinforcement learning with once-per-episode feedback.

Reinforcement Learning with Segment Feedback On the theory of reinforcement learning with once-per-episode feedback

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.119563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.859125Z digest=sha256:0776944a863aa5a7ddb86137054b1c31b1a319761612890b0f53f29708101de5

Observation 4254ac76-f449-49fa-b2f2-080697f08db9 · outbound

This paper cites Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning.

Reinforcement Learning with Segment Feedback Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.107444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.862300Z digest=sha256:6df05eb41331647d2e0e434c030453bb90bbf55717e2e4fd64b08a44c27e25d6

Observation e906830f-dac6-4f56-9ed3-365145e617dd · outbound

This paper cites Reinforcement learning with trajectory feedback.

Reinforcement Learning with Segment Feedback Reinforcement learning with trajectory feedback

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.096493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.864992Z digest=sha256:8419f5772904ae6d782f89a269dd6670532219ccb3690a956a6cbe89fb244362

Observation fb8f7803-616b-4706-98c9-9f1248569426 · outbound

This paper cites Improved optimistic algorithms for logistic bandits.

Reinforcement Learning with Segment Feedback Improved optimistic algorithms for logistic bandits

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.868220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.868220Z digest=sha256:ca51cc2b64f9e62d771612b67b48e31def706ad9d8653387dc61138ee87b650f

Observation d7d7eccb-65ee-42ee-bcd1-d734a6f62da8 · outbound

This paper cites Parametric bandits: The generalized linear case.

Reinforcement Learning with Segment Feedback Parametric bandits: The generalized linear case

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.080447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.871243Z digest=sha256:343e24441dd96b48a7d85eeb2b4ad109d4c905ba78dc8a21ba68436f595824fa

Observation a9985149-adc6-43e9-91ed-5dadc67c580c · outbound

This paper cites Harnessing causality in reinforcement learning with bagged decision times.

Reinforcement Learning with Segment Feedback Harnessing causality in reinforcement learning with bagged decision times

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.070439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.874352Z digest=sha256:5257ddcbad90777539b8f4e69602be58586acdbf53611ad71e07e05c658317b8

Observation c2cb0f95-f8d3-4aef-9dcf-5b58cf0598e5 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:19:52.058986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.877211Z digest=sha256:26a745d0d7d5164a0301b2004ff0d1af5a569e307359f7d4e97aacbfd4c5c86f

Observation d422c481-9aa1-411d-8d41-79dc6dd9e9cd · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Reinforcement Learning with Segment Feedback Near-optimal regret bounds for reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.048612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.879915Z digest=sha256:53e2d20c90cc913e691d6e3cbca21a9a6bd4d5482808de9d0147782894b4a409

Observation 5b0bfe69-5799-4b8a-8fd0-7f183172296f · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:19:52.038642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.882770Z digest=sha256:319c7ea2323e2ba516c64b03b0df11007f5a209d3f1d54c91edf4766dd956989

Observation b5dad29f-7bf6-47e8-92c7-26568c55f415 · outbound

This paper cites and Hutter, M.

Reinforcement Learning with Segment Feedback and Hutter, M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.028471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.886343Z digest=sha256:efe54f095b4c8ce17e259c5105568968a9c23814013d038f22577e106daf7c52

Observation 9be24f8c-4c6a-45fe-b684-70e86e12cc16 · outbound

This paper cites and Massart, P.

Reinforcement Learning with Segment Feedback and Massart, P

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.019473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.889347Z digest=sha256:751d5f036698c4ea4f8c294c3084c9efd8ca0632692a5cb4d14c98bba2408573

Observation 22347b28-8b4a-44a4-84d6-ac0ff433e6ce · outbound

This paper cites D., Jonsson, A., Kaufmann, E., Leurent, E., and Valko, M.

Reinforcement Learning with Segment Feedback D., Jonsson, A., Kaufmann, E., Leurent, E., and Valko, M

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.892457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.892457Z digest=sha256:672dc1543029d304f6afb525931e76f8eb756bb2a318ace2719324220563837b

Observation efd83852-8ee7-485f-a728-2325face270d · outbound

This paper cites and Moore, A.

Reinforcement Learning with Segment Feedback and Moore, A

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.003308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.895685Z digest=sha256:a6621e86a0aa0dfd29cc4ae649bcccc538570983b4dedcf262fff04a2b63ce8b

Observation 05dfcaf2-1f5b-4aa8-8b19-d3ae9b83aabb · outbound

This paper cites Optimal design of experiments.

Reinforcement Learning with Segment Feedback Optimal design of experiments

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:51.995256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.898605Z digest=sha256:6b1c3e5948eba56f3d44ec31f1d3fa8866b00cacb31e1dfd3a2a4e962636c5ce

Observation f4451481-f3ee-475b-82b5-6f603bcb7e04 · outbound

This paper cites Self-concordant analysis of generalized linear bandits with forgetting.

Reinforcement Learning with Segment Feedback Self-concordant analysis of generalized linear bandits with forgetting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:51.986161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.901392Z digest=sha256:181ef6711ca7b4d6145477264c4967b304c98765f49a0f63fa87aef505249504

Observation 19dcb54f-00ab-46af-96f0-463ae0cb0835 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.904455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.904455Z digest=sha256:21f2c608233aa680ab42ff796142db2a09f435f178aa26bf401a4d71edd1ad11

Observation ed638af5-d215-4e19-996b-97ec7b8aa83e · outbound

This paper cites Reinforcement learning from bagged reward.

Reinforcement Learning with Segment Feedback Reinforcement learning from bagged reward

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:51.970970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.908053Z digest=sha256:57e8f9915280fa715e45a1567da3b63f43e690d15a4dbf277919af116fb8221c

Observation 84586d59-6cec-4806-a0e8-c061cd2a6dd2 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.911511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.911511Z digest=sha256:fe4318dd00d2d466b455dd42b6a0f279f3b19bfbf452ea72d4464775029ae296

Observation 4cad5ac9-f212-46e7-903c-f6f62a6953da · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:19:51.954438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.914853Z digest=sha256:c2bd9c4d46e5a5007bbf6f8ff5843441f25c2e67e608242256c17d477d07c1cf

Observation 3bc721f4-5633-47d3-8fc7-93a12b879514 · outbound

This paper cites and Brunskill, E.

Reinforcement Learning with Segment Feedback and Brunskill, E

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:51.944380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.917733Z digest=sha256:b7d2c6501ec08cd52e42fc68a3bb5de21bf6938aef46629f1407e4f6c94e02eb

Pith citing papers

Observation 2f6a9377-259a-45e1-9173-21492ddfe58f · inbound

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling cites this paper.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Reinforcement Learning with Segment Feedback

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:02:45.752291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:02:45.617996Z digest=sha256:49aa70801696202bb64923ab9d8a39a0a23c5e2ba0caccb4ca43593e1d8f3145

Observation e5a35949-7b26-4bf3-b165-64b8ebb71333 · inbound

Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning cites this paper.

Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning Reinforcement Learning with Segment Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:33.975739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:21:33.975739Z digest=sha256:507145562846ff46e6825cc71155e88624084f67a73f1dd19713354e8d41a8c8