Pith. sign in

Paper Citation Record · LEDGER

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods

As of 16 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:1908.03263.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.03263 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T14:25:39.636614Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T11:33:20.892688Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T11:33:21.562600Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 191a5e9e-b75f-46ec-943b-b200e16d508c · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.924948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.177695Z digest=sha256:5aca783fa25b742bf618301b04fb8245a79617246428ab946f8f66ef29c3cb3c

Observation 7da8213a-70aa-4ee0-829a-f224c1aae600 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.897151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.187513Z digest=sha256:41f2f0c98539393f684146a646137775427be245150c0f7fab45522acb952c36

Observation b3ca9d18-1357-4dfd-a6ee-27d6b2e53085 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.869973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.196950Z digest=sha256:3b740267cac3e7eb30e39b9b77c57f77fcc4cdff40ee629668ae609333b418a7

Observation cc2a877b-8086-408d-9614-bd9979352332 · outbound

This paper cites Peters and S.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Peters and S

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.847301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.206498Z digest=sha256:b42577287d4e05aab89ac144fcc0878f3e084a12c29343b3157815ceffe7bd71

Observation 34cf5d5e-26cf-4e57-a13e-f01d62e27c77 · outbound

This paper cites Schulman, S.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Schulman, S

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.213978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.213978Z digest=sha256:bd2d6f28e17b8be79bf733e507cd51079b997c4080ae6d05f922577ff8339e99

Observation 57f71c96-e1ba-4f92-a4ec-cdf0b63a3ba2 · outbound

This paper cites Cheng, X.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Cheng, X

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.804769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.222566Z digest=sha256:ed1a0999db9f318c488f8a578454efc0484bafe18762871f9b541c704976241c

Observation d31b7ebf-c3c0-4b7f-a6d5-0fa7aa35bdee · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.784745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.234559Z digest=sha256:9ee3f4839ea08f0cd5da2a2829e577f15c4a73398f6da494fc297ee0206682c4

Observation 71803ba2-1efc-48fb-bd61-c9da035f4473 · outbound

This paper cites Cheng, X.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Cheng, X

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.750048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.245548Z digest=sha256:5a6fcf58a9b117bd7e6b7081be4a5c634d8cae9828bed30c06840cde3d85e60a

Observation 09f63648-3278-48c7-8939-de72168efcb4 · outbound

This paper cites Policy Optimization with Stochastic Mirror Descent.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Policy Optimization with Stochastic Mirror Descent

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:25:39.867071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.260813Z digest=sha256:7e9db5e6684b07cc097e4fdb2485ed4499fbe765ac802cc945dcd7d3fb67f997

Observation d626b641-55e4-4d66-892f-f924b8a01997 · outbound

This paper cites Ghadimi, G.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Ghadimi, G

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.725517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.274695Z digest=sha256:ce2970c9cad881ea31fb253de43c356d7f8f92ad1ae3ec6c0dfaf5ba4f180612

Observation a671989a-ef87-46c2-8c7c-a3c28e61d926 · outbound

This paper cites Kimura, S.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Kimura, S

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.702491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.282947Z digest=sha256:ec567491cabf70c4025b83801bff8d3e784cd0aa71317474a847f008e5c17cee

Observation 5c46c7f7-eb94-4c5f-bd5e-6990e74346a4 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.676747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.291379Z digest=sha256:cd1ddefa31b028bf88a140d353c4cc1f56107ea01823f25c5469e1f1b3b475bd

Observation 9aa4f302-9e25-4fe3-a30e-1ca7475dc66a · outbound

This paper cites Silver, G.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Silver, G

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.650996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.300911Z digest=sha256:170c3c34e69834c16c2729892f52d204ad082212462b86e4c17ac14826cfde54

Observation a7042fe8-6800-4ca3-9fbb-cd57f2745822 · outbound

This paper cites Schulman, P.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Schulman, P

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.616211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.306683Z digest=sha256:f39623cc9b32ec9df21c3d77eeb2a05505025cc7fd4296982cfe9151fec4e97d

Observation c6273e3d-8564-4baf-a028-7b905328105e · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.593088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.315864Z digest=sha256:af294b2fe262c70ad8a7bf29dae5f9010dff841bc75cfcb72b6ab962c778558f

Observation 80be2f38-460e-4332-b502-7de517abca12 · outbound

This paper cites Efroni, G.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Efroni, G

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.570376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.322420Z digest=sha256:a9b016329de58ff40d3d7b3b19e53cde9dbd697bf4aa85828d555452c58076f4

Observation f8d14759-c79d-4cfd-910c-c571bb326be1 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.551194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.328381Z digest=sha256:a73f586efaef9e32170689c8787ed6e6a3d33e315d961914709f20a0f585b7bc

Observation dd6fce16-93ba-4e1f-8598-cb6ec1a9da50 · outbound

This paper cites Greensmith, P.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Greensmith, P

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.531336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.337728Z digest=sha256:94730194f28604d8e64ad6ff1642cff7aee989861cbd313c4b0ef8699b32c5f2

Observation e1079437-cff7-41ef-9154-0d8c8d1d818d · outbound

This paper cites Jie and P.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Jie and P

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.508191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.346992Z digest=sha256:72561108ce0158f98b2372851ad7f5f8d6180f9b9e52669cf58b9182b0327902

Observation 24bf1f2f-8d04-49fe-b23c-be04665ed17b · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.483707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.363164Z digest=sha256:7c314e0d4225d25af45004238363b44e15608fabb9479ca10906365b2a421e9f

Observation 7ee92dd1-df7f-4d7e-a123-7650d1098d6d · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.462850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.370748Z digest=sha256:de43c0d9fc5eaa5a5bacead0e9a4a0b9d48c002c2826e61dba1180ef27a66ceb

Observation f864052a-b272-4a22-885e-ba3824d912c0 · outbound

This paper cites Grathwohl, D.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Grathwohl, D

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.441686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.386784Z digest=sha256:9319a152d2d403242b2d99f397a93ebbb32f799cae1c2db6840036acb853a6b5

Observation 8a087b20-2c97-45a4-8fed-e943814e64e2 · outbound

This paper cites The Mirage of Action-Dependent Baselines in Reinforcement Learning.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods The Mirage of Action-Dependent Baselines in Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.395346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.395346Z digest=sha256:0d4029fdf256e5104331b38ae51a534dd53a86293dc9ba1a616c2a6de58dcbb8

Observation 371e0090-7036-4b0c-a91b-ec496ff7e2e3 · outbound

This paper cites Reward-estimation variance elimination in sequential decision processes.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Reward-estimation variance elimination in sequential decision processes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.404839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.404839Z digest=sha256:11b13725194f2c4a81192eab8a12ce74a09a5d620ee56120f3c6b14d9d88fb2b

Observation a2db03ec-246c-485a-a25c-17b62980cfc2 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.417372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.418128Z digest=sha256:e246bead4862de4fdc19e5e6af62215ae9d6895ec51a25846bb2b9075f60d8ca

Observation c8c16567-bff4-4ed4-a284-9f49c2b7c3b5 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.397483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.425024Z digest=sha256:d6999dc0d4d50f5cd55b13567ecb27785dcdcc8eba81299e85057f1a98a21ede

Observation 83f44b55-f1fe-4b9b-9aac-c46a2fb6e61b · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.365337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.434584Z digest=sha256:9463be4e326b20ef2cf2a4b6a81d3e5616d7421262c701d72a4202c03837f438

Observation b4090b9c-e85c-46ae-a45d-ddb4d4f6ca6b · outbound

This paper cites Beck and M.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Beck and M

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.440612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.440612Z digest=sha256:a7bfb161bf04fd0fb20253ae468e347d3ffa4dc1db889bc78a4758c28bfd4570

Observation 07ba154f-3f7f-4ce6-831b-813832b46c4f · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.317137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.449838Z digest=sha256:68b4d3df4ffda6310df8d7dba4cd8ab39b253efa6eeb891e3b003c17432c447f

Observation 2c2a25d3-496d-4314-be2d-c416db887245 · outbound

This paper cites Vemula, W.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Vemula, W

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.284582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.464991Z digest=sha256:0cd76314da77486bb6dfae504a498b897992ac738fb711e8f3fb6e6ec0935a75

Observation 38408601-f664-4a24-bb12-efedbf9cadeb · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.258593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.482592Z digest=sha256:72b7974dca69789527f3543862558383fda17b60d4ffc0e4abfeea45aace0d87

Observation 895ec48f-b0a8-4842-8570-d3adb223217c · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.491821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.491821Z digest=sha256:751aff8d931821e8f41a14847255464de50ebf9aeb5cc53eae87a905016daccf

Observation 90fad94f-fd03-4d05-a2a9-3772dacca61f · outbound

This paper cites Schmidt, N.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Schmidt, N

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.499332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.499332Z digest=sha256:1dde0ea3cbc010e4833a0e11d9f730a138909cd9ffff952d15ca3e51e8916e2d

Observation bceff0ea-6228-4960-823e-e87165ca94b5 · outbound

This paper cites Johnson and T.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Johnson and T

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.199064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.506390Z digest=sha256:45d8d6296447cd126269f295a862ad66a098011aef12b02fb7c61e8854a99c9f

Observation 3b37b103-d5da-4510-b3d3-50c58d4bc5b0 · outbound

This paper cites Defazio, F.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Defazio, F

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.168026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.514204Z digest=sha256:18762625f8027b663b013fe7b70141928eb7ee1f31a3127235c7af4d2dc647a2

Observation 1df63345-774d-44a6-8c82-8065e0835dd4 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.132969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.520508Z digest=sha256:335639875380899ae4212c701aebcd40a6c514f04bbd77d79a59feb6008193af

Observation 9627ad71-871a-4bb5-8b6f-3f4df6e5262a · outbound

This paper cites Expected Policy Gradients for Reinforcement Learning.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Expected Policy Gradients for Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:25:39.759551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.540027Z digest=sha256:4330b027b0e169832449c55c14270cc396245e8d4f316cb8ad28035876bfb3f7

Observation 14a9e677-16d0-4d01-9b91-5cb3b06f84c7 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.105382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.547750Z digest=sha256:328d5473e3c5fcc3e8b5b07951e7bd0ff4c3efd1703e0dc9dc30c02f2756e40b

Observation 6716a93d-b318-4de0-9902-a434c6d3db5a · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.069970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.568347Z digest=sha256:c416d7822fc4eeb797cf143e4b5577c4911d4a72bc1d78c86108e12e36a85189

Observation af4d472a-85e8-4970-b13a-003c7530376e · outbound

This paper cites Baxter and P.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Baxter and P

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.038358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.579352Z digest=sha256:cdfc82f1eb2c86ae9336dc9de505ff523ae8d40efda015f73357f96dbd0a1bc7

Observation 0985f900-b8ab-4169-88f1-305275137236 · outbound

This paper cites Landau and E.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Landau and E

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.000894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.593945Z digest=sha256:b92d506bdaeaa21caa6a0261e3c5ecddb5768c660e7cb2a71280a02310b93c46

Observation 9e8624c0-204f-46eb-b5ab-034b6abb15d3 · outbound

This paper cites OpenAI Gym.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods OpenAI Gym

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.601235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.601235Z digest=sha256:050654b575146903e5c3c79b71558b468ef7d3554c9496bf7d3a9b7a757c12de

Observation 62138545-a281-4e10-87c7-5e8f3b0c6a9a · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:39.977301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.612620Z digest=sha256:1294e4fb52abc8f8f0aae43df759e3332b956542194ad073ec29cf6d6de01711

Observation bba63ec3-e7a8-4190-8c9d-92853e5157e5 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:39.946075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.619676Z digest=sha256:1763b1512ccc2f64f9d7ca0c8aef25f1f6d881e65e66949bcc16da812e006934

Observation b27d937a-603c-4b9d-bc23-642f3b55db4b · outbound

This paper cites That is, a feasible ordering must be causal at least in actions: the action randomness that causes a state must be arranged before that state in the ordering.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods That is, a feasible ordering must be causal at least in actions: the action randomness that causes a state must be arranged before that state in the ordering

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:39.922455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.627640Z digest=sha256:be6d568ec185e62bb63890a8c724af4bd2886e09bdbcb60f2cc845e9dc423e4f

Observation 4d2ba498-63b1-4513-bc65-d36c41ce8afb · outbound

This paper cites We consider the following operations (a) Suppose, in an ordering, there isSv→Su,v >u, then we can exchange them without affecting residue.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods We consider the following operations (a) Suppose, in an ordering, there isSv→Su,v >u, then we can exchange them without affecting residue

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:39.894577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T14:25:39.636614Z digest=sha256:236c04f6384662e7fd705da21ca745d8ac4c0c8e4b6a03f4d8c027e932584ff8

Pith citing papers

Observation 2ac05bc8-c0f4-4185-b19b-7e362f2cf794 · inbound

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems cites this paper.

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods

Reference 282

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:33:21.565277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-11T11:33:20.892688Z digest=sha256:75980b67b5516fec4193a409277bea1c2de6d2491239aecfa730d61477964acb