Pith. sign in

Paper Citation Record · LEDGER

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards

As of 11 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2501.14513.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14513 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:10:48.724636Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:27:25.150807Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T08:50:59.092753Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2a6c13df-86bb-4ab5-844b-fa927d35105c · outbound

This paper cites Loquercio, E.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Loquercio, E

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.353788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.602656Z digest=sha256:d37ec7e0a294189dfec50fef0f379fdcede77290502c8b17904ae1446e9f6a54

Observation f39976c4-480f-430e-901b-7313d4e1f02f · outbound

This paper cites Loquercio, E.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Loquercio, E

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.342872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.606745Z digest=sha256:51e753090d792e9c31fb82fd42ea5c0fd218825b1c423742dc21c323bc775915

Observation 45562607-49e2-4e8c-84af-fb7895d5b67d · outbound

This paper cites Kaufmann, A.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Kaufmann, A

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.333036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.609915Z digest=sha256:3023df9e96d445b95beec6d1a2d44e0ec684e8b204d2a7e96deda5667e240de0

Observation 76dcdfca-35c2-4c1d-9b97-52e3174b8cc8 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.321381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.613006Z digest=sha256:347617c7a444a19a9e3cd318958ab358ffaf127dbf81d39c8f1e5cf6154dca35

Observation 347fb54a-7ece-4caa-bd65-994b173fdad7 · outbound

This paper cites Back to Newton's Laws: Learning Vision-based Agile Flight via Differentiable Physics.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Back to Newton's Laws: Learning Vision-based Agile Flight via Differentiable Physics

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.616444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.616444Z digest=sha256:862850b63644b0e2c1e86a47aa9058f3e132831ce39aac13c779c584cf33f27c

Observation d4b199b2-798a-4c38-a549-84db8edd5f78 · outbound

This paper cites Wiedemann, V.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Wiedemann, V

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-10T15:10:48.620025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.620025Z digest=sha256:9ba6ad39f3b96c056781d341044a32915f278d59d60baa82e11ea35567a903c7

Observation cff916ca-fe64-46f4-8017-e934760916cf · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.311453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.623414Z digest=sha256:4fe888263bbf7303108963ae6a62d343c0db3b921fa74370bad6a0e1954f1fb4

Observation 8f7c4c75-7273-4fe9-a73d-23ee9732b71f · outbound

This paper cites Learning Quadruped Locomotion Using Differentiable Simulation.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Learning Quadruped Locomotion Using Differentiable Simulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.626404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.626404Z digest=sha256:9bac218d4f72c82dbb631d113cdb28e85c5a0448881bec3cc60360f66f6b92fb

Observation a3ed2d12-91a8-4c4a-b677-4296a8ba90f9 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.629534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.629534Z digest=sha256:79f5fe949c9e69f01606dc0cfa2e2ac5f8ed64776aca603229479bfd782c0b82

Observation 7add1830-6e13-4a1a-83ea-629e8f04fa1c · outbound

This paper cites Zhang, W.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Zhang, W

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.301955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.632272Z digest=sha256:f5d57142fb76bf9c95e55dbc55b2c351870282c14e5fe7be4ea01994a112635b

Observation 400c41c8-6704-48e0-a6bd-249558b1be33 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.293561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.635581Z digest=sha256:4a64b8f976deea2a06121f2d163892838a4b5fa54291e68bc549de10ba31ba0b

Observation fd32f32e-a259-4cee-a19a-26e41e0e3e02 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Playing Atari with Deep Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.638379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.638379Z digest=sha256:94256be235fcd32efc0e196099388c9d9cb1be0d2438180bc637567554c942f0

Observation 7541952c-6c51-45bb-bd5c-422ab98394e1 · outbound

This paper cites Continuous control with deep reinforcement learning.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Continuous control with deep reinforcement learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.641907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.641907Z digest=sha256:3878dc20c4dbc07c412a747bf04a19d39c06127898b0458efddb28b90e542d20

Observation 17aeecff-0047-4d6d-8480-85f28b681b0a · outbound

This paper cites Fujimoto, H.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Fujimoto, H

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.284169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.645239Z digest=sha256:a1a834267f9210d24f46eb5857f6f44fa1cd21bbe13458291c369fa87073968b

Observation 4a649408-30d0-4204-af3c-9a67ed14f5ab · outbound

This paper cites Haarnoja, A.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Haarnoja, A

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.648213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.648213Z digest=sha256:7592ceb3cf1d0cb134bcf574c98e1d44ea5f75b7960a1474692333160d005ee6

Observation 7a4a7514-f73b-4c86-aac7-ed59c30b87bb · outbound

This paper cites Trust Region Policy Optimization.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Trust Region Policy Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.651760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.651760Z digest=sha256:80bf0ddcabc6cf01a8a382e8fb35a86828f0781547b11e0367ad60de76210384

Observation 19e6577e-580f-4d35-b307-cbe404a74168 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Proximal Policy Optimization Algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.655628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.655628Z digest=sha256:96662492a6cb1ebcf183a47b11cb8b423ab471d24247386d2f828214a42dca8d

Observation 3282080b-e644-4e22-ae24-a23756b66830 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.268170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.659274Z digest=sha256:05a6ae1f55d546ca966c6952517939d7f0fdc7034d0068b77a1f81aac48fb89c

Observation 33b0d98d-fbac-4889-8b3d-984f6a4777a7 · outbound

This paper cites Deisenroth and C.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Deisenroth and C

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.663178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.663178Z digest=sha256:ef497caff554bc546233a6f10afd5e20285f0b61ee7648725d1cd73d854a9a4c

Observation 89d80766-acd2-4fc4-ab74-39af80008f1c · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.251386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.666518Z digest=sha256:30a67af5f48e04851225238c7fbaf58a6646c27fe9e26d481896adf870d275fa

Observation f128fc34-672c-413b-9455-c0091379df41 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.669930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.669930Z digest=sha256:61f1ae44f8089d3df70cfef5381ed398b052439d502f29b0f7f3b0eb0e7ee55c

Observation f80059c9-6d7c-4d3c-90b7-61df54726a2c · outbound

This paper cites Watter, J.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Watter, J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.673338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.673338Z digest=sha256:c77d899139c8cff3e9c62e4384c6fbb9950f4e6d33204c6d7a7b6d4a7a89777d

Observation 44269d92-9f4d-49a8-a6a7-8fa60d2121d6 · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Dream to Control: Learning Behaviors by Latent Imagination

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.676127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.676127Z digest=sha256:d4a77075da1c45ba3503fd0c5729c2b5a7423c3e501f43575a521548c99b0e86

Observation 16ba7ee5-3adb-4f6a-b810-88682f641455 · outbound

This paper cites DiffTaichi: Differentiable Programming for Physical Simulation.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards DiffTaichi: Differentiable Programming for Physical Simulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.679345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.679345Z digest=sha256:5660c59d75f3373425ef4d7122b6c91f8dacf01256334d97ed49a905dfbe6f73

Observation b24d34d9-1fe3-423f-88c8-e060a774792c · outbound

This paper cites Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.682411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.682411Z digest=sha256:a7628a5986b13ad8dfe628b206188f540a29b078e04e1c436219e48a6dd09ce6

Observation b65633a5-3bf1-4189-95f9-b7995508fb8d · outbound

This paper cites Todorov, T.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Todorov, T

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.685299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.685299Z digest=sha256:8f22ad6e0a4c2859f54ba401c7bc912d311f89713bf8e6b5d04f9159081a9afb

Observation d9036fca-56ba-4228-bf6a-1147d597a461 · outbound

This paper cites Heiden, D.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Heiden, D

Reference 27

Resolution
metadata mismatch
raw_fallback, observed 2026-08-10T15:10:48.855662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.688100Z digest=sha256:12eaa0a19fbfbcb83ea811aa83a85b4d4a6dfdb93444e58ad891a94c67bee5f4

Observation fe272e94-86ba-45a1-8198-9a65ea4d0140 · outbound

This paper cites Dojo: A Differentiable Physics Engine for Robotics.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Dojo: A Differentiable Physics Engine for Robotics

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.690868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.690868Z digest=sha256:348015063c70f37876060fb68fc2537323322941cec0b8f92f1ec864df8487ff

Observation b5ac3004-5f7c-4700-8691-2ef5ae39098e · outbound

This paper cites VisFly: An Efficient and Versatile Simulator for Training Vision-based Flight.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards VisFly: An Efficient and Versatile Simulator for Training Vision-based Flight

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.694231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.694231Z digest=sha256:65b149950ca13a772a10aff47d80867b21f7a0e9b02be13632b5b43ba7f81754

Observation c9b09ff2-12dc-452b-874c-933a5043cb92 · outbound

This paper cites Savva, A.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Savva, A

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.230830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.697111Z digest=sha256:6823efbd5cf0afda6fcec0808d588ad0cb149319e7326739044cbd4269b5e9b5

Observation bd1987eb-84d1-454c-915c-ba001b222dad · outbound

This paper cites Schoenholz and E.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Schoenholz and E

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.220966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.699836Z digest=sha256:715a1c0fa6653c014e533f11b8093344a9360e2526e3053159de54f62cb432b0

Observation 55011a73-a570-4ff0-a2b4-089610a5fbee · outbound

This paper cites Paszke, S.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Paszke, S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.702421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.702421Z digest=sha256:bb1501b32f423e83a4dc0b6b62af474de4acf391f1ab15ecaa15263d55404585

Observation 7239ef10-55ae-4a66-8fbe-5d8e30adfe93 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.204429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.705121Z digest=sha256:c067f739cd627f3a279cd3fed5efbfbb3cffaa3aeae65b1fd9792a2f223bd938

Observation f66c94d8-eb33-4d30-81ea-cca8aa979c40 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.194264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.707788Z digest=sha256:5f59e6a2d8ffea4012d23908618f2c0504c6479117112aa5c5beb419f0e620c9

Observation 7c01b7eb-38dc-4cb7-a07d-ee093e2e55d1 · outbound

This paper cites Accelerated Policy Learning with Parallel Differentiable Simulation.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Accelerated Policy Learning with Parallel Differentiable Simulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.710302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.710302Z digest=sha256:1203c98d773e6b84e224963b412fededadcf17eb073bbcee4917b7bc0b9ab0f4

Observation 35d460f6-6cb4-440d-b663-31922bd13f68 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.184123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.713293Z digest=sha256:6e4928378693228dd0183ad92908abb87a5b9678a58f9feaae13b91d55bb7145

Observation 914e794d-130e-4aae-8efb-3407f3df2b52 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.173553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.716317Z digest=sha256:51d47ff3e937e6fe7faedd7b805c0d557eb91cf95f184e6095e8b0639032a3e5

Observation 64f414f3-b693-4cc7-8005-2d0709843394 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.719001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.719001Z digest=sha256:8aa2e4aac702e75df3db99ecf6258e96cfe29be30b34ea19694a37f32cd92556

Observation f144207a-ff4e-4f5d-99db-4c575ce70b95 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.721707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.721707Z digest=sha256:bf85ce918be1b63a7b4e56dc991d41c21235b0d60ddc4892c3469b16483def2f

Observation 92b81cd6-9db5-4b6b-acce-458de32401a5 · outbound

This paper cites Raffin, A.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Raffin, A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.157346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.724636Z digest=sha256:29f0415ff8475f66580152efa357954e79636f1174af9ee0c0c4fe6701582d8d

Pith citing papers

Observation 9b8e006d-5641-4025-aac8-5e7acac6e7c9 · inbound

Simple but Stable, Fast and Safe: Achieve End-to-end Control by High-Fidelity Differentiable Simulation cites this paper.

Simple but Stable, Fast and Safe: Achieve End-to-end Control by High-Fidelity Differentiable Simulation ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:59.094277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:27:25.150807Z digest=sha256:e446d77538a4d52871bcfb9cdade5e5534644ddd7f48356f6f1393285c745b7b