Pith. sign in

Paper Citation Record · LEDGER

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards

As of 11 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2501.14513.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14513 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:10:48.724636Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:27:25.150807Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T08:50:59.092753Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2a6c13df-86bb-4ab5-844b-fa927d35105c · outbound

This paper cites Loquercio, E.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Loquercio, E

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.353788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.602656Z digest=sha256:9fd28e31151965e9c29fb943864b5e44283b8e13cd329e33fa6b537134e9a58b

Observation f39976c4-480f-430e-901b-7313d4e1f02f · outbound

This paper cites Loquercio, E.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Loquercio, E

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.342872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.606745Z digest=sha256:5c1b638b34d0cdbe6a2096dab880433e6df5949eb4a560c6e128cc42b96f7316

Observation 45562607-49e2-4e8c-84af-fb7895d5b67d · outbound

This paper cites Kaufmann, A.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Kaufmann, A

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.333036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.609915Z digest=sha256:a37d4963099cd85c6f63e21648dbe21fafd660b8d3e7e2d52720586168042405

Observation 76dcdfca-35c2-4c1d-9b97-52e3174b8cc8 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.321381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.613006Z digest=sha256:be0bfb84fb744581ae2ceabbdb05a3cc7f839fef5e8a33bd948f53fe0c5e08be

Observation 347fb54a-7ece-4caa-bd65-994b173fdad7 · outbound

This paper cites Back to Newton's Laws: Learning Vision-based Agile Flight via Differentiable Physics.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Back to Newton's Laws: Learning Vision-based Agile Flight via Differentiable Physics

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.616444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.616444Z digest=sha256:8b66be72c38e6736d03fa27daf6f7742dd463d7a4a5d2db9536b4e6182c862af

Observation d4b199b2-798a-4c38-a549-84db8edd5f78 · outbound

This paper cites Wiedemann, V.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Wiedemann, V

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-10T15:10:48.620025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.620025Z digest=sha256:148e6dd193448a3411ab0754f5467fe7594ec2dde8d6d686e31d617531d3e5ba

Observation cff916ca-fe64-46f4-8017-e934760916cf · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.311453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.623414Z digest=sha256:9a1539154b18a10d9375251ea8795a9da0871c3a0b7e2a35e3d372a9f571d56b

Observation 8f7c4c75-7273-4fe9-a73d-23ee9732b71f · outbound

This paper cites Learning Quadruped Locomotion Using Differentiable Simulation.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Learning Quadruped Locomotion Using Differentiable Simulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.626404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.626404Z digest=sha256:f35a17ac9ac30d44b83393cc7eb5c4a319a5fd05ab090e69205daf4cc7029b4a

Observation a3ed2d12-91a8-4c4a-b677-4296a8ba90f9 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.629534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.629534Z digest=sha256:af495fba1740625808c4b9a93a928d0c6285f037b5b6323887f806e90300841a

Observation 7add1830-6e13-4a1a-83ea-629e8f04fa1c · outbound

This paper cites Zhang, W.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Zhang, W

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.301955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.632272Z digest=sha256:03b27105c1005650dddad86ee0ecce909098b214ef5463e73056d3dd2bfc45d6

Observation 400c41c8-6704-48e0-a6bd-249558b1be33 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.293561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.635581Z digest=sha256:5d1885e7f60901e0c1f14798daf2a0a8edde4a5eb5eab54a9e33899ff77e1b3c

Observation fd32f32e-a259-4cee-a19a-26e41e0e3e02 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Playing Atari with Deep Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.638379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.638379Z digest=sha256:84b94f3da35d9a4e1b9b989a23597aa1561da808aeb08390b104ce05ca07c12f

Observation 7541952c-6c51-45bb-bd5c-422ab98394e1 · outbound

This paper cites Continuous control with deep reinforcement learning.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Continuous control with deep reinforcement learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.641907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.641907Z digest=sha256:e538e0c68001c000712c3c5de74870607bfe8ea8d1c28ed758d3f2bbe1fdcc37

Observation 17aeecff-0047-4d6d-8480-85f28b681b0a · outbound

This paper cites Fujimoto, H.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Fujimoto, H

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.284169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.645239Z digest=sha256:76b8263e92cacde5e52c23542bcef8f1044c8e0ebfa810fb4bb67f1dffbde5e6

Observation 4a649408-30d0-4204-af3c-9a67ed14f5ab · outbound

This paper cites Haarnoja, A.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Haarnoja, A

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.648213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.648213Z digest=sha256:56afaf147846f92ba5bf938ac01d25f06c5c0e7e632b848655181dbaf4c1d6b2

Observation 7a4a7514-f73b-4c86-aac7-ed59c30b87bb · outbound

This paper cites Trust Region Policy Optimization.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Trust Region Policy Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.651760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.651760Z digest=sha256:8211d48ae578ce655291f251a09fcdeb17343e36723b173f206a0077fdd7dd21

Observation 19e6577e-580f-4d35-b307-cbe404a74168 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Proximal Policy Optimization Algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.655628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.655628Z digest=sha256:cfaf0279bc6a3dd5a25a3f529690abfe715fca1370fca09a4f6a297863d3e94b

Observation 3282080b-e644-4e22-ae24-a23756b66830 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.268170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.659274Z digest=sha256:ab535a0f17f1c8749c157e17c438c7eeea01881c2bc2b04ef1c1d9419ad10dad

Observation 33b0d98d-fbac-4889-8b3d-984f6a4777a7 · outbound

This paper cites Deisenroth and C.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Deisenroth and C

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.663178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.663178Z digest=sha256:43e6b4c50cbeb71ed1be708dc53380433b2d8ad086cf656989a856d6e72e78e2

Observation 89d80766-acd2-4fc4-ab74-39af80008f1c · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.251386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.666518Z digest=sha256:bd3251208489c3d5f297edfa921214459193c89389476bf6cacc3a125e896dd2

Observation f128fc34-672c-413b-9455-c0091379df41 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.669930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.669930Z digest=sha256:c0e8ac2ffd718ba035f28415b3d41c90050db59e79d0505570a77b669b39a4ec

Observation f80059c9-6d7c-4d3c-90b7-61df54726a2c · outbound

This paper cites Watter, J.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Watter, J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.673338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.673338Z digest=sha256:e7d1ac7d5068a492067c9759b4b18a71762e4a6e0e2eb74e853588eeaf296a33

Observation 44269d92-9f4d-49a8-a6a7-8fa60d2121d6 · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Dream to Control: Learning Behaviors by Latent Imagination

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.676127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.676127Z digest=sha256:2ba530c4444d0fdf64cad10465cce86528bc19bc13454dd7fa4757c1b99dbe9e

Observation 16ba7ee5-3adb-4f6a-b810-88682f641455 · outbound

This paper cites DiffTaichi: Differentiable Programming for Physical Simulation.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards DiffTaichi: Differentiable Programming for Physical Simulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.679345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.679345Z digest=sha256:68bbe89dac8921e59bd5c9d19c2be72280a0f4850d75f1e9bf647a60eb0f3b19

Observation b24d34d9-1fe3-423f-88c8-e060a774792c · outbound

This paper cites Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.682411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.682411Z digest=sha256:da690ef7aa492cfa3d96cbf51a0ce36a5de465dbeb6f4b0d21b1311ea76bd80c

Observation b65633a5-3bf1-4189-95f9-b7995508fb8d · outbound

This paper cites Todorov, T.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Todorov, T

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.685299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.685299Z digest=sha256:e2583b6a43f882bb239fea7f8943db2e24845c908ca06e98884e6d3379581180

Observation d9036fca-56ba-4228-bf6a-1147d597a461 · outbound

This paper cites Heiden, D.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Heiden, D

Reference 27

Resolution
metadata mismatch
raw_fallback, observed 2026-08-10T15:10:48.855662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.688100Z digest=sha256:779686c13346c3ba3c3033b34bfded129abb8e180662d9a936a4d47bd2fed91c

Observation fe272e94-86ba-45a1-8198-9a65ea4d0140 · outbound

This paper cites Dojo: A Differentiable Physics Engine for Robotics.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Dojo: A Differentiable Physics Engine for Robotics

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.690868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.690868Z digest=sha256:2cb85e283957e67a0055aec9db90a9ed9f6f5a726c27535ec4a66a6691708004

Observation b5ac3004-5f7c-4700-8691-2ef5ae39098e · outbound

This paper cites VisFly: An Efficient and Versatile Simulator for Training Vision-based Flight.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards VisFly: An Efficient and Versatile Simulator for Training Vision-based Flight

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.694231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.694231Z digest=sha256:4e4b9ce7219e35345b5e7d07ae288ee4a15130039336fd43324fd763bb4fd88f

Observation c9b09ff2-12dc-452b-874c-933a5043cb92 · outbound

This paper cites Savva, A.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Savva, A

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.230830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.697111Z digest=sha256:f02b30370b9d4cc5a1e996207bc31ce72fb7d5c5e9a27607a0ab1432ab6e537c

Observation bd1987eb-84d1-454c-915c-ba001b222dad · outbound

This paper cites Schoenholz and E.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Schoenholz and E

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.220966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.699836Z digest=sha256:50d25531b4ed7d46a16cc465f6d791a5323a1195fb99c9567b22a7815ac9ecac

Observation 55011a73-a570-4ff0-a2b4-089610a5fbee · outbound

This paper cites Paszke, S.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Paszke, S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.702421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.702421Z digest=sha256:1d471d11210ad8141789e2895cf53425fa9170983acf40a116be61e5b6873e91

Observation 7239ef10-55ae-4a66-8fbe-5d8e30adfe93 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.204429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.705121Z digest=sha256:3fa284e444fab4d1dd0194dedc3fdf3f90d96e406b57564873fe23e1c825b3c4

Observation f66c94d8-eb33-4d30-81ea-cca8aa979c40 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.194264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.707788Z digest=sha256:8e076028a7fa884dd4c622e76c0cde5dc8f222da502ce4e43f32f78042159b88

Observation 7c01b7eb-38dc-4cb7-a07d-ee093e2e55d1 · outbound

This paper cites Accelerated Policy Learning with Parallel Differentiable Simulation.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Accelerated Policy Learning with Parallel Differentiable Simulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.710302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.710302Z digest=sha256:2352f2a3eb8a334858fce35a3fdd4069e018ebb8cc0d359b32ae67fcaaeadc71

Observation 35d460f6-6cb4-440d-b663-31922bd13f68 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.184123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.713293Z digest=sha256:541250cc090cd3b5b61961839f7d5d9c63227d7f2d530c410e8e0a1b936d344f

Observation 914e794d-130e-4aae-8efb-3407f3df2b52 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:10:49.173553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.716317Z digest=sha256:871809149720cee6f939ae5e853f4d9fb8cff88c0e1c6101ba96a2fb9b71d797

Observation 64f414f3-b693-4cc7-8005-2d0709843394 · outbound

This paper cites an unresolved cited work.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.719001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.719001Z digest=sha256:28a9252ebbbe2b73d2bfff36a6ff6a22f2116048848facf7ca74ef3a00fdee18

Observation f144207a-ff4e-4f5d-99db-4c575ce70b95 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:48.721707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:48.721707Z digest=sha256:1fbe61f0444211dfa48eda9d003a59c78243cf958f837658e40c6e2781266d2e

Observation 92b81cd6-9db5-4b6b-acce-458de32401a5 · outbound

This paper cites Raffin, A.

ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards Raffin, A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:49.157346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T15:10:48.724636Z digest=sha256:f0a635210a4f6838f0c9abdc4589e6e48a72e043681d3b34d1fd10a58c48eacc

Pith citing papers

Observation 9b8e006d-5641-4025-aac8-5e7acac6e7c9 · inbound

Simple but Stable, Fast and Safe: Achieve End-to-end Control by High-Fidelity Differentiable Simulation cites this paper.

Simple but Stable, Fast and Safe: Achieve End-to-end Control by High-Fidelity Differentiable Simulation ABPT: Amended Backpropagation through Time with Partially Differentiable Rewards

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:59.094277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:27:25.150807Z digest=sha256:caa1edd232650a21fa0ec30cae75e2bc6140cde05fe180d23637b6fffbbe4f56