Pith. sign in

Paper Citation Record · LEDGER

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2607.26509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26509 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:20:00.443820Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation feaae68f-2f0e-4845-96c1-d60edd8a5752 · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:56.345315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:56.345315Z digest=sha256:dc672ebff5ecfecb59403a2a853107aed35ef4da33127858440b153a3492bbce

Observation 7179c424-98b8-45a6-aa07-ee39bbc4cd2f · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:56.454068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:56.454068Z digest=sha256:7952ab8377bbb6e4934e7ae5b32ee7f4aee775c3f9ab859e44aa82882b74c4d0

Observation 2ff5af11-d22d-4f6a-b9d0-76adfb126434 · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:56.572836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:56.572836Z digest=sha256:b57048e2ee78ba31df154c8144316b6364718cc3649c41f57fc95d0498821ef3

Observation 6cfec29c-dee3-4690-86f4-6668e591827f · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:56.671393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:56.671393Z digest=sha256:0681ac41fb0042a219bcad475d12029c5afe20f95f95ebdb641c8339fae19737

Observation 8733a819-08ae-40a2-af3a-c21ef6abf65b · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:56.775235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:56.775235Z digest=sha256:c9dd520f718be70ce8f7353d05eb1ec61acf60acd180fc74827acace6cbb91f2

Observation 2f68c1fe-8df9-48de-8e5f-34156681f7b2 · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:56.879256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:56.879256Z digest=sha256:af5446d79cba3da60dd7df3ba8dcf03249d2389abc700f9f202eef67410b132e

Observation f09eb1e3-209c-4600-a507-4c430f2ef10d · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:56.981625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:56.981625Z digest=sha256:fe1ff2480bbf1425a6dccaf49eab124c15bd79bd6cf1d0a10abc287afef9bdba

Observation a0d53f79-8bc4-49ad-b68b-0cac5cce3db6 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:57.131516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:57.131516Z digest=sha256:91f2b9f68aabdd8fa41dc28fe9ef7873f5028d4b0adcfa326755471d818b2b3c

Observation fdcd6f40-9586-4647-930c-f773dab28eab · outbound

This paper cites Fujimoto, H.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Fujimoto, H

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:57.256589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:57.256589Z digest=sha256:a0a86589b4bcf93256165e7bd3b2b3d02da53d3d1786de19b24e32d729d4bc52

Observation 1aad6f97-5987-4dd1-864a-64798499ffe0 · outbound

This paper cites Tesauro, et al., Temporal difference learning and td-gammon, Communications of the ACM 38 (3) (1995) 58–68.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Tesauro, et al., Temporal difference learning and td-gammon, Communications of the ACM 38 (3) (1995) 58–68

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:57.342573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:57.342573Z digest=sha256:f2540b9ee30b3a2e54f6c749073b1e34d465fbeb0958c305cf4382dc5ffd2586

Observation 317eb993-e171-4c03-b25c-b77444ef3d48 · outbound

This paper cites Cetin, O.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Cetin, O

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:57.454575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:57.454575Z digest=sha256:a3a8151223a1b816b5417aa2923d4b6dc8f4581c55494ff28bfda73327397313

Observation 12875ec7-173c-46bc-83d5-b9ca0c3ca470 · outbound

This paper cites Nauman, M.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Nauman, M

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:57.554927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:57.554927Z digest=sha256:d42d189f49e310d8058bbd3caea8020922a0b28feaf44df003aedbcf32338a6e

Observation 84f51ae3-3d94-498e-8d57-07120340cc4c · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:57.722916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:57.722916Z digest=sha256:3a7fe45d6abeff2c6bbf4c2d5437e3ece7a9fdd022accf71965944820ffa7583

Observation 8ac3efae-3aa3-47d8-a294-b9eb8525c085 · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:57.870507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:57.870507Z digest=sha256:dd4c43be29ef74b6e26f6300e7b7c0cb5dbfdb18fac7335c9fa317964556c5b2

Observation 58169d2d-22b0-4e6b-90f3-7adc5deefa4c · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:57.980728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:57.980728Z digest=sha256:ff8620475657b5231bbd9f6956097188b22246e84844fb745c01585facc13957

Observation 12d79805-ed10-4f11-b540-7590531e0c31 · outbound

This paper cites Prioritized Experience Replay.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Prioritized Experience Replay

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.045890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.045890Z digest=sha256:3fa0bee1709cb232ab21875b9c35fb9043c5584b9a877910f847e5a86648b214

Observation 27520478-4065-40d0-b7d6-6f8cc0ddd8b0 · outbound

This paper cites Fujimoto, D.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Fujimoto, D

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.123100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.123100Z digest=sha256:137787ffc887bf95ecf39a0b888c17a3cb9faaed12ad41e641fc0b49962c636c

Observation f0d1f00b-6248-41b8-ac3b-0464fc02fc5c · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.198535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.198535Z digest=sha256:17e3f334ea5c2bf65582f0ef01247142255932e0e848a18563d294b8692903ab

Observation 24c3c813-1cb8-4079-8cb7-015bf8f45da7 · outbound

This paper cites Zhang, Z.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Zhang, Z

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.271877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.271877Z digest=sha256:dbad4f603ae152a7db113b55da335ec6587daf8b53464e00f5d48e21f6e19bac

Observation a074f1fd-fa65-4776-a518-e65e3bfced2f · outbound

This paper cites Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.343808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.343808Z digest=sha256:c09e1be423feeb873d55f3db4e477cc5ef12a0040cb3277f4f44477f0815ea1e

Observation ad3503b5-d135-475c-9d01-0b47fa6d4d41 · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.416980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.416980Z digest=sha256:c199107842c1fe3ea4259ba402a83c267784b2ef8f43acb403c5349cbd2dda1b

Observation d12f6818-06b8-496e-88cd-dba4daaa0aa0 · outbound

This paper cites Van Hasselt, A.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Van Hasselt, A

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.496770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.496770Z digest=sha256:093768205b87c76c4a30c33682a6ee544eeef52496c6d9c87aa4c35f603f04c3

Observation 5837f5c2-4c72-4f00-98e0-e0f509e345b8 · outbound

This paper cites Maxmin Q-learning: Controlling the Estimation Bias of Q-learning.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Maxmin Q-learning: Controlling the Estimation Bias of Q-learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.566038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.566038Z digest=sha256:12fcf1002a51f58726046c9063153ba96460ce62925d4879cdbc86a8bd0669f6

Observation f8defb5d-9ab7-4afa-b96a-58fe376965bf · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.665041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.665041Z digest=sha256:b7344d73f00e5ee67a2011ec5eacd118e5f5eb25331d96214233c55d873e7e41

Observation f4105ee3-6099-4726-b9be-fa4e0aae6c60 · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.740250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.740250Z digest=sha256:84320473fc76818180f6cc8db0ef0f18a1c8914cc3d973ce009957e5bb53a123

Observation 54f65f79-d978-4237-a220-46c624e43566 · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.801555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.801555Z digest=sha256:80073a2508608584e071dd3aeaac315ad90f3ca54d8df9bf9abdc119af59c5cc

Observation 4fcd855e-68aa-4d55-ab2e-2a862282c663 · outbound

This paper cites Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.893464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.893464Z digest=sha256:3f7754bd649fa53ae3a131237f276a2a80dd644f5f2c2757c91b30c248b33f15

Observation 87ba3c41-7aa8-4e41-bea0-6b54f0861db0 · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.988156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.988156Z digest=sha256:003336dc9b7abd3bab92247f1d1cd5806e81eb08e6778015d57050b654d9d45c

Observation 576af92a-a6a8-4fb5-918a-0a064d829d9c · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:59.077231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:59.077231Z digest=sha256:c5d0558ca1be68f5d2dcfc1d2104e96ba96baf288466cd66d587cd97cedf869b

Observation 303dc241-5205-4c4b-b91a-b14777f59d8a · outbound

This paper cites Moskovitz, J.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Moskovitz, J

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:59.169694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:59.169694Z digest=sha256:f0db0efd032feb6be386215a701925dd1d303d72336f80420099aa39d27b09cd

Observation c27d9a3b-3085-47ad-a5f4-a2a852990263 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:59.358820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:59.358820Z digest=sha256:1df29fb30dba8258864de6f025d51ab06e68254d9edb42ec65732a01a9e2fecd

Observation 89fc52d4-b86b-4934-b810-98dccf4fd844 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:59.537141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:59.537141Z digest=sha256:25da11a0f6c371c1e31e5c748d92d5d9a384f8677d9582ca45827768d88a1483

Observation 9d814c5d-4041-404d-8e5d-b0935e44b517 · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:59.674034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:59.674034Z digest=sha256:281a28b9d6ee53fb110b10cd33a0e2aa76fc5661d103f672f68cca9fa2e2a626

Observation 554a53ac-b034-47a1-9dc8-62d3eb95f108 · outbound

This paper cites Zhang, W.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Zhang, W

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:59.825409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:59.825409Z digest=sha256:0bddff70ce1f7280acfc686db6db66fb5cf780454e9e09fe7c6dba043aeadfc1

Observation 7a705d2d-e3e4-42e3-aacb-279a33792de1 · outbound

This paper cites FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:59.991253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:59.991253Z digest=sha256:7b06646b25d6be3264087343da6cf6f66aca31b23f58bc2513838e6090c17f31

Observation cae1e53f-0f96-4287-895f-b40f327a4ffc · outbound

This paper cites Hassani, S.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Hassani, S

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T14:20:00.129349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:20:00.129349Z digest=sha256:a49991615e6384a9ad8833982eac2320f2be406cc9c1dd2638865ce4c57b3876

Observation e55c1a13-c8bc-4282-96c1-3ad281306983 · outbound

This paper cites an unresolved cited work.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T14:20:00.219119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:20:00.219119Z digest=sha256:ae0a4912b7037789352edbf1ba156bc0c8058ae175f80afb53cbbda9edad12cd

Observation 03fe2268-d5bb-4a6c-b23d-dcfc61694285 · outbound

This paper cites OpenAI Gym.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning OpenAI Gym

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T14:20:00.288012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:20:00.288012Z digest=sha256:af07d109d3ae4d301dce671d9cbbfdd40d1e3f231296aa155bd53a118f8d62a6

Observation 26adee7a-c462-4769-8ad7-2596ef62d964 · outbound

This paper cites Coumans, Y.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Coumans, Y

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T14:20:00.353496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:20:00.353496Z digest=sha256:b4fe2db98f5c4c7133477f98ebc5c4b82c30514ad6cea4d1a4a2eeeefa06689e

Observation 1f059672-91dd-436d-9679-ae8180939ce2 · outbound

This paper cites DeepMind Control Suite.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning DeepMind Control Suite

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T14:20:00.443820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:20:00.443820Z digest=sha256:688540a57a7627ed039d204df43676b1a6eb953a8f42a1a09c2e11b08188836c

Pith citing papers

No inbound Pith citation observations are available.