Pith. sign in

Paper Citation Record · LEDGER

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

As of 8 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 9 inbound Pith citation observations for arXiv:2505.23150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23150 v1

Coverage vector

measured 100 of 118 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:58:54.542081Z

measured 109 of 109 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:39:05.516882Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T23:09:12.737900Z

Reference resolution

100 of 118 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved76
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c403830a-7572-46f5-8a4c-4e8f99131eac · outbound

This paper cites GPT-4 Technical Report.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:44.854747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:44.854747Z digest=sha256:d11e1d54ba8838b8176161043a62588a99d5f5ec6901fe8d495dcd72322ea08d

Observation 80c9da86-63a7-4a64-b95f-3a0d48e2c8c8 · outbound

This paper cites Provable benefits of representational transfer in reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Provable benefits of representational transfer in reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:44.967526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:44.967526Z digest=sha256:844ab77512346c28cc31c86e4439ff81d35a615823a6f88809f68dfab6c82fc0

Observation 713717d6-e6bd-4c9d-96bd-c26975730a9d · outbound

This paper cites S., Courville, A., and Bellemare, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Courville, A., and Bellemare, M

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.019350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.019350Z digest=sha256:ee21dbbb9005dc28c300e77e280ee6eb58e0cb776ca6840c14e8ef4aea1d0cad

Observation 97a075bd-184a-4409-aee0-3305870ecba4 · outbound

This paper cites S., Courville, A.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Courville, A

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.121469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.121469Z digest=sha256:0adce891dc4f301c101e36f67fa8725f1904fff1b088fdfe4e7ee55c56c7e885

Observation 7da3fa4c-e85c-4342-99ec-4195f8b9d31f · outbound

This paper cites Hindsight experience replay.Advances in neural information processing systems, 30, 2017.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Hindsight experience replay.Advances in neural information processing systems, 30, 2017

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.197775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.197775Z digest=sha256:e9e2466dcf305dc5d3f77d24f5c6ccf9c1ff538f084dc4276c91475e141ab8e6

Observation bf984bb8-9e57-4b9d-b244-bc6220398bcc · outbound

This paper cites M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.269538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.269538Z digest=sha256:8e4f13d208a5bc01747b29b8174809248a3cd3adc8b12be28313367990425b78

Observation 74c0d758-659e-43e9-ab11-00f55115d6d5 · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Video pretraining (vpt): Learning to act by watching unlabeled online videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.345247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.345247Z digest=sha256:32cfa87268bde477f81bc64e2a23be9389374baffe8acc8c9a171213df51d243

Observation 98169282-9aa4-4dae-8152-7ca68d2c59fc · outbound

This paper cites J., Smith, L., Kostrikov, I., and Levine, S.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners J., Smith, L., Kostrikov, I., and Levine, S

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.433639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.433639Z digest=sha256:3bf64fffbdbde95edf44f7c81bb74a6a473231c9cc9e4bed40fc09ddeadfaad0

Observation 3b1b6998-d6db-4859-ba31-625e5dc47236 · outbound

This paper cites J., Schaul, T., van Hasselt, H.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners J., Schaul, T., van Hasselt, H

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.571178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.571178Z digest=sha256:13f72394fd6497f5e2b571543106d46f53d5c478b661065c432bb1ce09a900f3

Observation d9949fcb-109c-4545-853c-02164a0e7727 · outbound

This paper cites G., Naddaf, Y ., Veness, J., and Bowling, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners G., Naddaf, Y ., Veness, J., and Bowling, M

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.663223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.663223Z digest=sha256:706e129989cabfd9ebddaf8e831935f83feb6971366e099c622641f95ac601e2

Observation c5f64815-4508-4e5a-ac03-e3cbf631d284 · outbound

This paper cites G., Dabney, W., and Munos, R.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners G., Dabney, W., and Munos, R

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.750044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.750044Z digest=sha256:75d4853da25e38522221c84b48d1245376cb36e89272bdba960102e4bf31fb46

Observation 08d1053d-46cb-4f71-880d-78bc3dd62be7 · outbound

This paper cites Dynamic Programming.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Dynamic Programming

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.804154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.804154Z digest=sha256:275f05c1175368a575288f7a8152b011dbc07651ebea211c43c55f9a4d312b9f

Observation 2486c5fb-c6b4-443c-8cdf-fd42dd633eb4 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.874874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.874874Z digest=sha256:c24ae50f12e6162e9a795f6c1aadd5737732ce5c6aae67130c5fe9feacf9f936

Observation 93b4fe42-9f73-4d76-82fa-1430bf502f95 · outbound

This paper cites P., and Weinberger, K.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners P., and Weinberger, K

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:45.968173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:45.968173Z digest=sha256:95f8974b51bb93f14528189a0bd4702ca05db41bfd3d09ac9f00db2dd7e8b609

Observation 68d41d36-8aae-423b-996d-11f1b1fdfef7 · outbound

This paper cites J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., et al

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.065866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.065866Z digest=sha256:96163fb45936e33a760bf72516b59a92facbbbe9972bb33093d1ae27cca015c3

Observation 63c3a67c-9aca-4c18-9056-0c246c26dc54 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners RT-1: Robotics Transformer for Real-World Control at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.158225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.158225Z digest=sha256:870fdf34e8c649c9ac24e07338d5c8ab6e3093b69bf2f7dc0c9322d69967f80c

Observation 778f8cbc-d862-4bf3-b45f-d40d66689849 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.274336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.274336Z digest=sha256:32353f7d612514f4d12b318f8887ed76421b60fdc70520535e670586b24cbc0d

Observation c8a57e69-4a22-4384-9865-e7bbadf3fcad · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Decision transformer: Reinforcement learning via sequence modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.379002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.379002Z digest=sha256:d89951a47a0ff4867c745f248476214b800e535079f0a6a93604ae986b6f3744

Observation 8bfdf1d4-ebbf-4fa2-969a-3b77060d52cf · outbound

This paper cites A system for general in-hand object re-orientation.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A system for general in-hand object re-orientation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.445275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.445275Z digest=sha256:47477d83a48380070d8cbba571cd649c31bb8ff1a8ccea858249034511951fe6

Observation 5c71d3a6-7cbf-4467-85f6-f06c0eb2f829 · outbound

This paper cites Gradnorm: Gradient normaliza- tion for adaptive loss balancing in deep multitask networks.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Gradnorm: Gradient normaliza- tion for adaptive loss balancing in deep multitask networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.566082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.566082Z digest=sha256:2548f1ab69c534da9966c6af5fddcd514d637dd4a16b87b9a7ca21908754d1b5

Observation b133e83c-fc66-447e-8f32-286d8c5f2621 · outbound

This paper cites Just pick a sign: Optimizing deep multitask models with gradient sign dropout.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Just pick a sign: Optimizing deep multitask models with gradient sign dropout

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.662012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.662012Z digest=sha256:22b24b749ae007a0a2f525ee2f5e08036efca4903e2178840554d14f26d2c19b

Observation 725d0859-c737-4840-9e5f-ba3f83d84306 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.707628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.707628Z digest=sha256:b676f752e3f307e8079ff0314f53a24b57d5ac4a2da87b5212c1bc0f2516956c

Observation e5752454-cbfc-4efa-9bae-76e8fa6c8afe · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.750339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.750339Z digest=sha256:a0cbfcb3136bb4fd9043c5d749250e8f4355ce84be0c499d04b0f5d9fc846a68

Observation f173b2ca-f1b5-45e1-a98b-0ac0ff439627 · outbound

This paper cites Better exploration with optimistic actor critic.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Better exploration with optimistic actor critic

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.835161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.835161Z digest=sha256:ed353d549dd61a6ace0b4205c6ccc315ac1ff66b800f79f3b445649ef53ea161

Observation a9579dbe-bdc1-4982-8b39-ed18dca7ba0f · outbound

This paper cites Magnetic control of tokamak plasmas through deep reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Magnetic control of tokamak plasmas through deep reinforcement learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.867259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.867259Z digest=sha256:f4d557609f2769cca4af4f1fd73f3e11b05c0df25fe17c761ab6e6a3ecb533ab

Observation 544c6b33-a6c6-45fd-ac93-60cd1d1f1f1d · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:46.915785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:46.915785Z digest=sha256:c1fdb6539c24357bb75d47ab8793e0037d235a0d7ec4de0d43eca55763e40bce

Observation cedc50c4-c18d-4625-9599-87af9b1a4611 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners An image is worth 16x16 words: Transformers for image recognition at scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.009204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.009204Z digest=sha256:cf33d49923ad444c844180f5baaa4e0b21c86a6ac8e4e19b265bdc1812557b73

Observation e88b3017-0087-4ffc-864a-353bd7bbd744 · outbound

This paper cites S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.060172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.060172Z digest=sha256:7196a0af6f756b6079aa87a39bc58f536753e8e566bcd3efe6248e833acd03b4

Observation 8156fac5-839e-4306-9452-eb846bba34f8 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.128470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.128470Z digest=sha256:283374c62496c51e24460553521b9bba4aa1ff67c73088b07ee782d22b34b6f9

Observation a990062a-6dba-4f00-b38e-c2f4518b63d4 · outbound

This paper cites A., Chebotar, Y ., Xiao, T., Irpan, A., Levine, S., Castro, P.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A., Chebotar, Y ., Xiao, T., Irpan, A., Levine, S., Castro, P

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.227015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.227015Z digest=sha256:ae6a18237d625294d5929f282122cd7aee5f7b4d542679837a697610f4710f11

Observation ba3cede9-4995-41b8-a536-f8b58699f1f3 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Model-agnostic meta-learning for fast adaptation of deep networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.334044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.334044Z digest=sha256:6ee1786f97b428d96d324cb1f6cc609280379b0b547f59e4a196548323815542

Observation 5e78924f-1e27-4136-9ecb-d46590b3ef89 · outbound

This paper cites and Gu, S.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Gu, S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.398149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.398149Z digest=sha256:3b261a7117a97fb3cd7a38685951df59316d6ee967399925f689d35d5e4a4bbb

Observation b3f8af50-5032-4c35-8ef3-b0e167b5c22f · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Addressing function approximation error in actor-critic methods

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.450522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.450522Z digest=sha256:e2c142a6d7cc57fbb015aae29989bf98722aab61ff74efb5058750e7c2bd5a07

Observation ad6cab91-fa5f-41a7-bd2b-995743cf13f6 · outbound

This paper cites S., Precup, D., and Meger, D.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners S., Precup, D., and Meger, D

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.514281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.514281Z digest=sha256:bca657aab86f4766c21c366046296d75bac2106566c98c03daa624407cadbe3f

Observation 3f5c7ad6-7ec2-46ff-b60d-2c762b03bdd6 · outbound

This paper cites Divide-and-conquer reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Divide-and-conquer reinforcement learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.599219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.599219Z digest=sha256:f73bb6fc30837a7ded9504e833a48bac074457809ed661a602c011c4c4226513

Observation 11695065-eb2b-43bf-8acd-b1b64740f6bb · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.702018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.702018Z digest=sha256:a6df5219998596d63aadacadbab13c7c1a5db9bbff33fb17579c2fed86c183ef

Observation 9b11e6db-448c-4902-ad8b-699cd749dc7e · outbound

This paper cites Mastering Diverse Domains through World Models.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Mastering Diverse Domains through World Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.810738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.810738Z digest=sha256:642bfdd455207887f96d3005e53d7cccedb2b032f75a4fc680b7b5ee74b35dab

Observation dcd824f7-6396-4ee8-87f5-7728212e1b21 · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.926005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.926005Z digest=sha256:68fbcb478ce9a4647b3c7652bb5f29acb0bd6a99b8010e739c5756902d4e6333

Observation ee899869-f4ac-402a-bf1a-243769f12730 · outbound

This paper cites R., Millman, K.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners R., Millman, K

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:47.989117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:47.989117Z digest=sha256:02b1592d14181b8f5a1e40716838120132cfcc6c33578da22a89776522be526b

Observation 83679e9d-c067-4d62-96aa-c79d0a711e38 · outbound

This paper cites T., Wang, Z., Heess, N., and Riedmiller, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners T., Wang, Z., Heess, N., and Riedmiller, M

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.058887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.058887Z digest=sha256:bedeac297ccb4275f1ffb0d10478dfda4dd2cee2a6d4cb46f45ceaa442816d86

Observation 0e9bb63e-08e0-494c-bd15-61e1e3d3a57a · outbound

This paper cites Efficient multi-task reinforcement learning with cross-task policy guidance.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Efficient multi-task reinforcement learning with cross-task policy guidance

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.115400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.115400Z digest=sha256:c8e1bf4f31266038443dd2c4b846d15e5fa8f8d6f1fa004104d10b3ef28b114a

Observation b8e672c4-b654-4d62-9126-c99e3324d0bc · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Scaling Laws for Autoregressive Generative Modeling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.198228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.198228Z digest=sha256:baf019d253fd7ee55b576617dc4a2248fc9ee289029e6ee7fb1aa61272a78570

Observation aa26013b-9414-44dd-8134-fb1b1af1cdd2 · outbound

This paper cites Multi-task deep reinforcement learning with popart.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Multi-task deep reinforcement learning with popart

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.306053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.306053Z digest=sha256:816b6754c7f6f5bbd25cc20a710585ea15dd8c81a9089720a3c2783dd9b2669e

Observation e10abf86-0652-4874-9896-dc073ebdf597 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Training Compute-Optimal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.423350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.423350Z digest=sha256:91ffa178ec3375796946dc193eab04d48617c4112f32c944e729f800d316d2ac

Observation da2ff812-3cf0-4a19-a9f0-32166cf81636 · outbound

This paper cites Otter: A vision-language-action model with text-aware visual feature extraction.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Otter: A vision-language-action model with text-aware visual feature extraction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.535699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.535699Z digest=sha256:326884c2c17938543ba3f487237608f79adf9adf97569e93198bedc4e62e29e2

Observation 1d70b86f-53ea-46b7-86c0-9f050d402e4c · outbound

This paper cites Generalization in Dexterous Manipulation via Geometry-Aware Multi-Task Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Generalization in Dexterous Manipulation via Geometry-Aware Multi-Task Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.664444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.664444Z digest=sha256:101497b051a7fca340085e0e4025e5fe4f40bd077b4a7f905c34f1d100bcfde0

Observation 433abc24-6485-4cf0-9535-847bfb6ac080 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:48.794678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:48.794678Z digest=sha256:0b3d6e6822de9d03c5f2045b70f6e38870d464ee2ed44c8a44552413b3d964b7

Observation 1c78482a-fe9b-43f3-9780-a093bedcb51a · outbound

This paper cites H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:06.510973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:48.918826Z digest=sha256:643ead6ef99f8ddde419b0f03ab014220177eb56b82c5ff530c6b3ce0f988d3b

Observation a618c6a6-b9fc-42b8-81af-bfefb5440eaa · outbound

This paper cites Scaling Laws for Neural Language Models.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Scaling Laws for Neural Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:49.042717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:49.042717Z digest=sha256:9b07b6e49c8456fc1541222d84c848814c088c86a56e87d1c1bfe6e3a0dbefe4

Observation 5a2632db-b1fc-431f-b978-3e0be2c69394 · outbound

This paper cites Champion-level drone racing using deep reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Champion-level drone racing using deep reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:06.350502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:49.165271Z digest=sha256:3df96854bf09b96159462c5b221f122e5ebd844863ca011a260b9594863518db

Observation 9dd9a88f-9958-4e08-a061-06a00151d3e1 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners OpenVLA: An Open-Source Vision-Language-Action Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:49.306614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:49.306614Z digest=sha256:c7a960b102e8d0519fd2adccb5905e1f3c51d7f057047039e376b8d4f0af79da

Observation 82ba696b-a068-4fdc-942f-828e6bd8b89a · outbound

This paper cites and Ba, J.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Ba, J

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:06.176243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:49.440203Z digest=sha256:49bf0ecb37a6ebca3095f90a88ff515e21326a5435f4271877b69e48946f8b19

Observation 25391644-5245-47ae-a423-bd5dfe19e699 · outbound

This paper cites C., Lo, W.-Y ., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners C., Lo, W.-Y ., et al

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:06.005942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:49.554910Z digest=sha256:ff320e0c9c796623b8a86b1afa2d67d5e5a1b550a5f99c65e23417c578b05aeb

Observation dabe2b8f-9758-4483-88a5-a239cd795e92 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:05.816415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:49.704891Z digest=sha256:5622f09cbd08eaa73b8090afa88d87d69df86b35f569a8ebd082bd869ccd0d69

Observation 0b377d3d-6be3-4aa6-b19d-e7f2245a7783 · outbound

This paper cites Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:49.816226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:49.816226Z digest=sha256:58ac9a40c28a886766aa3ae10d05b6624d3318b65e1ba2da883a88ab93585d38

Observation e0fce54b-56f3-4e46-9d78-a750c2e7257d · outbound

This paper cites DR3: Value-Based Deep Reinforcement Learning Requires Explicit Regularization.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners DR3: Value-Based Deep Reinforcement Learning Requires Explicit Regularization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:49.941083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:49.941083Z digest=sha256:0bdb5b410a86d1c2fe1b597522b1c5a443f31fa05a81327c0be68837b1e75ed3

Observation b92fdb5d-b34e-42b9-9dd8-0b79474c4011 · outbound

This paper cites RMA: Rapid Motor Adaptation for Legged Robots.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners RMA: Rapid Motor Adaptation for Legged Robots

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.021389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.021389Z digest=sha256:adae5dc78ab91ec373359ae8266816947cdf11228072ff63de79fc541861d8b7

Observation b7b884ea-e9b5-4d08-8c9d-2d6a9dafe2a9 · outbound

This paper cites Offline q-learning on diverse multi-task data both scales and generalizes.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Offline q-learning on diverse multi-task data both scales and generalizes

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:05.630424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.084622Z digest=sha256:7c0105061771ca4337a224a0fb046cd1175a56c282d82b9fc0705bf2e8cf8e6e

Observation e53f8be0-5c13-4972-8dd5-fe5ce4b48e33 · outbound

This paper cites Reinforcement learning with augmented data.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Reinforcement learning with augmented data

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:05.433016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.148267Z digest=sha256:8efd2c2e1afe2b4d29b1abb0acc41b3ad73e797f9963d72dd9acad4030a6012f

Observation 6e598f6e-02ff-41e2-80f0-7de53dca1f64 · outbound

This paper cites SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.224254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.224254Z digest=sha256:1a28818d8d45bee0ee4400eaf7a61ee5383bcf076ce806d3725a2d021441f823

Observation 0ac955de-f366-46dd-950e-f7f7e1d1f34f · outbound

This paper cites Hyperspherical Normalization for Scalable Deep Reinforcement Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Hyperspherical Normalization for Scalable Deep Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.269251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.269251Z digest=sha256:13d543ceb82b3613467095bc430320a635b1c8cd238b281b29ec4e98c56a8b3e

Observation ae14b0df-74e1-4d93-bc6c-6e594c8ee1ac · outbound

This paper cites End-to-end training of deep visuomotor policies.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners End-to-end training of deep visuomotor policies

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:05.163798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.321875Z digest=sha256:2cdd120c8ae1b734b222d55a09a653097bafcf0641c20a1efaaae91cf3033c39

Observation ba682875-3697-4628-a572-3f871e74de31 · outbound

This paper cites DeepSeek-V3 Technical Report.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners DeepSeek-V3 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.379474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.379474Z digest=sha256:da4bd9d92a63055742b1ef74c16d180b4cf822a53596d89b9b2bf018573081f8

Observation 391030fa-4064-41f5-a09d-32a19b9a97b0 · outbound

This paper cites Conflict-averse gradient descent for multi-task learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Conflict-averse gradient descent for multi-task learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:05.001905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.436638Z digest=sha256:35e7fc570d25eb2415ec96d7f64134e15a962f6d65117e3a7e8d3f5a213d90f5

Observation ff021dd4-6f61-420c-ad80-4985741505ef · outbound

This paper cites Moka: Open-vocabulary robotic manipulation through mark-based visual prompting.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Moka: Open-vocabulary robotic manipulation through mark-based visual prompting

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:04.743852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.496763Z digest=sha256:ac610c9521c3bba18374d5f905d97d32b2496aebd97556dff12f6db0bda352f6

Observation 91825c01-af53-4a35-827c-03fa771c2fe6 · outbound

This paper cites Scaling laws for fine-grained mixture of experts.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Scaling laws for fine-grained mixture of experts

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:04.481489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.558081Z digest=sha256:f9f0e747f4fc129757e1133b5caeaacb8c2799318bfad24cc5c7f78b79c0513a

Observation a5446eea-762d-424c-bc91-e942572d7a78 · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.612205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.612205Z digest=sha256:779d99be879911a7b90e0142cb91f7003829a9aa6319babbd7fa7a22a536b9a6

Observation 2a72955a-105a-4fd6-824e-1a6f81559308 · outbound

This paper cites An Empirical Model of Large-Batch Training.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners An Empirical Model of Large-Batch Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:50.670391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:50.670391Z digest=sha256:80330b1e12898b27c0798a46aae17cd32a78458aa0699173640b2a9503c930bb

Observation 0add11d2-7c60-4e76-9f95-c193c5bd944f · outbound

This paper cites A., Veness, J., Bellemare, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A., Veness, J., Bellemare, M

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:59:04.257681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.729363Z digest=sha256:6fd03dbff69e6848073cf6de523e43193c2f7f5c258b291c2cd14f66295d029a

Observation b5a37fcc-0fb6-4ddd-b758-ab4199734896 · outbound

This paper cites P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.997563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.786249Z digest=sha256:0a8bde9c131663340fd6eb89c315f28e2ffc89aa77e3140bc5463f9d9f1cbe8c

Observation abafc6a0-52de-4d8d-a1fa-c3a640fc15c7 · outbound

This paper cites Tactical optimism and pessimism for deep reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Tactical optimism and pessimism for deep reinforcement learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.761137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.835032Z digest=sha256:d33d8404d164d869ae5bec276b28ddd26b85c9cf0b387cd0205d79557c929d72

Observation 8f320d94-7ca4-4ed6-b067-77c05124fbfe · outbound

This paper cites and Cygan, M.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Cygan, M

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.545410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:50.913669Z digest=sha256:bc0633aa3ad94c6527488e16d8ccd1817e8de6a762c6ad74947498849c1f1311

Observation 19548d37-4bec-43bc-af16-66f226479d82 · outbound

This paper cites Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.075651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.075651Z digest=sha256:faf220973c0dd6f154cef95055ef88c48cc65056d5e62f3154d113986be572cb

Observation b79eba20-6d7e-4ddb-b599-54112b534780 · outbound

This paper cites Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.287792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.287792Z digest=sha256:baf152b79c1585f7f48cfc7ee25c6a6a120aaad34fd5b4734875703021319fe5

Observation cc34fa6c-1f32-40c8-b836-35a24ae65bc6 · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Dinov2: Learning robust visual features without supervision

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.335629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:51.408062Z digest=sha256:f7aa8a34c54052178e20b8a370187bc58f977503e11ad9c182706d3bc83ec353

Observation 08140eb0-776a-4c7a-be59-1307dba4a575 · outbound

This paper cites Training language models to follow instructions with human feedback.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Training language models to follow instructions with human feedback

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.565324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.565324Z digest=sha256:59f564da82e3b4ce891ac0710b249e2a3c7f749d0c3b0ce558c7b320ec756bfb

Observation 3bba43ec-1eaa-4744-9ece-92d7769e9d23 · outbound

This paper cites Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.719632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.719632Z digest=sha256:0972fc2d40016beeb10fed32a5ac5788355eb2317978a3c6335578224d18ae1a

Observation d4eaee66-7207-4e20-a551-212ec8395e2c · outbound

This paper cites Is Value Learning Really the Main Bottleneck in Offline RL?.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Is Value Learning Really the Main Bottleneck in Offline RL?

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.847608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.847608Z digest=sha256:58e2f2f5780e87e9cacda8b52970788051ce24e75c010a582ba4f1baee9300d2

Observation f4dbe154-c0a7-4194-aec4-404adbed9de7 · outbound

This paper cites Flow Q-Learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Flow Q-Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:51.955281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:51.955281Z digest=sha256:b1147700af31be900ed04d7cb95b51b53551339649b110a32595a29d41cae47a

Observation 8b01f985-6863-4184-b4fd-d48391dbff62 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.070645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.070645Z digest=sha256:7ed0f15541fe5789dd6a1c872576b2d483f197f56400779639260ed9fbf9f2f1

Observation 5c9ea69e-97fa-4277-973e-b05b52d26d1f · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:03.113363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:52.157484Z digest=sha256:df28039795919127ad12096a871c9bc104544548837ae32e5dc755d4b5260aa9

Observation 1f3c6d5b-3fbd-431d-b5ed-0507340cfac5 · outbound

This paper cites Efficient off-policy meta- reinforcement learning via probabilistic context variables.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Efficient off-policy meta- reinforcement learning via probabilistic context variables

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:02.900845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:52.275833Z digest=sha256:87d39e3667935b576efc02e00af31f1e87236cc79f33cdb0373573c90a2f2efe

Observation 0935f476-d90e-4a90-ab89-f9845ff239e6 · outbound

This paper cites A Generalist Agent.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners A Generalist Agent

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.357891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.357891Z digest=sha256:10da76c561c9887314fa4c98ca3cbf3ceb464fcb897b112f38a859dbc7016e8e

Observation fe4849cb-8398-425d-b46a-d6a609b6f0a4 · outbound

This paper cites Policy Distillation.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Policy Distillation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.473229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.473229Z digest=sha256:7c1e6afe132232b28ff043f91f264f62ce14f67613b1a7c111e5c9e17227e8d6

Observation c54498a5-69f8-4676-9eeb-e9dc48e64daa · outbound

This paper cites Value-Based Deep RL Scales Predictably.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Value-Based Deep RL Scales Predictably

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:52.599976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:52.599976Z digest=sha256:6e833d15f2081047d7d73c167bf0e2387f070ead0c3d3c2807ae470496b4eb18

Observation 312c1e26-ea4e-4218-b376-37e15e0b2dfc · outbound

This paper cites D., Courville, A., and Bachman, P.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners D., Courville, A., and Bachman, P

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:02.700604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:52.687060Z digest=sha256:a2537f49dff3f3f2711fe20b5a1104244eb237af46f65fd312fc05418c53e4c3

Observation 2bf15b1f-6df2-48bb-9ed5-19c29a58c920 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:02.436833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:52.799338Z digest=sha256:3c11898f93e9c62836f59961c8ca5b9296447e3decd8dd77a4069279b0f42567

Observation 980db95b-203d-45db-9dc7-f0f6d759858c · outbound

This paper cites and Koltun, V.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners and Koltun, V

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:02.224585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:52.948108Z digest=sha256:bdb3e106943d33d2107c0e208ab962cd00166906614064cde37e1ae409fe7666

Observation 26518396-25da-4193-9b3b-c8c1a5ad8717 · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:53.082154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:53.082154Z digest=sha256:02c9fadab1ccb1415a567532a864d642accfcb82713e23adad34152c479444af

Observation 2d99e650-1b44-4599-ae47-53f38f3ab586 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:01.908957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:53.239036Z digest=sha256:4e3d2f9f0ad2ce2b697005676d05c0150021c417ef6fe51dc16ff7885bad98cd

Observation d63ad454-a067-4c4b-8a9c-4c038b1d35fd · outbound

This paper cites Deterministic policy gradient algorithms.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Deterministic policy gradient algorithms

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:01.644985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:53.386932Z digest=sha256:68cee51a2753b8001be09021620ba64f313debb521b67c6d68217026590cce55

Observation 6390f0ca-ea85-4fef-9b4a-b50d88db26b9 · outbound

This paper cites Mastering the game of go without human knowledge.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Mastering the game of go without human knowledge

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:53.487874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:53.487874Z digest=sha256:d303ebe043815f169243847120b54002a1309409682fb16092b5003f00fd6acc

Observation efe21a7b-af83-4aea-b805-06006d46249a · outbound

This paper cites Multi-task reinforcement learning with context-based representations.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Multi-task reinforcement learning with context-based representations

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:01.310812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:53.631167Z digest=sha256:81d63c51561ac2ee91f73133c2adacece2ba6c356fff4d6475e19d9b235bee4d

Observation c9aa705f-9b4b-44d3-a5c5-4b0ef5788f1a · outbound

This paper cites T., Abdolmaleki, A., Zhang, J., Groth, O., Bloesch, M., Lampe, T., Brakel, P., Bechtle, S.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners T., Abdolmaleki, A., Zhang, J., Groth, O., Bloesch, M., Lampe, T., Brakel, P., Bechtle, S

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:00.982770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:53.780565Z digest=sha256:ab4df1f0b2d9ee6acd334b424c6c7f9101bf3a8167b2f777577b03f6e75e4fc3

Observation 517bebf6-d433-4734-b018-4b8d1cc626b7 · outbound

This paper cites Paco: Parameter-compositional multi-task reinforcement learning.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Paco: Parameter-compositional multi-task reinforcement learning

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:00.682511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:53.893424Z digest=sha256:a691ebc0f56695024a599eeed86182c83f3955c126cd7a1369662980a12a37f4

Observation 3638475a-9732-4c67-a418-49cf4febf075 · outbound

This paper cites an unresolved cited work.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:54.004334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:54.004334Z digest=sha256:bace1ffccbea35d60874258262b265c3302a6eb3c975b6fac28c871a19811055

Observation f4797f4a-225b-42be-bab4-c83fa7fde7b1 · outbound

This paper cites DeepMind Control Suite.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners DeepMind Control Suite

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:54.169006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:54.169006Z digest=sha256:10801d73820bc1fceeeb306720246d140d40b00b3a274e294d5b1af4b5d0f3a4

Observation 9ad56e27-cc51-49c5-973b-9397858f3258 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Gemini: A Family of Highly Capable Multimodal Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:54.270164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:54.270164Z digest=sha256:770e3e195fc6cd788bef36a26067657ec1a899fff5b1b4ee17632f81d66d8784

Observation 13c97a2e-7c15-4072-ada9-4c0f675e3410 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners Octo: An Open-Source Generalist Robot Policy

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:54.408075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:54.408075Z digest=sha256:a1c864ab16f27875405180e38941cc796a19d6eada49e11d8182ee436443f1dd

Observation 4abc9372-2b55-4050-b912-923702c19438 · outbound

This paper cites M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R.

Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:00.375399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:58:54.542081Z digest=sha256:f13e3e7038d1a0a2eefb159d0bf89d4f013f752f2a881bc2859d7847083d38ac

Pith citing papers

Observation bc77ead2-1f30-42fd-9916-5280d93a49ee · inbound

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies cites this paper.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:05.516882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:05.516882Z digest=sha256:7748e4c5335b2feb586515554118a967451200665064a40ac161e7a8ee2cbbd6

Observation 0b865dce-dccd-424b-8539-56b73b144627 · inbound

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning cites this paper.

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 1937

Resolution
unresolved
no resolver link, observed 2026-08-03T06:26:50.060301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:26:50.060301Z digest=sha256:d42ab6d35a63b78ade148aa2fb7b1031e7625ad371b60171c8db48e858c53fb4

Observation 60621c76-8d7e-4e50-99fc-2dba385bb7b0 · inbound

What Does Flow Matching Bring To TD Learning? cites this paper.

What Does Flow Matching Bring To TD Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:36:17.973880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T16:32:29.432272Z digest=sha256:ce8194901598e5e27aa4376d4ffaef6ffb1b11ab48c19f468ac34f6c175a7e7f

Observation 888cc10a-ff82-4837-a3e4-b20e78351f66 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:49.716390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:04:56.512544Z digest=sha256:46be430b9df7e14effd369fcdb90da5e2b2687f481a92d04eeb1b7310c19d79f

Observation 9c5a25c9-0ec3-42b4-81d1-2f0be6bf0658 · inbound

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control cites this paper.

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:12:41.363803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T17:08:31.770889Z digest=sha256:86a21a3ad83bfc52e6a4e306fe6af02c44d72c8ac26136a1260608d791bb7e03

Observation 309e830e-b3ea-45f2-9ca8-588e9c61e2ee · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.607354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:33:19.038889Z digest=sha256:407002ee392878b209e5bf94a375536dd37bdaa37af463e8aec17c16e1f88094

Observation 2664d873-3fe7-4957-8bac-b87353feff58 · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:32:24.804374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:27:38.643667Z digest=sha256:c277b9f228215e5d7d5253672bfd0101edd0eb2cca64d88aad77b1f77894fdd7

Observation 85285cc0-7eaf-4f1f-be3b-64a47fdd07a0 · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:09:12.743772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T23:04:12.943222Z digest=sha256:e5e2d886bc9c7b58546bd57c8962f2b79d2b479f183ce95dbccad4e4114f01fc

Observation 08c63a93-2769-4044-b758-df2313d4dca4 · inbound

Debiased Model-based Representations for Sample-efficient Continuous Control cites this paper.

Debiased Model-based Representations for Sample-efficient Continuous Control Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:12:28.425885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:09:01.132424Z digest=sha256:a70b7dff3617276e119c5240afed174b8682d9f207c0c4a25fe850d673924d62