Pith. sign in

Paper Citation Record · LEDGER

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution

As of 20 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2412.13492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13492 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:09:46.548746Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed7f5a96-23e4-476c-9472-1dbf3302eb33 · outbound

This paper cites an unresolved cited work.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:09:46.863199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.476073Z digest=sha256:20eaaae03ef1be845cba6572943a822866a46656869895348d0c183f771b1988

Observation 15149894-1abf-42a3-8bbc-b56c3608e1b6 · outbound

This paper cites an unresolved cited work.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:09:46.821947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.489993Z digest=sha256:5746cca89e738dad65bb68cf7757303c7f6df6967b0eb2bc55602e6ffc3bbab8

Observation 17e6cd46-e057-42db-96e0-217da7f9a3af · outbound

This paper cites Under no circumstance can you in- troduce new input variables.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Under no circumstance can you in- troduce new input variables

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:09:46.787211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.497074Z digest=sha256:82c98f92cc75e90ad601330f2ea92460944c411ac8d9a966c5335a07e0742361

Observation 5acad43c-9f9f-4d8a-88c2-6d4efb5cc494 · outbound

This paper cites Each transformed reward component should have its own temperature variable.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Each transformed reward component should have its own temperature variable

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:09:46.842675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.482494Z digest=sha256:b632f070f3a268717e10f33031424848dc13387fa754e3c4cca93f045c7f9e2b

Observation ade6ac28-ae59-4184-9fb1-63855f73ebc2 · outbound

This paper cites an unresolved cited work.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:09:46.765430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.504269Z digest=sha256:7cbefcc2c90d2336ec97fe88420d8ea5a6e9126e5518b4aa7c591a3f9b2be0e7

Observation d0377722-aeeb-41e6-9960-83f71f4bf7cc · outbound

This paper cites You may consider: (a) Changing its scale or the value of its temperature pa- rameter (b) Re-writing the reward component (c) Discarding the reward component.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution You may consider: (a) Changing its scale or the value of its temperature pa- rameter (b) Re-writing the reward component (c) Discarding the reward component

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:09:46.741325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.510691Z digest=sha256:1084ff4c21bf9c60a0b4deeb4a4983a31b593486e5153dd9f24a0a2621b9b4aa

Observation 7d02b0e4-dc15-41d9-8421-0637dc9537db · outbound

This paper cites Please analyze each existing reward component in the sug- gested manner above first, and then write the reward function code.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Please analyze each existing reward component in the sug- gested manner above first, and then write the reward function code

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:09:46.706990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.516657Z digest=sha256:483ea8e81a8ebde47b54b5f05f838a0081c243e31c44ef8a5768c106ccca996a

Observation 9dc958a7-a338-4c28-bca7-116918188d81 · outbound

This paper cites ‘‘‘python ... ‘‘‘.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution ‘‘‘python ... ‘‘‘

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:09:46.886098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.522384Z digest=sha256:c0bbc4191d9af56f66b33f4a201a5e2296c73bf61ab0df4e24987a32a12fd11a

Observation 614c3def-6066-4f28-8689-2793cd4614db · outbound

This paper cites an unresolved cited work.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:09:46.680829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.529365Z digest=sha256:a251168512c169123ab0f2d0dce38631fb7d56d13491e2cb3f23afea251cf504

Observation b3385918-2d74-410d-b0c2-5012f47370f1 · outbound

This paper cites Each transformed reward component should have its own temperature variable.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Each transformed reward component should have its own temperature variable

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:09:46.661971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.536082Z digest=sha256:6d39d972ddee60ce91a7ddff5fce275de37b6354eed0f1c3411286d704c30cdb

Observation a0b0bbf0-220e-41a4-bad2-3dc3705f9644 · outbound

This paper cites an unresolved cited work.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:09:46.640643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.543141Z digest=sha256:bfd588741937f3581126b78c4a7cfee50db7e0ca0c7dbb77285c10e2ba5e45da

Observation 08753ce8-d286-45f5-b2db-041e3dc195e9 · outbound

This paper cites Un- der no circumstance can you introduce new input vari- ables.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Un- der no circumstance can you introduce new input vari- ables

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:09:46.619628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.548746Z digest=sha256:de16995f67b8b0209052046297e007996923de38824b4ff50b7ebe910085d7bf

Observation 53867a7a-1d4e-4f45-aff5-9ade45212742 · outbound

This paper cites Maintaining Plasticity in Deep Continual Learning.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Maintaining Plasticity in Deep Continual Learning

Reference 318

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:46.454679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:46.454679Z digest=sha256:65be2b81165baeb72e608e63ea87e6c7e4428ff0303b47790e80968a8a0258af

Observation 64c91e9a-1dba-4cd0-b8a6-fb77167e6ab9 · outbound

This paper cites initial prompt.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution initial prompt

Reference 2008

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:09:46.906425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T13:09:46.462121Z digest=sha256:29625162351623348d99f00a019e448043131d5a6b46a88b2b4448fcafd6ef48

Observation dbae3db7-13b4-4f0a-9134-b85f42e72b46 · outbound

This paper cites In Conference on robot learning, 287–.

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution In Conference on robot learning, 287–

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:46.448315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:46.448315Z digest=sha256:de4383b9a0b76ad27258cfa7cbf7b5120585d84b2580f8c07fa5074fe195a350

Pith citing papers

No inbound Pith citation observations are available.