Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:09:46.548746Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2412.13492.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:09:46.548746Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ed7f5a96-23e4-476c-9472-1dbf3302eb33 · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 15149894-1abf-42a3-8bbc-b56c3608e1b6 · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 17e6cd46-e057-42db-96e0-217da7f9a3af · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Under no circumstance can you in- troduce new input variables
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5acad43c-9f9f-4d8a-88c2-6d4efb5cc494 · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Each transformed reward component should have its own temperature variable
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ade6ac28-ae59-4184-9fb1-63855f73ebc2 · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d0377722-aeeb-41e6-9960-83f71f4bf7cc · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution You may consider: (a) Changing its scale or the value of its temperature pa- rameter (b) Re-writing the reward component (c) Discarding the reward component
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7d02b0e4-dc15-41d9-8421-0637dc9537db · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Please analyze each existing reward component in the sug- gested manner above first, and then write the reward function code
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9dc958a7-a338-4c28-bca7-116918188d81 · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution ‘‘‘python ... ‘‘‘
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 614c3def-6066-4f28-8689-2793cd4614db · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b3385918-2d74-410d-b0c2-5012f47370f1 · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Each transformed reward component should have its own temperature variable
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a0b0bbf0-220e-41a4-bad2-3dc3705f9644 · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 08753ce8-d286-45f5-b2db-041e3dc195e9 · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Un- der no circumstance can you introduce new input vari- ables
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 53867a7a-1d4e-4f45-aff5-9ade45212742 · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution Maintaining Plasticity in Deep Continual Learning
Reference 318
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64c91e9a-1dba-4cd0-b8a6-fb77167e6ab9 · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution initial prompt
Reference 2008
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dbae3db7-13b4-4f0a-9134-b85f42e72b46 · outbound
Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution In Conference on robot learning, 287–
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.