Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:24:26.960285Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.14448.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:24:26.960285Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T18:48:30.813878Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T22:37:25.557248Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7f894705-7617-445f-8c1b-2b17ce80d21a · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dd77321-6e4e-4a94-aace-179706d84337 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb8ea895-684c-4635-99f3-d9721e5a0e1a · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4d3df70-0895-4a22-b42d-a4c059332ce1 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c29089f0-f925-4735-9590-1650f110ef0c · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75ab87cd-d0fa-4021-9bcb-abe44b8215b0 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac3da65f-7370-49dc-b9df-67f934ddbf01 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Gonzalez, Ion Stoica, and Eric P
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f39ecf-b494-496b-9b11-0469089520fe · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cfcd0bb-fadb-4e6b-8ccf-fde2171ef875 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae836f5d-2ec9-4190-ab73-636145995dc8 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9002957-21d9-4875-a897-03d251502dee · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e67fbb3-4309-435e-a7bb-998a4539be9f · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Amago: Scalable in-context reinforcement learning for adaptive agents
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba936084-338a-4836-8f68-876ad2fc79ea · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e48c53b-f9ff-4c69-9468-c13145611a1f · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison GPT-4o System Card
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad5dd136-440b-4ddf-8c69-b71d2f8a32f5 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison OpenAI o1 System Card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12608b17-a6aa-4d7f-97d5-7625df16776e · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison SelfEvolve: A Code Evolution Framework via Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b5855c9-6d5c-4dd9-8bee-06200e41d7e9 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bba2ab71-24b2-488c-b42d-8d7e8f632fa1 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b1caedf-ded2-4acc-a86c-4754fe1e3c96 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison In-context reinforcement learning with algorithm distillation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c6af786-b0d3-4876-890e-8458e295b795 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47ec9c2b-cedd-4086-a6e1-e12d4bed5258 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison DeepSeek-V3 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 573640db-7a0d-4cfd-a009-6ce75db25cc0 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5071ce91-e83c-46c5-91d4-ad1c1eda275a · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73be322b-c6bf-45bc-9a77-51b52d489d37 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16619bf5-2e33-4f78-b84c-67bd1957301d · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison WizardCoder: Empowering Code Large Language Models with Evol-Instruct
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ec30262-2949-4f68-9c9c-9d176f8d17af · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50a679f9-b410-483b-b832-9b35f0a99539 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 923dfd12-a7d8-45a1-ab01-8732785c259d · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison POPGym: Benchmarking Partially Observable Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02ef6cc0-cd2b-4919-9ef7-840083d83377 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b7c9f4d-3da1-4bed-b0ac-24bc2df53085 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8023024d-db22-4845-85ff-8ddb260776a7 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f30a8df-16de-4e8e-86ce-a5ff348464a1 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe21e204-5cd9-4e77-9bd2-4bc281f6dd26 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2ad8fd8-3413-4166-9fde-dfe714fb6a3d · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf691fe-db50-4fee-9b98-96c3a4b4752c · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison A Survey on Self-Evolution of Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b4784a3-2c23-40a1-b635-1fd39b2e067a · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1ec01af-17f8-4dd0-8250-045b5c5082d9 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e475ccbc-afae-471d-ad57-36119217c0b6 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0cfb8b71-f14a-44d6-b8f3-c7a381c3cc12 · outbound
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bada8048-acc7-44a9-9366-d65717504109 · inbound
From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.