Pith. sign in

Paper Citation Record · LEDGER

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison

As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.14448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14448 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:24:26.960285Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:48:30.813878Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:37:25.557248Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f894705-7617-445f-8c1b-2b17ce80d21a · outbound

This paper cites online" 'onlinestring :=.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:23.786452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:23.786452Z digest=sha256:055960308fba6c2ee838f0d30547714241931ceb3297a4d9da446ad800bb8fc6

Observation 3dd77321-6e4e-4a94-aace-179706d84337 · outbound

This paper cites write newline.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:23.859578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:23.859578Z digest=sha256:117af0597395c1f13388fbfddce903ca8d0c052c18acd59a4fed59091d8d8f75

Observation bb8ea895-684c-4635-99f3-d9721e5a0e1a · outbound

This paper cites LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:23.916722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:23.916722Z digest=sha256:07c1d8d8f3bcd4a3c7e006adee2e3628ac027f7181154ea55d649d1462bf18d1

Observation b4d3df70-0895-4a22-b42d-a4c059332ce1 · outbound

This paper cites GPT-4 Technical Report.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.013600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.013600Z digest=sha256:1c4b3e531e089f7cebb78fe2c479df893e8ca252d2912dfc3049719b7b6bb589

Observation c29089f0-f925-4735-9590-1650f110ef0c · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:30.153815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.080497Z digest=sha256:b425cd69ca422ac77345124f89949b47f6906b11f1b3cc93106e01a3875068e0

Observation 75ab87cd-d0fa-4021-9bcb-abe44b8215b0 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.170968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.170968Z digest=sha256:825f933cf17d86d701d292ce69e2f1fe07a6b9c94864c80ca21cc7933a8cc17a

Observation ac3da65f-7370-49dc-b9df-67f934ddbf01 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Gonzalez, Ion Stoica, and Eric P

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.254343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.254343Z digest=sha256:e37c4da9647e7ec145221e526780524cc16ae2c5e30b72259701376c426b8c2f

Observation b7f39ecf-b494-496b-9b11-0469089520fe · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.333492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.333492Z digest=sha256:744d943ffc77f7cbbd5b8fc29e286542a328fe8dd6b4ff5c9689fef966680ef8

Observation 9cfcd0bb-fadb-4e6b-8ccf-fde2171ef875 · outbound

This paper cites RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.397362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.397362Z digest=sha256:470443edaa4f3903bebf8e2f0574f84ad523cc82009a153b8e6467b3849e4596

Observation ae836f5d-2ec9-4190-ab73-636145995dc8 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:30.005234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.455525Z digest=sha256:7965b3913f00e912851b687329a629a0bacd75db806be2654541cea1f2eb1ad5

Observation d9002957-21d9-4875-a897-03d251502dee · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:29.877594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.511202Z digest=sha256:5fde8d3cd9965305828c83ba6b33c5915226f815025b0b7765b3fc698cceaddd

Observation 9e67fbb3-4309-435e-a7bb-998a4539be9f · outbound

This paper cites Amago: Scalable in-context reinforcement learning for adaptive agents.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Amago: Scalable in-context reinforcement learning for adaptive agents

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:24:29.723566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.568343Z digest=sha256:6ee67a74fe18ac4a06ae5ac21362f2a50faf3dc5eb16c3ec00215baff65e6f4b

Observation ba936084-338a-4836-8f68-876ad2fc79ea · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.623902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.623902Z digest=sha256:9ca2335057e504ee02462dbf410654765d1a13842aeaef2a37266277e735345d

Observation 4e48c53b-f9ff-4c69-9468-c13145611a1f · outbound

This paper cites GPT-4o System Card.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.682267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.682267Z digest=sha256:24107a1e6f0a56757e103599bd2e0dd0fe92199fb17a967353b348577fa91cd3

Observation ad5dd136-440b-4ddf-8c69-b71d2f8a32f5 · outbound

This paper cites OpenAI o1 System Card.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.783283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.783283Z digest=sha256:197deb9353f388cf872f86215733f4ca8a4ed9e3518114320260be56ab8ac048

Observation 12608b17-a6aa-4d7f-97d5-7625df16776e · outbound

This paper cites SelfEvolve: A Code Evolution Framework via Large Language Models.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison SelfEvolve: A Code Evolution Framework via Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:24.831453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:24.831453Z digest=sha256:774a5b5bc2287b2ac5b30952c13ca47649e5d12b2d3c43027dc3a2b6a6b44888

Observation 2b5855c9-6d5c-4dd9-8bee-06200e41d7e9 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:29.538934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:24.881597Z digest=sha256:d5776137fa82bf8c375e40ef9930b85829e4fa756b8c669cba740595cb66fe4b

Observation bba2ab71-24b2-488c-b42d-8d7e8f632fa1 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:29.358637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.014205Z digest=sha256:c630a4ae2a11e906bf289c1d617d1e5da2f0f7d81bdac5c6fe1997b87afb8412

Observation 3b1caedf-ded2-4acc-a86c-4754fe1e3c96 · outbound

This paper cites In-context reinforcement learning with algorithm distillation.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison In-context reinforcement learning with algorithm distillation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:24:29.146080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.090138Z digest=sha256:65c812700b9adc79287294c74b5b635a53d77e614188eab430384842e87f63e7

Observation 4c6af786-b0d3-4876-890e-8458e295b795 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.960005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.170461Z digest=sha256:06c9598d3f4f86b3e707d6b1e6c05de9a34b03143ddead01315ab1d3d73f6ba6

Observation 47ec9c2b-cedd-4086-a6e1-e12d4bed5258 · outbound

This paper cites DeepSeek-V3 Technical Report.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison DeepSeek-V3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:25.256888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:25.256888Z digest=sha256:59502cd732d8fb21dfdf1ecf670549c7122365892686144461909e02c298d7fe

Observation 573640db-7a0d-4cfd-a009-6ce75db25cc0 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.760291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.350340Z digest=sha256:ddad9ceea257ed3c54b83057b333ffaf42310373736ed3ac0f938088804ee018

Observation 5071ce91-e83c-46c5-91d4-ad1c1eda275a · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.498765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.479508Z digest=sha256:ed5e9d3bde16769893413c37575c9c1294bae5ce262a20eff329aadcb87232e1

Observation 73be322b-c6bf-45bc-9a77-51b52d489d37 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.249737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.552776Z digest=sha256:9ca619fea4d5cc77bf074b15979882c25004b98ab7d75039dc703fc541bb3801

Observation 16619bf5-2e33-4f78-b84c-67bd1957301d · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:25.657136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:25.657136Z digest=sha256:46c9d07f843d1c0a84596248cc593d721badc4d8fba24d12faee822ae8fcfe7c

Observation 9ec30262-2949-4f68-9c9c-9d176f8d17af · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:28.070041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:25.734272Z digest=sha256:be6d17a6c254dd29b74f73753012f563eaefc1c4ea34c19b653cdb3121043dc8

Observation 50a679f9-b410-483b-b832-9b35f0a99539 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:25.851359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:25.851359Z digest=sha256:2a3a5f635a72ded970111db561163d8fdab324a29c981175d567c304fb1562ea

Observation 923dfd12-a7d8-45a1-ab01-8732785c259d · outbound

This paper cites POPGym: Benchmarking Partially Observable Reinforcement Learning.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison POPGym: Benchmarking Partially Observable Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:25.955282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:25.955282Z digest=sha256:d2485fbb56f9bf9d79f559789634932ab5d0626eea3a1b8de989337eb004393c

Observation 02ef6cc0-cd2b-4919-9ef7-840083d83377 · outbound

This paper cites Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.043298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.043298Z digest=sha256:0b0dcc2f3be408e32327f3f5a43c20bdd1eca8ed342497921cb0c7a7b38a1be9

Observation 5b7c9f4d-3da1-4bed-b0ac-24bc2df53085 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:27.925468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:26.131603Z digest=sha256:ee9b0a4ba183d42c788b59693e6f77ff2306944d4ad01ea27fc5bdad0f6132dc

Observation 8023024d-db22-4845-85ff-8ddb260776a7 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:27.763206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:26.237012Z digest=sha256:a44783cf82f3a14e6d4f263d368292761c221f3bc132f87b14c7f85b25eed337

Observation 5f30a8df-16de-4e8e-86ce-a5ff348464a1 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.328652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.328652Z digest=sha256:2229b77492b06c9ef06188031da5575d408166f123c3067200ebd235068d68f1

Observation fe21e204-5cd9-4e77-9bd2-4bc281f6dd26 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:24:27.572140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:26.426949Z digest=sha256:e28b72820a0dd848a6e7bbfdd4ee8042fd047ebbeb502084d7c4bf7353b5895d

Observation c2ad8fd8-3413-4166-9fde-dfe714fb6a3d · outbound

This paper cites Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.491541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.491541Z digest=sha256:a078fb1bf47e1229a0037d9026ceba83eb618dadc0b2c2dbea704c800c2db1b5

Observation 6bf691fe-db50-4fee-9b98-96c3a4b4752c · outbound

This paper cites A Survey on Self-Evolution of Large Language Models.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison A Survey on Self-Evolution of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.554434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.554434Z digest=sha256:3c083c8206ec2d151271e1270d527962a42c738bdb63c6f1ca98546d62edc6e5

Observation 6b4784a3-2c23-40a1-b635-1fd39b2e067a · outbound

This paper cites MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.690889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.690889Z digest=sha256:253c704cb453fabe08da190bf692a8e52487112d8e3e6b8cc4dbf2f99fa2be2d

Observation f1ec01af-17f8-4dd0-8250-045b5c5082d9 · outbound

This paper cites an unresolved cited work.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.798146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.798146Z digest=sha256:92dc52e378e67e180fb9f292084013097d6bb896f417a7aac744d334a7067e4f

Observation e475ccbc-afae-471d-ad57-36119217c0b6 · outbound

This paper cites PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:24:27.227107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:24:26.878378Z digest=sha256:1f47685177049a039743cb6ed69efc9897bcca64f435df70f79d074de35ecba6

Observation 0cfb8b71-f14a-44d6-b8f3-c7a381c3cc12 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:24:26.960285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:24:26.960285Z digest=sha256:5b7c7e8cfb49c72e74167b18f85e29f857037fdcc82a0e2a791e1322c70109dd

Pith citing papers

Observation bada8048-acc7-44a9-9366-d65717504109 · inbound

From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory cites this paper.

From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:25.558920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T18:48:30.813878Z digest=sha256:e607a605cd989a691e391b36e6a6b6427f24e1f14559ed435ccc6107a614308e