Pith. sign in

Paper Citation Record · LEDGER

Training Language Agents to Learn from Experience

As of 7 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2605.20477.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.20477 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T07:11:09.642275Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact11
  • verified fuzzy16
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0575f69-4592-4cdf-b54c-b1ac1b1a13f0 · outbound

This paper cites GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.

Training Language Agents to Learn from Experience GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:14:02.609943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:4c7f159c4605d28932e749bbf80b19dca74abf0c0542e1073736fd5b6d09771a

Observation 7335257d-2e04-41c7-8091-86fc361f1099 · outbound

This paper cites Openai gym.

Training Language Agents to Learn from Experience Openai gym

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.731454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:889e1335351555bf9a03ca59db089f61b33ea7a3523dc3d76b160560cc5967fa

Observation 9bc94e4c-2619-41f1-a1a8-b67f6717569e · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901.

Training Language Agents to Learn from Experience Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.733794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:4c5c985ad07ae550067c9ff7179a85fc496805bb2af0ad6e63665b640b699173

Observation c464cd0b-aa18-4ada-89c9-67c9997ece16 · outbound

This paper cites System prompt optimization with meta- learning.

Training Language Agents to Learn from Experience System prompt optimization with meta- learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:14:02.620371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:e703d3bbae64ee0e7d8bc149ab7e9005a80968e87ea816838c0c7109fada1d16

Observation 08f9f6e7-465d-41d0-a05e-41a2911fae6d · outbound

This paper cites Improving retrospective language agents via joint policy gradient optimization.

Training Language Agents to Learn from Experience Improving retrospective language agents via joint policy gradient optimization

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.727285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:927470395e27f09c8a1a5b9717244b003316691ea852fedec543435d7ffc4c98

Observation d2d0aab2-8033-4101-858f-74a3afbb3930 · outbound

This paper cites Samule: Self-learning agents enhanced by multi-level reflection.

Training Language Agents to Learn from Experience Samule: Self-learning agents enhanced by multi-level reflection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.729416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:d63381178d928e75fe42d58023b0d9438a9623023cec535d6dd5405da7a3e03b

Observation b2792107-6ad0-473d-9d75-99635800c804 · outbound

This paper cites Meta-rl induces exploration in language agents.

Training Language Agents to Learn from Experience Meta-rl induces exploration in language agents

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.725235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:ea5485943cda6eee9a67ceda8274bd18aa3f76a36cef1e284a9ec0106234190b

Observation fb825063-655f-4ad6-98ce-cc7398056407 · outbound

This paper cites LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations.

Training Language Agents to Learn from Experience LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:14:02.612567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:4dac045b763bee49542d3d50abc077e67d16bd3787f4c68b4c7382965832a90f

Observation b56f7cea-0bc9-4139-8a11-4423d9966f1c · outbound

This paper cites Minihack the planet: A sandbox for open-ended reinforcement learning research.

Training Language Agents to Learn from Experience Minihack the planet: A sandbox for open-ended reinforcement learning research

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.718180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:19fea3490a62d0d95c912e4f2cf193ddda5eb67b7609891f25399a2112bdb1a5

Observation f3822e6a-619b-4f23-9b07-9fa736463b64 · outbound

This paper cites Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks.

Training Language Agents to Learn from Experience Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:14:02.632692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:2cb6bac32cc23fb8b658c2fa5f09950c2e53fc6df51666105b0fc0868f08afea

Observation 3779a81b-68be-4745-beeb-8cebec9963ee · outbound

This paper cites Can foundation models actively gather information in interactive environments to test hypotheses?.

Training Language Agents to Learn from Experience Can foundation models actively gather information in interactive environments to test hypotheses?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:14:02.635780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:22207ca8dd7db077befc3a1b08962c5b1a98d98d4a9a50c46a6f2acf01ed4d32

Observation 7c08bb24-47fe-4f83-9c00-96af5b80d4cc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Training Language Agents to Learn from Experience DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:14:02.622820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:77bfb7600a30d01d3c4676e614510900b3d2c8f31b1b2ada93693029c16a451d

Observation 4f481cf5-6a9b-4367-b737-de41b94ac7d9 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Training Language Agents to Learn from Experience HybridFlow: A Flexible and Efficient RLHF Framework

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:14:02.625298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:ef49b530d4c02c6443835ae535b675238ea583069ac2a2c5965e458cfa9585af

Observation 82a2cfd3-3ddc-4405-aafd-be070015b8f4 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652.

Training Language Agents to Learn from Experience Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.720600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:e15ea43acb287476c65c673c07e59e2a30abc07e01043b5f981e3bbf595df91c

Observation 61856e92-1d27-48ee-ace9-0b2603b1ddf9 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Training Language Agents to Learn from Experience ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.712234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:592e3558c8495165da178be930526a15a4336ad00d70f56ff371b547af5a9641

Observation 81080f43-7494-4373-ba7a-f76f71833c8b · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Training Language Agents to Learn from Experience ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:14:02.617944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:b85d35976dbb1e5650ff32597cfddab827da771ca370c5fddf35de31be7c2e51

Observation 2781f0a4-ce1d-443d-8178-5212b9c6351d · outbound

This paper cites Welcome to the era of experience.Google AI, 1.

Training Language Agents to Learn from Experience Welcome to the era of experience.Google AI, 1

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.714201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:255f77482a088d696b90832dd55027ab8575fa4dead31c95ef253d6fe8dc1450

Observation 79819df9-1c0c-4542-acb8-32aeab412892 · outbound

This paper cites Cognitive architec- tures for language agents.Transactions on Machine Learning Research.

Training Language Agents to Learn from Experience Cognitive architec- tures for language agents.Transactions on Machine Learning Research

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.705714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:4805a0d732d20dec54a2cb773dd552d4294b08ba9c07f640029f0caad7337599

Observation 9a165cc2-7fc3-4e0f-afde-191b8ac2a5ce · outbound

This paper cites A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345.

Training Language Agents to Learn from Experience A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.708114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:b2641ef25f65918a8e44cb77a65bd44d2321a454ca92ca6d960dfc5403c86c1e

Observation 26220ed4-5bac-409c-a97f-eb8cd1f98c7a · outbound

This paper cites Qwen2 Technical Report.

Training Language Agents to Learn from Experience Qwen2 Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:14:02.630011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:fd1a3720a075ae2200b34691fb7adf657368be80058dae1771d2122c8443f8b1

Observation 044bf3c0-48eb-417c-91c5-12ad3182d313 · outbound

This paper cites Large language models as optimizers.

Training Language Agents to Learn from Experience Large language models as optimizers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.716175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:1b3bace6826b5cf5b08aff1db1921bfed08260267af7e7af92c0c352fb502f33

Observation 6d60c726-6090-4c48-84da-7651d43b299c · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Training Language Agents to Learn from Experience React: Synergizing reasoning and acting in language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.723025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:1f1bda6de92538ffd41911fbd7e5632a84b5274dbe932d92afc1c73b74cfb69e

Observation b931cb18-5ab5-47e2-b9d4-e2fa4fe981fe · outbound

This paper cites Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization.

Training Language Agents to Learn from Experience Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:14:02.615242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:5fb250ec6a4d4b30fc5d676016bdfb312e5cd13032005e949fe86cbecca57433

Observation 9b13a92c-9b9e-4214-9dff-29c3405f7973 · outbound

This paper cites Assessing Adaptive World Models in Machines with Novel Games.

Training Language Agents to Learn from Experience Assessing Adaptive World Models in Machines with Novel Games

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:14:02.627802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:789c573fbf35a3f0c9ae6bd5f2b09d558d50819c18bf3daa83ccfaab426e84ac

Observation 5f50fbc8-9a26-4340-87c2-cf49a37a18d3 · outbound

This paper cites Expel: Llm agents are experiential learners.

Training Language Agents to Learn from Experience Expel: Llm agents are experiential learners

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.710188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:ce5907e6d05993a5a1fb6aa2338238bc370bd7c55f160eb642d3dcd4c5294861

Observation 50480d2a-a712-4630-b675-14547b666410 · outbound

This paper cites an unresolved cited work.

Training Language Agents to Learn from Experience Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-21T07:14:46.703743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:c66e15522a6eef28faf6dbab967337533c3bdc597f7c81fc4fba5bcfd9f82d69

Observation 39e026f5-c9a7-40d2-9d95-62e76332d934 · outbound

This paper cites an unresolved cited work.

Training Language Agents to Learn from Experience Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-21T07:14:46.700145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:a3f182634d0be953c49d136b88029045f50d42ba73c4ed4dedaecca36046e241

Observation 188e03b2-acb2-4554-b34c-9ad0f97d2d11 · outbound

This paper cites Qwen/Qwen2.5-7B-Instruct.

Training Language Agents to Learn from Experience Qwen/Qwen2.5-7B-Instruct

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.701969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:bf2693fe7c467afe7912a91176a2be51efeab0f61a45fb489c32ff3aa3e710e1

Observation 1f25f9fa-c319-40ce-b8b2-b7ee59cff123 · outbound

This paper cites an unresolved cited work.

Training Language Agents to Learn from Experience Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-21T07:14:46.696464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:8351404817636405c84b4b4c7859fa749617cbda5af4e6edf73a6dd225069224

Observation 9c149829-0e1b-4dde-b3de-8e51b1eed1bc · outbound

This paper cites an unresolved cited work.

Training Language Agents to Learn from Experience Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-21T07:14:46.698343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:80e3ee75cf27c8ef9e05f43c76b9f5e788f5d29114e3eccf1577afa46452cd82

Observation e5f32a59-ed69-4c71-a145-49271e60f199 · outbound

This paper cites an unresolved cited work.

Training Language Agents to Learn from Experience Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-21T07:14:46.691346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:e2e6ef0e20b5f50347394705cc268c73270a13ce29b29e8ad68385702abc6f58

Observation e97f10c0-e480-44c2-8264-38e5828d5b3a · outbound

This paper cites <" symbol. I will move east to investigate further. Action 2: step e Observation 2.

Training Language Agents to Learn from Experience <" symbol. I will move east to investigate further. Action 2: step e Observation 2

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:14:46.693384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:11:09.642275Z digest=sha256:cb76ad018aa2d6ba3f96571aec5043bbc7b54835eedaf2f1408b5e54e1393997

Pith citing papers

No inbound Pith citation observations are available.