Pith. sign in

Paper Citation Record · LEDGER

Disentangling Exploration of Large Language Models by Optimal Exploitation

As of 11 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2501.08925.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08925 v3

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:18:58.530774Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:18:58.266458Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T20:18:59.380251Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact6
  • verified fuzzy9
  • unresolved48
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c89830e-10d2-4c37-a231-a2a07db3797d · outbound

This paper cites Surprise-Based Intrinsic Motivation for Deep Reinforcement Learning.

Disentangling Exploration of Large Language Models by Optimal Exploitation Surprise-Based Intrinsic Motivation for Deep Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.173568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.173568Z digest=sha256:42e8681cd1fc759437b0e856fa180155f21ad96a099431fd05bd6c8bc2b24259

Observation 874763f3-8e68-4487-8ec9-6e6b1e88c452 · outbound

This paper cites GPT-4 Technical Report.

Disentangling Exploration of Large Language Models by Optimal Exploitation GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.179775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.179775Z digest=sha256:e0607298154ddae30676e0dcd23d21eaa801a20c5cfe2c8bce1261f6291d8062

Observation a32f13e5-e1e4-4c02-9dfe-fd2616a85bcd · outbound

This paper cites an unresolved cited work.

Disentangling Exploration of Large Language Models by Optimal Exploitation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:18:59.798920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.185892Z digest=sha256:6cec5b4f9c00d6148e0035bfe70f2950c5691e3b91f11e2127cbdae2ef5fd628

Observation 0e8f0cb4-ded5-4b0e-949a-fed2354d680c · outbound

This paper cites an unresolved cited work.

Disentangling Exploration of Large Language Models by Optimal Exploitation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:18:59.765360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.196787Z digest=sha256:3b070b154acd0af1d9821d90bc6fa6ba79fcb95aa38fe13cfae72be9ef65da60

Observation 9b49b0bd-d869-49cd-8799-23cd752ef79b · outbound

This paper cites Claude 3.5 Sonnet.

Disentangling Exploration of Large Language Models by Optimal Exploitation Claude 3.5 Sonnet

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:18:59.749185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.207925Z digest=sha256:773cb56e5864d06b35464adad45ae54ac894d69fd711c0cc9061ce8acf25cfba

Observation 1f53e069-f486-444b-afd7-614d579c7beb · outbound

This paper cites an unresolved cited work.

Disentangling Exploration of Large Language Models by Optimal Exploitation Unresolved cited work

Reference 6

Resolution
parse uncertain
raw_fallback, observed 2026-08-10T20:18:59.782006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.202791Z digest=sha256:f1e67179ca8180d7b214fe1f2165c40936589043b6a5d9fc154b49d6254ad056

Observation 9c4dc64a-6f48-4590-ac21-b9a8c12a10f2 · outbound

This paper cites Bellemare, S.

Disentangling Exploration of Large Language Models by Optimal Exploitation Bellemare, S

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.220781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.220781Z digest=sha256:eb55f6dcc3d33345c226a82c46fff27eb4e3db98e34b12f7087c6b738e81f35d

Observation a6aabf19-a49c-4207-bac2-5613d30d567f · outbound

This paper cites Decoupling Exploration and Exploitation in Multi-Armed Bandits.

Disentangling Exploration of Large Language Models by Optimal Exploitation Decoupling Exploration and Exploitation in Multi-Armed Bandits

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:18:59.485496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.214412Z digest=sha256:f2e24e36c34842fd0607dcb9ac8ee2ba92149d1abf621f4d2ab15e0cc4997c03

Observation 347ead80-e5c4-4621-829e-4e98a9ceb1ae · outbound

This paper cites Efficient Sequential Decision Making with Large Language Models.

Disentangling Exploration of Large Language Models by Optimal Exploitation Efficient Sequential Decision Making with Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.235151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.235151Z digest=sha256:f905770c36a97a71807974eb0c750f84f191cc9df74d1fafd8d6f04ba2a432fb

Observation cb8c605a-9b60-4b3f-bc7c-8f6cdb4b7128 · outbound

This paper cites Exploration by Random Network Distillation.

Disentangling Exploration of Large Language Models by Optimal Exploitation Exploration by Random Network Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.227604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.227604Z digest=sha256:df05b2997c7a651aaf9c0dc05b79096a91359e4d4f339663e4c2304716ed0c36

Observation 8aaab6d8-369a-4d56-9696-92deea6f4c9d · outbound

This paper cites Chevalier-Boisvert, B.

Disentangling Exploration of Large Language Models by Optimal Exploitation Chevalier-Boisvert, B

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:18:59.719749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.246134Z digest=sha256:d4063bc132ad591a2f7cc6d87e240f181fbac9ec860f34c5c0d38b0ed10f6182

Observation 809b04b4-4169-4548-80bd-e8bf0d814cc7 · outbound

This paper cites BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning.

Disentangling Exploration of Large Language Models by Optimal Exploitation BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.240737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.240737Z digest=sha256:35462485f93756052beceb260b0371da82969e59006b6943e41d366a6066f6cd

Observation 0a4710df-e13a-4980-bca6-e52c5ba6cd47 · outbound

This paper cites The Llama 3 Herd of Models.

Disentangling Exploration of Large Language Models by Optimal Exploitation The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.255893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.255893Z digest=sha256:bd7191286d8ec75db2221bfa8354abc57b2e5f8019bf318a88382a77c5ea7f66

Observation 486a8b31-40a1-4d5b-ad7c-0c363023fcea · outbound

This paper cites an unresolved cited work.

Disentangling Exploration of Large Language Models by Optimal Exploitation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:18:59.702888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.250970Z digest=sha256:d8c3a6b27d149fea89a6901ae0af15a8cb8f069cdb77b052e9ee1a0c337209cd

Observation fa4690cf-90ec-4c9a-9a49-e4cb0294a079 · outbound

This paper cites Disentangling Exploration of Large Language Models by Optimal Exploitation.

Disentangling Exploration of Large Language Models by Optimal Exploitation Disentangling Exploration of Large Language Models by Optimal Exploitation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:18:59.385875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.266458Z digest=sha256:3caff73d7005ff4531e9a34082f8ea4ce58f36d0bf713d7aef8c6f963d9448b3

Observation 1b64d18c-6134-4684-b248-88d1f45436a7 · outbound

This paper cites Go-Explore: a New Approach for Hard-Exploration Problems.

Disentangling Exploration of Large Language Models by Optimal Exploitation Go-Explore: a New Approach for Hard-Exploration Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.261205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.261205Z digest=sha256:c02b69dce1578921cea0cc3f5cea758a4b21f9ad16afbca7ab0a9abb4953a11f

Observation 80b17858-6ed7-4884-8aa1-45062e23bc9c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Disentangling Exploration of Large Language Models by Optimal Exploitation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.276424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.276424Z digest=sha256:d6bf59839f1559473094268b119b79c3a00cc129357d957e432d034ceceb6fc2

Observation 1707be96-f40c-454b-af6a-dc5723a0ef14 · outbound

This paper cites AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents.

Disentangling Exploration of Large Language Models by Optimal Exploitation AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.271222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.271222Z digest=sha256:616f071fb859426441b0fa67547d3c71c59bff6430cec4e725dfaedef9fc1024

Observation fda2ffc1-b5cb-4199-8ecb-6cc226be7279 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Disentangling Exploration of Large Language Models by Optimal Exploitation Reasoning with Language Model is Planning with World Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.286376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.286376Z digest=sha256:6b761338ab78ae613c9eecc92865394c18d464b555a5e1312e00996e47d056a4

Observation 3a91a203-1efd-4d38-918f-ef197a4e667c · outbound

This paper cites Benchmarking the Spectrum of Agent Capabilities.

Disentangling Exploration of Large Language Models by Optimal Exploitation Benchmarking the Spectrum of Agent Capabilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.281444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.281444Z digest=sha256:e4352579b8e19f8e01d7574926831c3dcbfc1c3d3a650a6a04e24906c11c7fd8

Observation 0778d96a-f16a-4e59-a6e0-40890b108f29 · outbound

This paper cites WESE: Weak Exploration to Strong Exploitation for LLM Agents.

Disentangling Exploration of Large Language Models by Optimal Exploitation WESE: Weak Exploration to Strong Exploitation for LLM Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.297123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.297123Z digest=sha256:b3c2272ca3bc0ba62cbd199d4e8ad2b02ee1deb8b7e90a5013176e36e32d8cb7

Observation d723556d-ac2a-40f0-a59a-023ca1f118f8 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Disentangling Exploration of Large Language Models by Optimal Exploitation Large Language Models Cannot Self-Correct Reasoning Yet

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.291538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.291538Z digest=sha256:1ad32bca11263ec1031a6246f19ab43ea81efc00fc1e4e01d0ee2f636cc7776b

Observation 1f2968bd-a1b4-4896-afcf-b6590903dec3 · outbound

This paper cites Mistral 7B.

Disentangling Exploration of Large Language Models by Optimal Exploitation Mistral 7B

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.309398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.309398Z digest=sha256:78637cfc6aca3350a3fa7733323a445de683c51497d85541fd71f18bb3d8253d

Observation 189c52e6-20cc-48a8-8141-865ae5157366 · outbound

This paper cites MathPrompter: Mathematical Reasoning using Large Language Models.

Disentangling Exploration of Large Language Models by Optimal Exploitation MathPrompter: Mathematical Reasoning using Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.303099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.303099Z digest=sha256:cb54815621db6fcf019ecab1505dea8a717dedd205ab8577a1b25074405b908b

Observation 26910b9d-19f2-4e9f-b9d3-109937cdb717 · outbound

This paper cites an unresolved cited work.

Disentangling Exploration of Large Language Models by Optimal Exploitation Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.320667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.320667Z digest=sha256:0881761af821ab6c1506ceed6c373e15a7a4f0735f99f77e230b247914ebe686

Observation 9396374a-002a-4bdc-971f-95e14ebec8a8 · outbound

This paper cites LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks.

Disentangling Exploration of Large Language Models by Optimal Exploitation LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.315070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.315070Z digest=sha256:09245124b88c7e4db3231cded78680eeb2100dbbbdf6f5823123feee12b5fddc

Observation 4ab7306b-2fd0-40d9-820a-bbd08cc26b17 · outbound

This paper cites Can large language models explore in-context?.

Disentangling Exploration of Large Language Models by Optimal Exploitation Can large language models explore in-context?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.331917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.331917Z digest=sha256:73b71d9f601834b75e23f7df6ff2da790052d96f7d452d8ded828f8e917dcfe1

Observation 956721f7-63ce-4a66-ad9f-ff2cb0d6ce3f · outbound

This paper cites an unresolved cited work.

Disentangling Exploration of Large Language Models by Optimal Exploitation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:18:59.685078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.326052Z digest=sha256:4cc1f399c4c24ded4565632afce2f91e1d3e9401d4b18fc1b926d706fd51fc12

Observation d8cf034a-8ea2-4753-86a0-2fb3f50e44b9 · outbound

This paper cites In-context Reinforcement Learning with Algorithm Distillation.

Disentangling Exploration of Large Language Models by Optimal Exploitation In-context Reinforcement Learning with Algorithm Distillation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.342404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.342404Z digest=sha256:51051e50c6310f654c87289d5e6405fc39affc39bb75b9ef417369a4ad7462da

Observation b8b17ff5-3e43-472c-b878-20add7c71e4f · outbound

This paper cites Küttler, N.

Disentangling Exploration of Large Language Models by Optimal Exploitation Küttler, N

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:18:59.668808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.337487Z digest=sha256:b47799d1863e3471604faeee99eb42c59d3745291a6e36d7670baedb1c63a52f

Observation 226d3299-5e91-47a9-81cf-b4170ee0e970 · outbound

This paper cites Continuous control with deep reinforcement learning.

Disentangling Exploration of Large Language Models by Optimal Exploitation Continuous control with deep reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.352768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.352768Z digest=sha256:02e9412b01654861ba908bc0a11585ef79f441174b43ca766af7f76c4f505b46

Observation 0c56a136-b35b-4cfb-80a5-29a92aa4bfd6 · outbound

This paper cites A Systematic Investigation of Commonsense Knowledge in Large Language Models.

Disentangling Exploration of Large Language Models by Optimal Exploitation A Systematic Investigation of Commonsense Knowledge in Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:18:59.127093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.347424Z digest=sha256:fadc5d94128fd19ed32b51c94fab52394787d39baf8a8c2818d972d0d02b54ec

Observation 65a75fa9-caf1-47ea-8989-419b5f6b589c · outbound

This paper cites Intelligent Go-Explore: Standing on the Shoulders of Giant Foundation Models.

Disentangling Exploration of Large Language Models by Optimal Exploitation Intelligent Go-Explore: Standing on the Shoulders of Giant Foundation Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.363672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.363672Z digest=sha256:f345d3694987572ece2eba477a695359d4e4884340a527872b7718ea2de32898

Observation d19a85b6-b431-4043-956d-92d26a775b71 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Disentangling Exploration of Large Language Models by Optimal Exploitation AgentBench: Evaluating LLMs as Agents

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.358394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.358394Z digest=sha256:dd908088814d66c48cf85ad40c81021f3858b43383bd2c7068af6af0c33f310e

Observation 58486615-9a0c-4c02-8d55-6663342b4a54 · outbound

This paper cites Accelerating exploration and representation learning with offline pre-training.

Disentangling Exploration of Large Language Models by Optimal Exploitation Accelerating exploration and representation learning with offline pre-training

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:18:59.039540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.373995Z digest=sha256:ce13700363310de1c26921d97543791294ea9c4dc798dcccf6249fd4d441ecdc

Observation 95902126-9dc4-42b6-a28b-47794fead0fe · outbound

This paper cites LASER: LLM Agent with State-Space Exploration for Web Navigation.

Disentangling Exploration of Large Language Models by Optimal Exploitation LASER: LLM Agent with State-Space Exploration for Web Navigation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.368973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.368973Z digest=sha256:69814a0d0440be6229be8b5519be63941f68e265bc01ad9cae86cabff418a6c8

Observation 5055f052-c402-4c9f-b2e6-6520b3950480 · outbound

This paper cites EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration.

Disentangling Exploration of Large Language Models by Optimal Exploitation EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.385039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.385039Z digest=sha256:50eda0698983c14d24b7ebd3397f242abeb10b7eb283bac8df72c5f1d4226815

Observation 31db0b48-286b-4995-9191-cdc66f0167ae · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Disentangling Exploration of Large Language Models by Optimal Exploitation Playing Atari with Deep Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.379372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.379372Z digest=sha256:8713123eecdaf199ef1a291290a8e8a1d3c7b9d4185e245522e5d368faaee7c4

Observation e8bb7481-8b1d-4b95-ad1c-9b0f54f0edf7 · outbound

This paper cites OpenAI o1-mini.

Disentangling Exploration of Large Language Models by Optimal Exploitation OpenAI o1-mini

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:18:59.652265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.397494Z digest=sha256:73a53499cd497aa0f7b534a36978ab70fc3ad7dd68e75c69f22f08b86edda95a

Observation 049b0a02-6458-4065-a8e0-0c367025c292 · outbound

This paper cites First-Explore, then Exploit: Meta-Learning to Solve Hard Exploration-Exploitation Trade-Offs.

Disentangling Exploration of Large Language Models by Optimal Exploitation First-Explore, then Exploit: Meta-Learning to Solve Hard Exploration-Exploitation Trade-Offs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.390579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.390579Z digest=sha256:947d302842d9e3325a5bf28ba9b6aad92116b505ccefe91507017f48b6734343

Observation 4b7b7de8-3265-4285-bca6-db07e16b2947 · outbound

This paper cites an unresolved cited work.

Disentangling Exploration of Large Language Models by Optimal Exploitation Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.408716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.408716Z digest=sha256:35c7282b13d0979adc0950993bfbca77a976ece42020e57ded0ac54eaa993b7a

Observation e1ded799-6ef0-4fd9-a49d-f408ec13da6a · outbound

This paper cites BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games.

Disentangling Exploration of Large Language Models by Optimal Exploitation BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.403115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.403115Z digest=sha256:f3b15e15811eca101d37331a07727e6777e891b70cd087ab9bc05f73b792e704

Observation b0e487ac-8e5b-4eed-b7a0-6aecfcbd5b7c · outbound

This paper cites Sequential Planning in Large Partially Observable Environments guided by LLMs.

Disentangling Exploration of Large Language Models by Optimal Exploitation Sequential Planning in Large Partially Observable Environments guided by LLMs

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:18:58.840397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.420211Z digest=sha256:9e5b2d737f9abf9cc1aed18d74a6592e261c8ddbcd4d28a3f41ab034a5172f23

Observation 64042dd7-7d05-4964-b7d1-31ec9a94ba0c · outbound

This paper cites Pathak, P.

Disentangling Exploration of Large Language Models by Optimal Exploitation Pathak, P

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:18:59.634863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.414108Z digest=sha256:94eb07036257cf35ce1d4c5fc82ad72990c0a80042c112d744c3787e355de21a

Observation e6abf4a5-82f5-48b3-b248-ead0e00d8691 · outbound

This paper cites LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations.

Disentangling Exploration of Large Language Models by Optimal Exploitation LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.431750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.431750Z digest=sha256:014477ff9b67e2ca13656fc496d9f27b88e276a0abcbe6fb892722d50a7eae57

Observation 75696448-fa7a-477f-9490-2dda829c9369 · outbound

This paper cites Doing Experiments and Revising Rules with Natural Language and Probabilistic Reasoning.

Disentangling Exploration of Large Language Models by Optimal Exploitation Doing Experiments and Revising Rules with Natural Language and Probabilistic Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.426352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.426352Z digest=sha256:f66be57ad06d8206ec35a6f3e42da3fcc9b6498bd9a6823390edcb32ec715426

Observation ff11d04f-a8a3-42ac-a124-2b9c94c822e6 · outbound

This paper cites Schäfer, F.

Disentangling Exploration of Large Language Models by Optimal Exploitation Schäfer, F

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:18:59.617844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.441956Z digest=sha256:89cd26d59e4bd0128b5dbb27342521c50f3c1d23bcfe5097dc7f854d9919c128

Observation a36c3d67-9610-49f4-8e4f-4fd5afe62610 · outbound

This paper cites MiniHack the Planet: A Sandbox for Open-Ended Reinforcement Learning Research.

Disentangling Exploration of Large Language Models by Optimal Exploitation MiniHack the Planet: A Sandbox for Open-Ended Reinforcement Learning Research

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.436794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.436794Z digest=sha256:350fcc441db4d67c58f7210bfafdd2989a276f251b9b9aadab9e2542c85c3c42

Observation 5acd3950-caa4-4486-abf0-d07a019efcbc · outbound

This paper cites Retrieval-Augmented Decision Transformer: External Memory for In-context RL.

Disentangling Exploration of Large Language Models by Optimal Exploitation Retrieval-Augmented Decision Transformer: External Memory for In-context RL

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.451394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.451394Z digest=sha256:f1fe06c5e114f82562e078f506e6aaef3a8e35207a826d1e8e6ed330756d74de

Observation d04cbe40-5f26-4459-b67c-d23c2cfe1eff · outbound

This paper cites Schmidhuber.

Disentangling Exploration of Large Language Models by Optimal Exploitation Schmidhuber

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:18:59.599640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.446609Z digest=sha256:aaaafcf03fd408fdf3b74850f63b402d19d9197d672dc6472106ce50206271d7

Observation 5e0ba4ff-9fd4-422e-9b81-d32da5d24a80 · outbound

This paper cites AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories.

Disentangling Exploration of Large Language Models by Optimal Exploitation AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.461847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.461847Z digest=sha256:3101a8777f03590c7b973350d918b5c2a6009c243b856433cd1b45d493e0de18

Observation 48f9fbcf-9989-4e78-b061-5d47869b4d0e · outbound

This paper cites Shinn, F.

Disentangling Exploration of Large Language Models by Optimal Exploitation Shinn, F

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:18:59.583090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.456530Z digest=sha256:bf398a8e08352d20941c0a71ff79f05afab6677a84caffd4139e0cf47d670369

Observation 3742c7b5-f72c-4136-8590-d617fffcaaa3 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Disentangling Exploration of Large Language Models by Optimal Exploitation Gemma 2: Improving Open Language Models at a Practical Size

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.472360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.472360Z digest=sha256:be1bcc1ad9e0916f4e1290e666d224e9fcc9b88e60be4fb7b0d53a66ab536278

Observation f5c28aab-d922-42ca-9e83-a7afcee2b9cb · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Disentangling Exploration of Large Language Models by Optimal Exploitation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.467337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.467337Z digest=sha256:065f969a5adc0f952ec0464be9f846f443a3190932b6fd5bb5ede06ce0bf90b0

Observation 94190651-513f-46fb-aa18-7561668ba091 · outbound

This paper cites an unresolved cited work.

Disentangling Exploration of Large Language Models by Optimal Exploitation Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:18:59.552068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.483576Z digest=sha256:a55624676e469c52478a0d12248186e47cfe48b16411b500d0f5c0e17839836b

Observation f3b4ab18-3592-4c6c-ad1d-f3ff7254ec16 · outbound

This paper cites Valmeekam, M.

Disentangling Exploration of Large Language Models by Optimal Exploitation Valmeekam, M

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:18:59.567823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.477916Z digest=sha256:b203a552f84718f662ecabef0cc0a7a5fc9a23e1edf53fc292663feaea1aa232

Observation b6e492cf-cafb-4b0a-a162-0297a0209807 · outbound

This paper cites SmartPlay: A Benchmark for LLMs as Intelligent Agents.

Disentangling Exploration of Large Language Models by Optimal Exploitation SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.494246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.494246Z digest=sha256:fa445484945e07023933724e01fbbe2f16aa401e93c48f976040d2266a653828

Observation 96f40b38-d6bf-4eba-a551-55f232d75e5b · outbound

This paper cites Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning.

Disentangling Exploration of Large Language Models by Optimal Exploitation Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:18:58.698860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.488803Z digest=sha256:16b615b95908e70e405575675eebaf0cfccf5af181febe97aab6310047815a4f

Observation 37cc3db9-8ca4-425b-a299-271f004f579c · outbound

This paper cites an unresolved cited work.

Disentangling Exploration of Large Language Models by Optimal Exploitation Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:18:59.536075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.508986Z digest=sha256:37807bc985f7000276adc43fcc4db846b048217ec70c7cc768e0cad2f6a983bc

Observation e68580c5-533a-4706-954a-bbd6c8dcf546 · outbound

This paper cites Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond.

Disentangling Exploration of Large Language Models by Optimal Exploitation Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.500173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.500173Z digest=sha256:f3797bf91ae71163d17d57e53d2b20f91a66383d78084b8889f1be3b69161ac5

Observation fc104b07-5fd9-468e-a9ac-1e7c96accf81 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Disentangling Exploration of Large Language Models by Optimal Exploitation Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.519583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.519583Z digest=sha256:7a597ed3aa9fb7321e64f54ed68665952f00e1e1ef8a710dbeeb9e4499d87428

Observation df5710bb-4147-4536-ad86-55341e7c88d9 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Disentangling Exploration of Large Language Models by Optimal Exploitation ReAct: Synergizing Reasoning and Acting in Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.513808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.513808Z digest=sha256:d89b3a252af0d8bdd95a5a3f4093ccc2dcaa3babc00b9b73d39b3ccf321fcb64

Observation c5dbbabd-c9be-4b86-b56c-801f111b2655 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Disentangling Exploration of Large Language Models by Optimal Exploitation WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.530774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.530774Z digest=sha256:9843ea3a1f8d2cb68e0b1935db76910e058d606710c2a53728f6a94de5a9acc3

Observation 8923fd79-695d-4e7a-8674-71bac52aaf57 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

Disentangling Exploration of Large Language Models by Optimal Exploitation AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T20:18:58.525337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:18:58.525337Z digest=sha256:6628119b25da47cc129d514f2c0a3fa11b33b75b65302b141ebf43a65ba30943

Pith citing papers

Observation fa4690cf-90ec-4c9a-9a49-e4cb0294a079 · inbound

Disentangling Exploration of Large Language Models by Optimal Exploitation cites this paper.

Disentangling Exploration of Large Language Models by Optimal Exploitation Disentangling Exploration of Large Language Models by Optimal Exploitation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:18:59.385875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:18:58.266458Z digest=sha256:3caff73d7005ff4531e9a34082f8ea4ce58f36d0bf713d7aef8c6f963d9448b3