Pith. sign in

Paper Citation Record · LEDGER

Large Language Model-Enhanced Multi-Armed Bandits

As of 18 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 4 inbound Pith citation observations for arXiv:2502.01118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01118 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:35:28.447005Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T23:51:50.319798Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:45.028347Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7ac073ff-9d07-4bd0-a484-31bec8094c44 · outbound

This paper cites Exploring Large Language Model based Intelligent Agents: Definitions, Methods, and Prospects.

Large Language Model-Enhanced Multi-Armed Bandits Exploring Large Language Model based Intelligent Agents: Definitions, Methods, and Prospects

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.399651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.399651Z digest=sha256:c018f5dae3ce180e5a22d8544f86d79062c87207f3375807cadd97a7396da167

Observation 6ae5361e-2716-490a-ac7a-b3e893acc484 · outbound

This paper cites In-context Exploration-Exploitation for Reinforcement Learning.

Large Language Model-Enhanced Multi-Armed Bandits In-context Exploration-Exploitation for Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.402832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.402832Z digest=sha256:d23b9604eb59278e92aa58f1f7229159d3a10c12477432908918e40a95fd1599

Observation 9ef088c9-3e97-4d6a-9054-6f8144f7cdf2 · outbound

This paper cites Efficient Exploration for LLMs.

Large Language Model-Enhanced Multi-Armed Bandits Efficient Exploration for LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.405762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.405762Z digest=sha256:434b8af785073b03a8cf8f73cf2ca4f294bf603e90f19962a3e4715764afb085

Observation b3b1b16b-31b2-4072-ac64-9ff18668e5ef · outbound

This paper cites Can large language models explore in-context?.

Large Language Model-Enhanced Multi-Armed Bandits Can large language models explore in-context?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.412559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.412559Z digest=sha256:db137727189f988226d4032bfc52eda33d7a86ac46ef57d9be00693ff4a95f7c

Observation 11fce7a2-85a6-4237-a1a5-511c898a5531 · outbound

This paper cites In-context Reinforcement Learning with Algorithm Distillation.

Large Language Model-Enhanced Multi-Armed Bandits In-context Reinforcement Learning with Algorithm Distillation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.414618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.414618Z digest=sha256:c32075165ec5ae0820f82b7250ff7fc2fffdd2427c0f6e14e1b84ec74619d20b

Observation 00a4f45e-0efa-4836-b504-e9d32e48dcd5 · outbound

This paper cites Feel-Good Thompson Sampling for Contextual Dueling Bandits.

Large Language Model-Enhanced Multi-Armed Bandits Feel-Good Thompson Sampling for Contextual Dueling Bandits

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.416638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.416638Z digest=sha256:5829e211e4e005e2c7a63b68b9ea90326ce788575e127a3a326fde04bbeeb1fd

Observation 13f79b30-3afd-46c1-8894-1d1359c4d537 · outbound

This paper cites Prompt Optimization with Human Feedback.

Large Language Model-Enhanced Multi-Armed Bandits Prompt Optimization with Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.419518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.419518Z digest=sha256:124f4f742e960599ca341d99a316f11e37969c949b71684a8a5573d64fb5f5dd

Observation e19a04e7-8d86-4d17-8419-f2bd12ee4e85 · outbound

This paper cites DeepSeek-V3 Technical Report.

Large Language Model-Enhanced Multi-Armed Bandits DeepSeek-V3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.422266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.422266Z digest=sha256:604b1685f8c4fa793b81a6741da980813480d253f6ed98479da7cac4f5466fa7

Observation d9640052-2ea4-4ec3-ac20-e542a866f8f9 · outbound

This paper cites GPT-4 Technical Report.

Large Language Model-Enhanced Multi-Armed Bandits GPT-4 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.427581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.427581Z digest=sha256:580a7f894b3078e278103cbcc612864cf00c3c2c71061dde843be8b105739692

Observation ef8e30af-bcb0-4fed-aa18-796d22b2f942 · outbound

This paper cites Wikilinks: A large-scale cross-document coreference cor- pus labeled via links to wikipedia.

Large Language Model-Enhanced Multi-Armed Bandits Wikilinks: A large-scale cross-document coreference cor- pus labeled via links to wikipedia

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:35:29.317407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T16:35:28.430202Z digest=sha256:19497c633a8f4f6323c556762793553a9789b11d9b5db5511e6bae3e3c802914

Observation 8e777ba9-046d-4891-8604-e21799a6a796 · outbound

This paper cites SmartPlay: A Benchmark for LLMs as Intelligent Agents.

Large Language Model-Enhanced Multi-Armed Bandits SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.435740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.435740Z digest=sha256:6141756765cc67926aabed28c8ef1fafcc1cf66ecdd214c5000a45d8d4ab2e97

Observation 81caac2d-8707-4acf-ac4f-aa8752b83d8c · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

Large Language Model-Enhanced Multi-Armed Bandits The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.438561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.438561Z digest=sha256:a9ccec00c7ea51ba9ed5813bbf2fbf3f6a68744ceaefadb862e25fec309e7da1

Observation 721aa724-5544-4a75-a6b6-cd19e0770c73 · outbound

This paper cites AgentGym: Evolving Large Language Model-based Agents across Diverse Environments.

Large Language Model-Enhanced Multi-Armed Bandits AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.441317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.441317Z digest=sha256:2f4f43b92b10b3e000c934dcc32c45dc4a5c5b3375929eadd3849fabb0750848

Observation da32c694-8115-4e1f-86e5-32a9ef833f86 · outbound

This paper cites Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents.

Large Language Model-Enhanced Multi-Armed Bandits Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.444193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.444193Z digest=sha256:3c0b7d3b1a20b1bfafdbb95bc3c3a7057e96f35dd7cd232e3f4cbc5ac0133a8b

Observation 4da292a3-76d7-43e0-83ae-7134394e7c3b · outbound

This paper cites Y ., McAleer, S., Fried, D., and Salakhutdinov, R.

Large Language Model-Enhanced Multi-Armed Bandits Y ., McAleer, S., Fried, D., and Salakhutdinov, R

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.410526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.410526Z digest=sha256:456948222f3d6e7c595572e8966b864f6f18f05da59ccef7537a7526709f88a1

Observation 4bfc6eeb-f2ca-4590-b52d-89aa034c0b8f · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

Large Language Model-Enhanced Multi-Armed Bandits LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.447005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.447005Z digest=sha256:650bfa497221253c4d5292987ef552ffbd3175803a46e6c6d93697b8a3eb4a4b

Observation 9f9b5af0-f6c9-4a12-b53c-52a27ba770c1 · outbound

This paper cites Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning.

Large Language Model-Enhanced Multi-Armed Bandits Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.392673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.392673Z digest=sha256:c05c6e9812a20e7259483c522c937a9d3330579de305f7d3ebc3e4969c283ce7

Observation 6fa7b74c-07aa-4730-814e-0f0e3f15ad62 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Large Language Model-Enhanced Multi-Armed Bandits Reasoning with Language Model is Planning with World Model

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.408114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.408114Z digest=sha256:aef718b486ee6e960a4c63b9b916f98f13f7460629a5a5e05a76dc94645334f0

Observation 60bec8d0-2692-46cf-91bb-7a64c7da2c42 · outbound

This paper cites Neural Dueling Bandits: Preference-Based Optimization with Human Feedback.

Large Language Model-Enhanced Multi-Armed Bandits Neural Dueling Bandits: Preference-Based Optimization with Human Feedback

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.432870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.432870Z digest=sha256:f52b444fda61302fb8ee7222c263d9de42580a2477e8ee840c4d69c1f6d4b7b0

Observation 14a80937-72c7-4b81-8c85-3dde3647387a · outbound

This paper cites P., Xie, Q., and Nowak, R.

Large Language Model-Enhanced Multi-Armed Bandits P., Xie, Q., and Nowak, R

Reference 2023

Resolution
verified exact
arxiv_id, observed 2026-08-09T16:35:28.672114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-09T16:35:28.425078Z digest=sha256:bd5e1d19cc27926a506e99559e6819541da33f4bdd196cf16f91e6c598079936

Observation 562233f7-1593-4733-bd40-d1f9d4fb1cd3 · outbound

This paper cites Efficient Sequential Decision Making with Large Language Models.

Large Language Model-Enhanced Multi-Armed Bandits Efficient Sequential Decision Making with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T16:35:28.396588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:35:28.396588Z digest=sha256:cc4659474a34094feec8816aa452729ededbe66127e93b5ca0545d95bc1af725

Pith citing papers

Observation 7e85714f-417c-4951-a850-a56ad7447dad · inbound

When Do We Need LLMs? A Diagnostic for Language-Driven Bandits cites this paper.

When Do We Need LLMs? A Diagnostic for Language-Driven Bandits Large Language Model-Enhanced Multi-Armed Bandits

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:25:51.129269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T18:32:12.396568Z digest=sha256:c8fa52c03223074809beb390b32c463eef1824afd0550b70832ff6dae4b8d0bb

Observation e24fc470-732f-419d-8a18-c5e79c456709 · inbound

Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits cites this paper.

Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits Large Language Model-Enhanced Multi-Armed Bandits

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:25:22.579124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T12:22:28.655926Z digest=sha256:f8bee2f4d423f3ab35f375bb5efae27100858642a381842cf98a2615e65cceb7

Observation 925a4193-bd9b-4dae-80d2-88e79e406c0b · inbound

GRIMIP: A General Framework for Instance-Specific Configuration of MIP Solvers Using LLMs cites this paper.

GRIMIP: A General Framework for Instance-Specific Configuration of MIP Solvers Using LLMs Large Language Model-Enhanced Multi-Armed Bandits

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:45.029930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T09:05:56.687400Z digest=sha256:80940e83708cb1ed667d17fe722612a2a30ccc6e624c6676ee8365856860e9a7

Observation 72a679c9-9451-436b-84b2-a0ae00b3e3ba · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Large Language Model-Enhanced Multi-Armed Bandits

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T23:51:50.319798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:51:50.319798Z digest=sha256:64b99e001b81e051fb721512e7f5aa3d9653efe4eac8e6874f417669e4915ea8