Pith. sign in

Paper Citation Record · LEDGER

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

As of 17 August 2026, this Paper Citation Record lists 100 of 122 outbound references and 4 inbound Pith citation observations for arXiv:2504.14520.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14520 v1

Coverage vector

measured 100 of 122 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:50:28.171568Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:50:30.476760Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T08:58:13.216074Z

Reference resolution

100 of 122 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e6c8538-f2b3-46ea-9dbb-bbd4b8ec88a8 · outbound

This paper cites Is creativity without intelligence possible? a necessary condition analysis,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Is creativity without intelligence possible? a necessary condition analysis,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.711553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.711553Z digest=sha256:6b916d58da156ba93594b98afb80f3da93fedcd50ca32662958d740b49634cbd

Observation d882530a-a2bd-4cb8-9c77-a800c9bc2dc6 · outbound

This paper cites Thinking LLMs: General Instruction Following with Thought Generation.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Thinking LLMs: General Instruction Following with Thought Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.716555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.716555Z digest=sha256:3857c9355005ad380a631f1b541ecbb7ad653f19d2bf14a71591c122e7bd7e85

Observation 096de597-32ce-4b2d-9bc7-38d428171b72 · outbound

This paper cites LLMs for Explainable AI: A Comprehensive Survey.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey LLMs for Explainable AI: A Comprehensive Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.721261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.721261Z digest=sha256:5c30580b738b4ce47e353a306a3da901c4b955361cc4ed3778465062ee596ea2

Observation 1dd879aa-9cc7-4df7-a203-09ba2a5c835b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Chain-of-thought prompting elicits reasoning in large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.726139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.726139Z digest=sha256:a0191341261aa057d782564128ff2f8c6f64c981e7547f66dc4c5d69cae4035f

Observation 42affbab-670e-46d7-b223-df2d468a7cde · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.730766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.730766Z digest=sha256:1d51f63fdad6ae3ec42f9131732597448dcd21bf480da18c35df8a0bfd4ae47f

Observation 71dc5145-b59d-4802-a989-688a1b3bb8a9 · outbound

This paper cites Retrieval Augmented Generation with Multi-Modal LLM Framework for Wireless Environments.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Retrieval Augmented Generation with Multi-Modal LLM Framework for Wireless Environments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.735660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.735660Z digest=sha256:2c24ff01488341ff4372d947de1145c68f2c7e2a8f0ed0887b2d9a33c8a9ba5a

Observation 154bcae4-8dbc-42c0-a481-b5108176eb91 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.741423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.741423Z digest=sha256:f3f4b9908f9134d9622734c03c584b6c8d6b73f79abf1fda62385eb1fd880872

Observation c9515db6-f872-4c55-b877-87afbf98d45d · outbound

This paper cites Assessing llms for high stakes applications,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Assessing llms for high stakes applications,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.745978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.745978Z digest=sha256:8fe2f52c6e0c5f58b7f07abd3d842ed479808d4e55b2893eaddab53f144539e9

Observation b65b79f8-6cb2-43d9-945e-cac4f7eaac68 · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.750368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.750368Z digest=sha256:b2acc7e2d60c48ece980c86c2e10c1c95fc254101b76a3d8ffc6d63c5434fcdf

Observation 44773207-9de3-461f-a6dc-b6efd21e2ec9 · outbound

This paper cites A review of methods for alleviating hallucination issues in large language models,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A review of methods for alleviating hallucination issues in large language models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.755117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.755117Z digest=sha256:222b3f923dad0552ac9949d6168caf98bc6ee7ea372d44208c875c0c9bc59109

Observation d7c67a13-6121-40a0-80a4-fe585826d8a8 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A Long Way to Go: Investigating Length Correlations in RLHF

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.759524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.759524Z digest=sha256:2e236fb1a54ad3b938de988e5d9c7a0928d86d803767e622fa655c834b04112d

Observation 1b9787b9-0b1f-4fe6-ab13-7722ba6d720b · outbound

This paper cites Human-level control through deep reinforcement learning,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Human-level control through deep reinforcement learning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.764165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.764165Z digest=sha256:d11f42d7fb95d679ec7a458a4f09e9a8863077174718c73957d6bea2c986a056

Observation 6c73aaa1-87b3-4b79-ab48-28cf63316f5a · outbound

This paper cites Continuous control with deep reinforcement learning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Continuous control with deep reinforcement learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.768476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.768476Z digest=sha256:2d7ccf38c323bcb1fabf1fc6906366e4cd5b3bfb766ac8cb7976e6a62c52ec2f

Observation 6e385f08-7479-43b8-a3d1-aed2a29b7ed7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.773074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.773074Z digest=sha256:9c3c48e5adf08b9c5c0a108e6326fe4e2d5b811dd88f7e125b86e4fe10aeb2af

Observation 427aca7f-f001-4028-8712-34c0f4cc7b30 · outbound

This paper cites Mixcl: Mixed contrastive learning for relation extraction,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Mixcl: Mixed contrastive learning for relation extraction,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.777510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.777510Z digest=sha256:700bdc08592a7dc9edb45c35d54e242fbb84077ed13aefcbc5f93bfd2be072d6

Observation 051a03b3-4b26-40c7-a85c-9253f5c82c3a · outbound

This paper cites Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.781720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.781720Z digest=sha256:e2115c6b10e0e2977f531ed7011620defc90ff2a0e2be70240e7a47631f58f5b

Observation abe910a9-7111-4836-a2d8-1d669142b0d8 · outbound

This paper cites Mitigating large language model hallucinations via autonomous knowledge graph- based retrofitting,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Mitigating large language model hallucinations via autonomous knowledge graph- based retrofitting,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.786475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.786475Z digest=sha256:e8b692ae86165e1bca8c2329d27b5e979214a09ef377363227d902771a4e6e25

Observation 6cfc1447-a83b-465b-bcfa-f0b303f4ff94 · outbound

This paper cites A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.791031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.791031Z digest=sha256:096cb7f9efff5db606aa75761909f8f351228b1831b8b590de98a8973029c6b9

Observation 18d80309-24b0-459c-9331-679058f8a2bd · outbound

This paper cites Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.795731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.795731Z digest=sha256:9f8446f5dd972b2ce5a9dfd1ec288414527d5dba41273bdb3dcdc9f85f84719b

Observation f1e05a96-e1a7-42d2-a8b3-725ffe62fe8a · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.800254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.800254Z digest=sha256:2938edd795e6444393374e35591ced658d8feec06d8e76248dd96fe97fe1b0a1

Observation b177ac6c-635b-4070-9139-0957ba51eeec · outbound

This paper cites SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.804684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.804684Z digest=sha256:0c69b2759ca2c2520ff62fca4e8a3372ca86c77753b68a5f8df76594ec636637

Observation 03780251-b86b-445b-b0fd-433f624b25a3 · outbound

This paper cites Chain-of-Verification Reduces Hallucination in Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Chain-of-Verification Reduces Hallucination in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.809408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.809408Z digest=sha256:41eaf31478290d3399812cb4c7a2c1988802bb7c60c301ae6faaf5f85a9edca9

Observation eb624d59-d1c8-426d-8410-550fd622c07e · outbound

This paper cites Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.814258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.814258Z digest=sha256:de2b99a8f9a9ef610690489b0f8954d4f0a33d60f0827916c21055090e608d7e

Observation fa6d9464-ecb5-4ab5-8836-a1589b16fbc5 · outbound

This paper cites Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.818881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.818881Z digest=sha256:79cb4e9b3004c12d5b0407c7cea62eed5e3ced30a1875c90f15f7b55ad746727

Observation 394cbf54-979d-498f-b796-80434bc5d06a · outbound

This paper cites KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.823504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.823504Z digest=sha256:43ee81a07c9597ed35fd452bf2359f4116b45e01bdd24b98d2cef9a0b6d05152

Observation f8a06022-c61e-49da-b9bb-3b649acac5ac · outbound

This paper cites Inference- time intervention: Eliciting truthful answers from a language model,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Inference- time intervention: Eliciting truthful answers from a language model,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.828277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.828277Z digest=sha256:859cf96bba14ecda8f0cec93a596bb1c0036407b224ceab495609c62d31f0d7e

Observation 66daf7b3-f725-4b40-81f1-135f56b87328 · outbound

This paper cites DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.832665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.832665Z digest=sha256:418fac4ab0b304da2d67e0b77286721e1fbff040d91934755ba0b7ac292a2363

Observation 11b83328-ce46-42b9-bb2b-30733779bead · outbound

This paper cites Self-refine: Iter- ative refinement with self-feedback,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Self-refine: Iter- ative refinement with self-feedback,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.837478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.837478Z digest=sha256:d7707fffa8bf96f2f166e4ba7e796d17963d311cac64c5ed0a9484be7224d97a

Observation 5e1110ee-275e-4c93-8f5c-966c12c56193 · outbound

This paper cites Can Large Language Models Really Improve by Self-critiquing Their Own Plans?.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.842026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.842026Z digest=sha256:84f5b4590030c5029e979b23968759a4dafa322f636b07390f298ae3575546de

Observation 49e02cfd-ec9c-4204-917e-a3e98acae603 · outbound

This paper cites Large language models lack essential metacognition for reliable medical reasoning,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Large language models lack essential metacognition for reliable medical reasoning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.846740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.846740Z digest=sha256:9f412e128f27df965e1f09d30776c5c339d834c10cdeb238837a42ebe54f6e4f

Observation 42b023a8-cbf4-4719-901a-76c742b31881 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Large Language Models Cannot Self-Correct Reasoning Yet

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.851212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.851212Z digest=sha256:d6c25759c8e31182ffe0ffd94fb779990b0d6e9bc95e3093ae6a3d30fa81a96e

Observation 1be6a8c3-4756-4598-ad56-93b346f8fc36 · outbound

This paper cites LLM-based Multi-Agent Reinforcement Learning: Current and Future Directions.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey LLM-based Multi-Agent Reinforcement Learning: Current and Future Directions

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.855907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.855907Z digest=sha256:40f6c00229d7c7af54af32ce295d4f23842961302bf2728cd4eb98084418d22a

Observation e3a1c42c-2c14-46da-a5a2-310886184be1 · outbound

This paper cites A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.860924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.860924Z digest=sha256:43e0d73815fd90e745129cce1b79d61d74f4a21fa76574c09faeda9a40f018d3

Observation 698330db-ad30-4eb7-8bbd-fab455087f69 · outbound

This paper cites Leveraging large language models for optimised coordination in textual multi-agent reinforcement learning,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Leveraging large language models for optimised coordination in textual multi-agent reinforcement learning,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.865663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.865663Z digest=sha256:2b0ea6739d290c9b35a8ae91a13d8daf126438a5d084ab89e76a601d56202352

Observation 364cf4fc-5cf4-4a74-8ac4-69dfc1e583aa · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.870208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.870208Z digest=sha256:fde5c960c28b80f4398f91a2114ed1dca3a41a0de37a2201b00935f95bc47394

Observation 8777bdbe-28ec-468b-8ee2-64ba41735139 · outbound

This paper cites Building Cooperative Embodied Agents Modularly with Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Building Cooperative Embodied Agents Modularly with Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.874878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.874878Z digest=sha256:b90faad4b06e8283aed569b9d5962e84e9503297dd228b2d9d04cd853981a076

Observation 1ee36bf9-6b7b-4bdc-bdfe-5c28d25f82b1 · outbound

This paper cites Smart-llm: Smart multi-agent robot task planning using large language models,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Smart-llm: Smart multi-agent robot task planning using large language models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.879546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.879546Z digest=sha256:c36bd01956128d62cf8b76feb12bb4aff97a15f689e81f5a59d6f90514922ece

Observation 1e6c3d32-4502-4e7e-bd49-c845961de962 · outbound

This paper cites Roco: Dialectic multi-robot col- laboration with large language models,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Roco: Dialectic multi-robot col- laboration with large language models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.883728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.883728Z digest=sha256:fcadde6dbb3764fe9c41372b19c7d33a83254f760aebc1cfdb69f4fa8ff0d306

Observation ffa46914-3744-4eb5-9038-cdd6a3550848 · outbound

This paper cites Co-NavGPT: Multi-Robot Cooperative Visual Semantic Navigation Using Vision Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Co-NavGPT: Multi-Robot Cooperative Visual Semantic Navigation Using Vision Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.888658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.888658Z digest=sha256:1d69f2a8c9b6082ccb44c2118139402fee8a4905d1030da84604a6dc99d96765

Observation 45e3ee58-f44c-4e3f-887b-0270a1e01346 · outbound

This paper cites Embodied LLM Agents Learn to Cooperate in Organized Teams.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Embodied LLM Agents Learn to Cooperate in Organized Teams

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.893387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.893387Z digest=sha256:322afc755cdbd33b2c76ec8a49335f176d433af6ba1ee8af7476bf08f9740e80

Observation 43151392-5bfa-46e7-8c12-d548975aae4d · outbound

This paper cites Multi-Agent Consensus Seeking via Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Multi-Agent Consensus Seeking via Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.898289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.898289Z digest=sha256:fd4f3064efd2f89d3e340567c94de6a2fe00d7a1b882706f084fa805f24ce9df

Observation ebe41552-9243-42fa-bf8e-cd80ba91705b · outbound

This paper cites Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:50:29.185967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:50:27.903048Z digest=sha256:b1c7ed9d8275ad62a92d99de735e9357ec4170ed695bcc5180b1f342f609ebe6

Observation 352db76e-7349-40df-9c3b-bbb078ccf2b1 · outbound

This paper cites ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.907723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.907723Z digest=sha256:fb945712715ffff8adfd1a1c9b320688d65bb57bafa90ca832b19b5acd7bc904

Observation e8b33906-5f13-4b0c-acb1-ef29da2f2776 · outbound

This paper cites Theory of Mind for Multi-Agent Collaboration via Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Theory of Mind for Multi-Agent Collaboration via Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.912446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.912446Z digest=sha256:9ea2a0b8558ba050f63615a89fa3f6e151d63508a078f95da543e36760c7fb63

Observation dd710fb3-b6b7-4f08-a328-aa2b344bf477 · outbound

This paper cites Buffer of thoughts: Thought-augmented reasoning with large language models,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Buffer of thoughts: Thought-augmented reasoning with large language models,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.917164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.917164Z digest=sha256:a18e1b8d1c3e5aeca506ab368f080deff9965e0b4d66c410efd8c2ada929a388

Observation f5bf7050-ed37-4906-9f77-cf7cb5e6d5a3 · outbound

This paper cites Meta Learning for Natural Language Processing: A Survey.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta Learning for Natural Language Processing: A Survey

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.921574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.921574Z digest=sha256:b3d64627e3b1fd30685cca611f0710d1e885cef28b6e12e7f09f61d5a36df1ee

Observation 7c7d0086-2289-4c4b-ab98-b24412aa91f7 · outbound

This paper cites Continual Learning for Large Language Models: A Survey.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Continual Learning for Large Language Models: A Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.926559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.926559Z digest=sha256:fd5f7f6dff1f260acf2d68073f17d040bbd6f5d1f4f0dc010feb74faa99ca09f

Observation 889eb9f5-b774-44aa-be5a-eee1a1c129a3 · outbound

This paper cites Continual Learning of Large Language Models: A Comprehensive Survey.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Continual Learning of Large Language Models: A Comprehensive Survey

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.931335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.931335Z digest=sha256:dd55971ac21f622dd811c8e78ca2ae6f3fd89ac3595c6881830cdf527e1e77f0

Observation 529f355d-e1ec-410b-8770-7601458b3a14 · outbound

This paper cites Towards Incremental Learning in Large Language Models: A Critical Review.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Towards Incremental Learning in Large Language Models: A Critical Review

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.936058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.936058Z digest=sha256:10aa60193b3527dd6ee27ccb3cea3ba71b8f30490deddaceaffffe87e0822642

Observation 97d70b5f-788b-4f6f-af2d-ea48749a2aee · outbound

This paper cites Learning from models beyond fine-tuning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Learning from models beyond fine-tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.940822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.940822Z digest=sha256:67d85b513a76ec3b99a95dd71cf2387888976efb6fb28fbcbdabcce1bb3bc747

Observation 5dc7f790-083b-4d13-a459-46dc0a0c4f84 · outbound

This paper cites Meta Reasoning for Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta Reasoning for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.945469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.945469Z digest=sha256:9f520a73c45c97295822688c8aa07dc30db87f6f49912f38aac0d4d11ee15c80

Observation 68c5b9f5-1cba-499a-b9c0-edb8a2a9bacf · outbound

This paper cites Reasoning with large language models, a survey,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Reasoning with large language models, a survey,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.950019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.950019Z digest=sha256:c3842516558ae1684da0f15a164f7ccdfabeb3b3f60a762f8dde88477c125441

Observation 8ef3dec7-8bee-4817-900b-2b2c8f74c115 · outbound

This paper cites Gizaml: A collaborative meta-learning based framework using llm for automated time-series forecasting.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Gizaml: A collaborative meta-learning based framework using llm for automated time-series forecasting

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.954474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.954474Z digest=sha256:92e19b476332e2c8be31a4bba7c80034428eb60b33823fea5c7c1e7baa8c9e59

Observation 034a4aa5-03f0-460b-9991-e56b5462e6f8 · outbound

This paper cites Towards lifelong learning of large language models: A survey,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Towards lifelong learning of large language models: A survey,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.959408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.959408Z digest=sha256:859c091373e5a0a2bf455104e67efb627cecc957009c1c62cdd3e16732139594

Observation 35511f4b-50f6-4ed2-b0f4-fea4370ae594 · outbound

This paper cites Thinking Machines: A Survey of LLM based Reasoning Strategies.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Thinking Machines: A Survey of LLM based Reasoning Strategies

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.964075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.964075Z digest=sha256:bcd3b087184a4e7cce0917e94229b6ddcaa3371e4383e90b5f3796aa7af89ef8

Observation 7cb7659b-e411-4e9a-a034-897ef2f3b2a9 · outbound

This paper cites MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.968815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.968815Z digest=sha256:6b45ee717609a295055796c93c0e18f15fd9b8460928286faeb4eecd2b1dc0e3

Observation 46850007-b9cd-4d20-b553-92a5208124ac · outbound

This paper cites Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.973380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.973380Z digest=sha256:bcbc1a36adc539d85287093d2e849c15ab91dd7f7fb89f6bc7bec66d6d7acd4b

Observation ebcbbd79-9a65-47e0-9a1f-9bff521d5042 · outbound

This paper cites Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:50:28.936737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T11:50:27.977789Z digest=sha256:2757df701dfeadb8fcd51717776b9301d8adad1b6fba3f34c851785cdb51bb9d

Observation 431dabcf-b987-43c7-a8e3-e51b2555374d · outbound

This paper cites Malalgoqa: Pedagogical evaluation of counterfactual reasoning in large language models and implications for ai in education,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Malalgoqa: Pedagogical evaluation of counterfactual reasoning in large language models and implications for ai in education,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.982603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.982603Z digest=sha256:b70832a72999d97fb3eacf76e18fda25f4acc353d8f6d03c24fa892d9e8f4a1a

Observation 377461d5-1bdc-4e12-9799-a738d93b855c · outbound

This paper cites METAL: Towards Multilingual Meta-Evaluation.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey METAL: Towards Multilingual Meta-Evaluation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.987474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.987474Z digest=sha256:d1c1a2e1e1a65dbddcdfaad3b00aa5ac7fbf6397101b85b3fda4405cab485744

Observation bd47ff9c-9ebf-4db1-a278-a2dc24002790 · outbound

This paper cites MetaICL: Learning to Learn In Context.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey MetaICL: Learning to Learn In Context

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.992183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.992183Z digest=sha256:0e38ecf594f250b9fed305d135a2d17bfa9c58bb362dd7c53112c762f4aa02c9

Observation 4e4b7159-712f-4537-ba66-a7ee4299d250 · outbound

This paper cites Towards metacognitive clinical reasoning: Benchmarking md-pie against state-of-the-art llms in medical decision-making,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Towards metacognitive clinical reasoning: Benchmarking md-pie against state-of-the-art llms in medical decision-making,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:27.997065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:27.997065Z digest=sha256:d75b67b83516ab7e15ea2b500d9a021a2bbb47d271d67e15bb2641130e7ce9db

Observation dc0c558f-ff7e-4608-b522-ec34abe6f280 · outbound

This paper cites Do Large Language Models Know What They Don't Know?.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Do Large Language Models Know What They Don't Know?

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.001877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.001877Z digest=sha256:6e458091d7e250855fafe0e38a9182cb2919a6dd53a88a0636827640c32d6f44

Observation d9ae61b9-778e-49eb-9f06-cb89bc74c7c0 · outbound

This paper cites Language grounded multi- agent reinforcement learning with human-interpretable communica- tion,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Language grounded multi- agent reinforcement learning with human-interpretable communica- tion,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.006674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.006674Z digest=sha256:f8a38eea32724d902ecb96d9ec547dfbfaffc812ffb3a6749d4722a8fea7ed2a

Observation 67fcd80a-e994-4467-af7e-ddc518990c26 · outbound

This paper cites Hypothetical Minds: Scaffolding Theory of Mind for Multi-Agent Tasks with Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Hypothetical Minds: Scaffolding Theory of Mind for Multi-Agent Tasks with Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.011118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.011118Z digest=sha256:577a5880ef1b0913306ec49d10a4a47de1881bc0942403c295847a383af54a8d

Observation ebbc836d-4e0a-40dc-aacb-1fc3ce636dcd · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Distilling the Knowledge in a Neural Network

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.015604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.015604Z digest=sha256:ab9d7e7334ff71ff5c295a927c327736656e4488dcdb0b94aac2a4d63c2df0d4

Observation 92a9ccc0-e665-4136-b76a-460f8febf5ee · outbound

This paper cites Knowledge distillation: A survey,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Knowledge distillation: A survey,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.020070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.020070Z digest=sha256:a9e934a003805026d11e35ded411eaa90e9ad2d0ca7129121d9bf990ccd8b524

Observation e13eeaa1-c6fa-4272-84fa-c432c205caa2 · outbound

This paper cites Large Language Models have Intrinsic Self-Correction Ability.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Large Language Models have Intrinsic Self-Correction Ability

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.024588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.024588Z digest=sha256:65f8cb00446db00d20a1212208c1aff26a16e829471f0088c516bebd2106c140

Observation 8e8ead3d-6cf2-4ef4-9157-ea712b0ae964 · outbound

This paper cites Can Rationalization Improve Robustness?.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Can Rationalization Improve Robustness?

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.029078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.029078Z digest=sha256:0ec919d136ca678ebf081908eaa35a993642fb90738492ebfdd426e40e925acd

Observation 41c59bf3-1b97-4413-8d7a-e4df063ee9e4 · outbound

This paper cites Measuring Compositionality in Representation Learning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Measuring Compositionality in Representation Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.033705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.033705Z digest=sha256:4f8478b3f2c839a8ef59b0de38919e121a6a9625081cb419fb8c7297a4bbb8f3

Observation 6affd4e1-55c9-4651-bf20-c75c71cdc5e3 · outbound

This paper cites Multi-agent Reinforcement Learning in Sequential Social Dilemmas.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Multi-agent Reinforcement Learning in Sequential Social Dilemmas

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.038753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.038753Z digest=sha256:0cd9fa8b17354b30161c005a118722a1b8aeee1b51b3e012432021f5aa513bba

Observation 4045999f-60d3-410e-8cd0-9c3db45adbaf · outbound

This paper cites Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.043495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.043495Z digest=sha256:3c835fcdd3510e2b4c9cdfce6d89802a993ece6e71350137ecfcc97f46850963

Observation 27c89c8d-3764-4f1c-aae9-f40bda59ecfa · outbound

This paper cites AI safety via debate.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey AI safety via debate

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.048060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.048060Z digest=sha256:e1e8790b5f55fea1aa357e23dc3f2af56ecadb50029a53020cf4d8963de2f2b1

Observation 5941701a-ef45-4456-82a8-9b9cbb91d422 · outbound

This paper cites Adversarial training for high-stakes reliability,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Adversarial training for high-stakes reliability,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.052767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.052767Z digest=sha256:0ae003c0c649e2fb4f4f85c36041f7347997622e1a2ef8cd9e79548c01841300

Observation 6612d0f1-855f-48ff-8898-5be4bf4fe428 · outbound

This paper cites AD-AutoGPT: An Autonomous GPT for Alzheimer's Disease Infodemiology.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey AD-AutoGPT: An Autonomous GPT for Alzheimer's Disease Infodemiology

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.057392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.057392Z digest=sha256:1361d609c846a92518d3597b5710ec7b145ddff6d68eb7034ba27cebd20f1515

Observation d9024881-cdfb-4788-91f7-2dbf9b92107c · outbound

This paper cites Deep reinforcement learning from human preferences,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Deep reinforcement learning from human preferences,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.062051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.062051Z digest=sha256:6822af19604f37a7665e222e5f0c5af11d3a4d143dcacafeff25fb95269a4bed

Observation 2b65ccd2-3f01-410c-a3de-2de81a16e07d · outbound

This paper cites Training language models to follow instructions with human feedback,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Training language models to follow instructions with human feedback,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.066324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.066324Z digest=sha256:af3265f5b354fba4902bf65a24b927f4dde40b88f584b4415f5cafff19948faf

Observation 961f0250-0a70-4bf2-b01b-fdfd2a0a40fa · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Training Language Models to Self-Correct via Reinforcement Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.070616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.070616Z digest=sha256:1e4e31c5d7119062516552730035aa20b23f47862a89d804f70717f82e405409

Observation ce8f0fa8-a1ad-4dbf-8c38-a30ac0bcf42d · outbound

This paper cites A Tutorial on Meta-Reinforcement Learning.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey A Tutorial on Meta-Reinforcement Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.075344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.075344Z digest=sha256:8a31ba7f49b88a5656ca1c40bd32027d81c1f39266f687b7fc167ff9bb13095a

Observation 85be3b41-3738-43e8-9fcb-d5c5900d92ad · outbound

This paper cites Eureka: Human-Level Reward Design via Coding Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Eureka: Human-Level Reward Design via Coding Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.079809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.079809Z digest=sha256:ff5c2f3e800888fca5932fdef48490ac88d4e496c82ec66b4ecdd77b794fe206

Observation 7e57c22b-0edb-45fc-b1af-92cf449c920b · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.084177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.084177Z digest=sha256:108d106d4c523c1b3cd4399e12ffe115a2395ed5d0327d035f5bd71c59c920ca

Observation 93397ad0-a6c0-4119-b5ef-f58d4b6bbc0c · outbound

This paper cites Learning to summarize with human feedback,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Learning to summarize with human feedback,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.088784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.088784Z digest=sha256:9f20e94b78a6fb8b0e2a9aae2fc60c5119a08bd73e598c1b8a647ff05cdf408b

Observation 0a0131c6-e27b-4d0c-ba80-8fb8cc73294d · outbound

This paper cites Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.092823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.092823Z digest=sha256:1ab946b8532df536832754695fb527873374c6cbbea28f7c54bd3df06ae91737

Observation c1f33125-130b-4524-aa1c-2a2cadc76e4a · outbound

This paper cites Curiosity-Driven Reinforcement Learning from Human Feedback.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Curiosity-Driven Reinforcement Learning from Human Feedback

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.097413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.097413Z digest=sha256:4ac4cfa7895e36d994d6c8cc1938155a7dc9fd376ab55110a848fdde1e516858

Observation 4e39e6a5-dffc-4217-b1c4-ae0d3fd4e97a · outbound

This paper cites Online intrinsic rewards for decision making agents from large language model feedback,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Online intrinsic rewards for decision making agents from large language model feedback,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.101951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.101951Z digest=sha256:9f5a61c0f374a2a193bec5ab99018b94900bfee528f6120c5532d8b01a1ce37b

Observation 45c41508-2731-41d1-9af1-c041ccf42002 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.106527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.106527Z digest=sha256:80e69538b5b82c23d73c2edf47844a3259ef488a87e2488003e764f8cf19d816

Observation 4290a4a5-6c3b-4d78-b8e2-93a8058b923c · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Grandmaster level in starcraft ii using multi-agent reinforcement learning,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.111357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.111357Z digest=sha256:a222297adb94d4829b7c0fdebdc12078790f1b2575dbe49a7aef3dcbd58a5463

Observation 961f9518-3979-41e8-95cb-bd8151638423 · outbound

This paper cites Human-level play in the game of diplomacy by combining language models with strategic reasoning,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Human-level play in the game of diplomacy by combining language models with strategic reasoning,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.115812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.115812Z digest=sha256:4753262ac0e34851e2585e94f2d4502ea9f4842bfb8c572ed13f489bfd95075d

Observation 77bbbd09-8f67-474a-8f63-6f5ed77d072c · outbound

This paper cites Shall We Team Up: Exploring Spontaneous Cooperation of Competing LLM Agents.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Shall We Team Up: Exploring Spontaneous Cooperation of Competing LLM Agents

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.120233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.120233Z digest=sha256:b1cee490b06d50e13ae1c68f735c708cc15d704e3308746412be6557d3c83df5

Observation d6b594b8-4cf9-4770-9318-2f3ae99696b8 · outbound

This paper cites Self- playing adversarial language game enhances llm reasoning,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Self- playing adversarial language game enhances llm reasoning,

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.124703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.124703Z digest=sha256:0e3b1a4e2603c4a221a9c1832dd18e138aa85d39e2951d11f67bf0ba9944d626

Observation b50ec277-b147-423a-8e93-544c3680f91e · outbound

This paper cites Logicattack: Adversarial attacks for evaluating logical consistency of natural language inference,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Logicattack: Adversarial attacks for evaluating logical consistency of natural language inference,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.129150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.129150Z digest=sha256:372b05bfdc80e08afe159402b69ad254062cdab8dfe94d76367a204349d89ef7

Observation 4d0e1425-408a-4eee-8266-419855a42db0 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Model-agnostic meta-learning for fast adaptation of deep networks,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.134129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.134129Z digest=sha256:6044368b0f0006f4c8e418a47b8307dcc5ef80d2b8cbf5ee058214c272040ccf

Observation 54c4ffe3-a2f7-474c-ba71-1c2918a801bc · outbound

This paper cites Meta-learning for Few-shot Natural Language Processing: A Survey.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta-learning for Few-shot Natural Language Processing: A Survey

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.138491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.138491Z digest=sha256:dee23380a13e68b042e97472e4509c88f24eb706323e8330727190278307952c

Observation b4a7406c-bca5-4b9d-b880-910dfc948c4c · outbound

This paper cites Meta In-Context Learning Makes Large Language Models Better Zero and Few-Shot Relation Extractors.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Meta In-Context Learning Makes Large Language Models Better Zero and Few-Shot Relation Extractors

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.143364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.143364Z digest=sha256:d57c235164499301e0f5b7b036921a7d5cc2b8407fb309fa439bcc546eb7baa6

Observation df4e08ac-3fce-4d24-877c-f38553618c9b · outbound

This paper cites Improving Consistency in Large Language Models through Chain of Guidance.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Improving Consistency in Large Language Models through Chain of Guidance

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.148408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.148408Z digest=sha256:2bb3c8661986378e0eaed89e0570a21ae6d6537dfaed0147e7eae12a21152ca9

Observation 3633bd04-b82f-4a56-8a7a-09f3aeff3b4f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Training Verifiers to Solve Math Word Problems

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.153101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.153101Z digest=sha256:59b7dadaf874f94f6f6c607ba11e46f375e69883fea884980582353e7416ae87

Observation f1d73e9d-158b-46df-adae-97258db6d2ae · outbound

This paper cites Red teaming language models for contradictory dialogues,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Red teaming language models for contradictory dialogues,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.158013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.158013Z digest=sha256:c98be838affaa94d26e3cbe4fa390cf6431d8c790a5f66d58b073f982e3e75e5

Observation 104758c2-e148-4106-b2f1-69ceb8b023b1 · outbound

This paper cites OpenAI o1 System Card.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey OpenAI o1 System Card

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.162336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.162336Z digest=sha256:4532911ff3aa7f995c3ec78735d97bbd780a68d924e17ee322c693d27051d20e

Observation b0de2eb8-aa5b-43e9-b20b-4346b35a212a · outbound

This paper cites Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies,.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.167260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.167260Z digest=sha256:11c3a2f9c869228e4aedf67c9eae00c440a2a3273c55bedbbd2f276f672d5785

Observation 219f7fc9-5318-4c73-bf61-8388a04b39f8 · outbound

This paper cites Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate.

Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:28.171568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:28.171568Z digest=sha256:f77633ac96738e5146aa1963e342c34a828eea6e602cf9c43bc2c697d6c05381

Pith citing papers

Observation d795cdaf-31c7-4e31-b172-fad1c89479b0 · inbound

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models cites this paper.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.476760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.476760Z digest=sha256:136f569b12d56504c288007d4f2c4a4ae8fc3b48bd5fcd67e6d731b52033d65d

Observation 594d600c-b10e-46ab-b09a-d5a8219ea514 · inbound

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration cites this paper.

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:58:13.218271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T08:53:03.420926Z digest=sha256:319dad61c8e2693f84db87b948ccff4ac1983235f321906ad0bd1983682c7dd3

Observation 0b30f91a-1da6-43b0-b8d6-2a8131d40932 · inbound

Human Cognition in Machines: A Unified Perspective of World Models cites this paper.

Human Cognition in Machines: A Unified Perspective of World Models Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:12:25.995829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T08:12:15.663761Z digest=sha256:3e9596084b805bcacfa4f1a7b28c34953ea51a7e456655ea7a1248560546eed5

Observation c3a4fb96-ec94-4311-8ce7-2106133bb173 · inbound

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs cites this paper.

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.570592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T00:51:37.506096Z digest=sha256:c18d9782f4f85a1bf358f7a228be26f5faa0f68a63df45efd22f4381663ca876