Pith. sign in

Paper Citation Record · LEDGER

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One

As of 10 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2505.15306.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15306 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:25:00.054246Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7ad5cac9-df3a-44a8-bfda-9bd22a7c4f8b · outbound

This paper cites Reinforcement learning: An introduction.A Bradford Book, 2018.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement learning: An introduction.A Bradford Book, 2018

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.681461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:56.234369Z digest=sha256:56e1209713273bd8896ac7e4adbd0de363edf2cb88799d55f9e3640dd749a21c

Observation b65d7d72-dd7a-4714-becf-e4a3fe7b93d8 · outbound

This paper cites An introduction to deep reinforcement learning.Foundations and Trends® in Machine Learning, 11(3-4):219–354, 2018.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One An introduction to deep reinforcement learning.Foundations and Trends® in Machine Learning, 11(3-4):219–354, 2018

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.660069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:56.288962Z digest=sha256:ebee9bed0abca812bc92263b803729e1ffda8fdb510a92a604f7131f3afb9fb3

Observation e379d43f-4d11-4265-8313-801cd0839fad · outbound

This paper cites Mastering the game of go without human knowledge.nature, 550(7676):354–359, 2017.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Mastering the game of go without human knowledge.nature, 550(7676):354–359, 2017

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:56.346937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:56.346937Z digest=sha256:192e943cc0c35801bcb7a47a1fb5dfe36984a9cd4cb38d4983c404faeab87b68

Observation 2db00582-0d4f-4bac-856b-d4dcff09eaab · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:56.431165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:56.431165Z digest=sha256:c7f9ebd3ee63088fbea4b57aac4e6055da67f50f34ffb1063b85482c89819a21

Observation b5c85c06-0d0d-4240-a7f2-d61777df09e3 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Dota 2 with Large Scale Deep Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:56.541701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:56.541701Z digest=sha256:89e42e68d969f7b1d5325472f357a196c7b6b9376a9a76c2f0c62263f86e07b0

Observation eb94c72c-b50b-45d6-98d8-4ab6433b0bf5 · outbound

This paper cites Mastering atari games with limited data.Advances in neural information processing systems, 34:25476–25488, 2021.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Mastering atari games with limited data.Advances in neural information processing systems, 34:25476–25488, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.606430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:56.598174Z digest=sha256:4dadf0e5f68359d53a2af9a85afd1ab366c960effebe6761acfee8c7fd66936f

Observation 78b34cf3-9b9d-4b81-85d1-48c022f7a2b8 · outbound

This paper cites A graph placement methodology for fast chip design.Nature, 594(7862):207–212, 2021.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A graph placement methodology for fast chip design.Nature, 594(7862):207–212, 2021

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.584564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:56.693310Z digest=sha256:3cbae42203c72c7f44c072f39cedc1fa1015496b062aeea1fc8ec2b80c0468d4

Observation d14d986c-aa0b-4a88-a7b6-4179af0f8d06 · outbound

This paper cites Hierarchical reinforcement learning for scarce medical resource allocation with imperfect information.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Hierarchical reinforcement learning for scarce medical resource allocation with imperfect information

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.558171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:56.767959Z digest=sha256:12f06e982ab57a01157789f96f2f862afb6a954b0058691db4c65a4ca7fa22ce

Observation 99e60500-87cd-4591-81ca-4282dfb822a5 · outbound

This paper cites Reinforcement learning enhances the experts: Large-scale covid-19 vaccine allocation with multi-factor contact network.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement learning enhances the experts: Large-scale covid-19 vaccine allocation with multi-factor contact network

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.538975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:56.855259Z digest=sha256:e9537d0a1d0f82fcd19881ecf64bcfaa74b8196db428dc92a96e1653991d20ea

Observation f0e78dcd-1a14-4ab8-9db4-436492db1d43 · outbound

This paper cites Gat-mf: Graph attention mean field for very large scale multi-agent reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Gat-mf: Graph attention mean field for very large scale multi-agent reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.520522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:56.945786Z digest=sha256:63e6446258997956f75de6f258d71cbd5da5d7353c959ead216fce254abcb266

Observation 1638ec47-1810-47be-9ac0-132adbf5f3bb · outbound

This paper cites Spatial planning of urban communities via deep reinforcement learning.Nature Computational Science, 3(9):748– 762, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Spatial planning of urban communities via deep reinforcement learning.Nature Computational Science, 3(9):748– 762, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.047385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.047385Z digest=sha256:27c8a5c64d85c0a14a745c384afab30964c2be3f34836aae5d8df66c72125dfb

Observation 15e054f8-c016-49d9-a301-117d0711c46d · outbound

This paper cites A survey of machine learning for urban decision making: Applications in planning, transportation, and healthcare.ACM Computing Surveys, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey of machine learning for urban decision making: Applications in planning, transportation, and healthcare.ACM Computing Surveys, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.491251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:57.099196Z digest=sha256:3a97aa4b30db65f27a8abac12c37a4a2ce0572341da717e30f2d75d5eb4bc8b4

Observation 39339997-605b-4318-bbbf-bfec067c9de8 · outbound

This paper cites Dyps: Dynamic parameter sharing in multi-agent reinforcement learning for spatio-temporal resource allocation.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Dyps: Dynamic parameter sharing in multi-agent reinforcement learning for spatio-temporal resource allocation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.471478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:57.212195Z digest=sha256:60804ecf5ca2d479c3aec12bcd145e952b7f13f3ac0fbeb701ca362f7909518c

Observation 3652024f-61f0-4ba1-b24b-bf63788475f8 · outbound

This paper cites Coopride: Cooperate all grids in city-scale ride-hailing dispatching with multi-agent reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Coopride: Cooperate all grids in city-scale ride-hailing dispatching with multi-agent reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.453447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:57.301068Z digest=sha256:e89567c0fc8c444ed3d813e66efdd5a01450b1ebaa970f571e258b29059e952f

Observation a86bb074-2a9b-47c2-90ba-95720d7c2bec · outbound

This paper cites A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.402374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.402374Z digest=sha256:4993698f0c03854cf3d36b0a7a9458c2cce1a2694a52b52c4fa15fd648f30896

Observation 52045072-bd79-4cbd-a6a1-4dafc2302c9e · outbound

This paper cites How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.516895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.516895Z digest=sha256:85d1d8ca579f380da83f12f11d3e97eab70893982db50b5fcdb1d4700f192517

Observation 73c5c68b-f8d5-41cd-840a-450a3b80d005 · outbound

This paper cites Ensemble deep learning: A review.Engineering Applications of Artificial Intelligence, 115:105151, 2022.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble deep learning: A review.Engineering Applications of Artificial Intelligence, 115:105151, 2022

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.437638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:57.614737Z digest=sha256:9cb34be4c554867f9b5d7b81c33362aac66693ee5d0fcf156e1115a85079494e

Observation 85e2220f-f902-4ad9-a9af-b91f988f89e4 · outbound

This paper cites A survey on ensemble learning.Frontiers of Computer Science, 14:241–258, 2020.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey on ensemble learning.Frontiers of Computer Science, 14:241–258, 2020

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.421082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:57.701960Z digest=sha256:e9b271dad70559348572cfb3d4199a826996fab8992111f178484af344f17bf8

Observation 422ef0c2-f114-4985-bf1d-8e18ca44d06a · outbound

This paper cites Ensemble reinforcement learning: A survey.Applied Soft Computing, page 110975, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble reinforcement learning: A survey.Applied Soft Computing, page 110975, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.405209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:57.775550Z digest=sha256:6e3f396c808fa74ead72c357af21fbcfdd5e7bdac0361251ff8d41d7270a3ce3

Observation 69160df6-5acd-4af3-b01f-1599a827a21e · outbound

This paper cites Neural network ensembles in reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Neural network ensembles in reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.389623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:57.856542Z digest=sha256:8c58f207dbfa44a092b59bc612ce4b956d7f187575e9cf1e5125e3b627b70c91

Observation 499cc08f-ef3e-4042-b788-fac959861c37 · outbound

This paper cites Reinforcement Learning based dynamic weighing of Ensemble Models for Time Series Forecasting.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement Learning based dynamic weighing of Ensemble Models for Time Series Forecasting

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.963590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.963590Z digest=sha256:babc50fe4d0e04d75c5a152b7f4b00919503d3a6738a4f601b9dff767dfabdc4

Observation f5a597ba-0a2c-4a3b-b84f-36281b50f99a · outbound

This paper cites Ensemble algorithms in reinforcement learning.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 38(4):930–936, 2008.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble algorithms in reinforcement learning.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 38(4):930–936, 2008

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.374081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:58.083139Z digest=sha256:ed8a3d44e0f91407f61e334885e0720dd4fb4bbbcc3f6bd8fe68350267342899

Observation cc9b4d82-1585-4b67-a68d-64dd57681529 · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.Journal of Artificial Intelligence Research, 47:253–279, 2013.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One The arcade learning environment: An evaluation platform for general agents.Journal of Artificial Intelligence Research, 47:253–279, 2013

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.358572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:58.191168Z digest=sha256:92c61cb7c557376f2b3ced81f579c67a7337570912bdd601cff2eb39c1df1d62

Observation 9e8f1398-449e-4f9f-a24d-73586e8337f3 · outbound

This paper cites Model-Based Reinforcement Learning for Atari.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Model-Based Reinforcement Learning for Atari

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.285479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.285479Z digest=sha256:e718818c36c3d69c054e12fa6f77ba1dc2fc840738aa7664c2d93cf0bf061051

Observation 80b1eed5-579f-40ec-9f56-ebcb17ade291 · outbound

This paper cites an unresolved cited work.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.360163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.360163Z digest=sha256:295bcd6f783fa35998e7f39e8bd0687b97bdf32e08e40dc0f0e46223ad2734e6

Observation e4ac7fd4-1ae1-47eb-bc30-7a4d805c4141 · outbound

This paper cites A survey of gpt-3 family large language models including chatgpt and gpt-4.Natural Language Processing Journal, page 100048, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey of gpt-3 family large language models including chatgpt and gpt-4.Natural Language Processing Journal, page 100048, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.330048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:58.423192Z digest=sha256:8fa91fb1d9f3c931ce02e37aacb1b088bf94759026f1714fc1c32ddde43690ed

Observation 2cec5757-1406-40ba-8e63-a0a23a7a4af5 · outbound

This paper cites GPT-4 Technical Report.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.507475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.507475Z digest=sha256:bf21ea368fda5f4e7b856747aa4e9703cafad16b36f738cfcbbcf2eb19d5814b

Observation 9ed98770-0f4d-4d91-a149-51d27d2c74ac · outbound

This paper cites Evaluation of openai o1: Opportunities and challenges of agi.arXiv preprint arXiv:2409.18486, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Evaluation of openai o1: Opportunities and challenges of agi.arXiv preprint arXiv:2409.18486, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.596746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.596746Z digest=sha256:70fb3f4a4904c74bdf8f85de96f6d9c964751a41a5596ca5d68069ebaadcef63

Observation 0bd3492f-c5fb-4285-91ba-dc3f233a93e9 · outbound

This paper cites Early access for safety testing.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Early access for safety testing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.719213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.719213Z digest=sha256:a2f71825221af28334ed12b8c85a1b35d3760d1661a1892b8abd55a8250d42cd

Observation 7e419586-5f54-4b2f-97a6-d767d7a28875 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.773436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.773436Z digest=sha256:8340c9c1971a6c375754ab7b3d10c3a71cb50f1f086c04bcb2dba8ebc549ced6

Observation 81c4a7cd-d620-47d9-aa37-efa2acf441b9 · outbound

This paper cites The Llama 3 Herd of Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.858777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.858777Z digest=sha256:9baa410a688e8d46c40fc9bd48276817bf631410212694a4cc43730fed1ba467

Observation e50db835-66ed-4bc0-be9d-8dd885ffb078 · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.296489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:58.905319Z digest=sha256:3b2ae035bb528ed7753fe8abc3e68804f70a2fc80137c60f31d294c41705a983

Observation 75835048-2848-449e-b468-48470fc8d4d8 · outbound

This paper cites HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.977041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.977041Z digest=sha256:bae43674f2689d219e1eefd1a03e4e0b79b11a3708659b877da32c0ae97c6312

Observation 208f9e6d-4ddf-402e-904d-6970c1895359 · outbound

This paper cites Stance detection with collaborative role-infused llm-based agents.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Stance detection with collaborative role-infused llm-based agents

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.274908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:59.013226Z digest=sha256:9f9831096b1fa7304318aeb0c0692a97555cd18473aad24bde32e4478f4e6930

Observation 9c59abd1-1c3a-417f-917f-fa74e16f8d15 · outbound

This paper cites A Survey of Large Language Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A Survey of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.120276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.120276Z digest=sha256:98d4f8d9bf94c27dae9937f1f1a02f02d10d0465385f1ef8acae0fa002c59704

Observation 2b89717b-1b06-4ca5-a502-4431b39f1141 · outbound

This paper cites A survey on evaluation of large language models.ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey on evaluation of large language models.ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.187554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.187554Z digest=sha256:0535b8be83b84277b99b82876c477f7af1c52431aab196798191dc6afd043800

Observation 5e1fc835-6489-454b-9871-983b131ae05f · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.239623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.239623Z digest=sha256:fd56a5b41e54529f641c694e0b6fff2bea0641764a2217136fb7c6549e79d046

Observation 0feaeef7-b1aa-4336-a23b-c0835c105e9f · outbound

This paper cites Human-level control through deep reinforcement learning.Nature, 518(7540):529–533, 2015.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Human-level control through deep reinforcement learning.Nature, 518(7540):529–533, 2015

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.336489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.336489Z digest=sha256:59cfb105ff369111def8b76364d0eae5b1c9b6e64b39356aa32b0852bbc6afa1

Observation 64033e28-5714-48c8-ba06-23aa04f9f757 · outbound

This paper cites Deep reinforcement learning based ensemble model for rumor tracking.Information Systems, 103:101772, 2022.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Deep reinforcement learning based ensemble model for rumor tracking.Information Systems, 103:101772, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.226796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:59.406512Z digest=sha256:85d868dfd6584ff1793f3ec9a089ac00f37670be6028c6b3e986947973680ef7

Observation af3c0484-5644-417c-9949-b2722bfb93fe · outbound

This paper cites An oppositional-cauchy based gsk evolutionary algorithm with a novel deep ensemble reinforcement learning strategy for covid-19 diagnosis.Applied Soft Computing, 111:107675, 2021.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One An oppositional-cauchy based gsk evolutionary algorithm with a novel deep ensemble reinforcement learning strategy for covid-19 diagnosis.Applied Soft Computing, 111:107675, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.179323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:59.516768Z digest=sha256:13672353d4b63f976ef2b807bc38e37ff2ae402120473bc20fddb535435f5584

Observation fe683cb7-9a2a-4882-8a66-49e4a15e6b1d · outbound

This paper cites Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.577571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.577571Z digest=sha256:b60a5f49fa84661d6566dde2bf587fa825d111aea225807c8eda960a4ce2b814

Observation d068ec1c-548d-4d1b-8f5e-9195ddd2e1f5 · outbound

This paper cites Augmenting autotelic agents with large language models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Augmenting autotelic agents with large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.080388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:59.675486Z digest=sha256:805cb783703170436b5c1484f2b5e637eb40d8bb18ee51f498d7ea5c597a8b60

Observation 5674d13e-b4ba-43fb-b7cc-48b26df7d801 · outbound

This paper cites Read and reap the rewards: Learning to play atari with the help of instruction manuals.Advances in Neural Information Processing Systems, 36, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Read and reap the rewards: Learning to play atari with the help of instruction manuals.Advances in Neural Information Processing Systems, 36, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:00.823614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:24:59.806577Z digest=sha256:80519e02441f3c35ce28eeab1acf464fe2064870d82e779070e94244b0e608b1

Observation 1bfd96f7-721e-46bc-821e-f6528fd842e8 · outbound

This paper cites Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.961027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.961027Z digest=sha256:09fd02fc8ff9afbd492aa3ce6e27ab94456aa7be4f94e175e1029aefb41c227a

Observation e6e1b00b-aa8c-4a42-a3a7-d870bb862034 · outbound

This paper cites Text2reward: Reward shaping with language models for reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Text2reward: Reward shaping with language models for reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:00.632744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:25:00.040054Z digest=sha256:42560162709ce86ea58f5acbc802de6805427fffc3d9519fd67a6560b5358036

Observation 61197aeb-f728-4ac8-bfb1-da8bba5d30b2 · outbound

This paper cites LLM-Empowered State Representation for Reinforcement Learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One LLM-Empowered State Representation for Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:25:00.044677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:25:00.044677Z digest=sha256:83d8e71d4858a737980ee32065142a95d6642f9580cf8511cdf576b8b0f9ce90

Observation dcc67497-61f2-4cda-bb79-433cc4f087b4 · outbound

This paper cites Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:25:00.049554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:25:00.049554Z digest=sha256:b3378bb239a4c83fb9d0305615b3aea9556bd8a5ab2205499b8fc8923ed7e52c

Observation b0b4ac66-6707-458b-a5e8-9cec8977cf8d · outbound

This paper cites Exploration Situation.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Exploration Situation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:00.520873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:25:00.054246Z digest=sha256:ecd97ca46ea48c901eef2ef2f2166ee6adeadd60be9cae996bb70e43cf1ccf8c

Pith citing papers

No inbound Pith citation observations are available.