Pith. sign in

Paper Citation Record · LEDGER

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One

As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2505.15306.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15306 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:25:00.054246Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7ad5cac9-df3a-44a8-bfda-9bd22a7c4f8b · outbound

This paper cites Reinforcement learning: An introduction.A Bradford Book, 2018.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement learning: An introduction.A Bradford Book, 2018

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.681461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:56.234369Z digest=sha256:d1b6a019bf4d238ea981a0012cf1c2fd74ead1be7c05f987e2b0939d8887d429

Observation b65d7d72-dd7a-4714-becf-e4a3fe7b93d8 · outbound

This paper cites An introduction to deep reinforcement learning.Foundations and Trends® in Machine Learning, 11(3-4):219–354, 2018.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One An introduction to deep reinforcement learning.Foundations and Trends® in Machine Learning, 11(3-4):219–354, 2018

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.660069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:56.288962Z digest=sha256:2c7adb41b47338b13cd5f627cd7c6d65ac0fe81062c6721fd3bac3a9f9483211

Observation e379d43f-4d11-4265-8313-801cd0839fad · outbound

This paper cites Mastering the game of go without human knowledge.nature, 550(7676):354–359, 2017.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Mastering the game of go without human knowledge.nature, 550(7676):354–359, 2017

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:56.346937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:56.346937Z digest=sha256:87a2f5d2a35c65ef522e00e1d61dc3d36f3cf846bd7ca09cca2eaca31a56c24a

Observation 2db00582-0d4f-4bac-856b-d4dcff09eaab · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:56.431165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:56.431165Z digest=sha256:ebcd83df2e9f3309a5e66d5c862cc498e95124b184d08fc2f25447deb8f8bcd4

Observation b5c85c06-0d0d-4240-a7f2-d61777df09e3 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Dota 2 with Large Scale Deep Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:56.541701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:56.541701Z digest=sha256:38f52a78a389e001cf7a21d190450c98134b4f4ae17d079e2e70e9af8a685ad1

Observation eb94c72c-b50b-45d6-98d8-4ab6433b0bf5 · outbound

This paper cites Mastering atari games with limited data.Advances in neural information processing systems, 34:25476–25488, 2021.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Mastering atari games with limited data.Advances in neural information processing systems, 34:25476–25488, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.606430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:56.598174Z digest=sha256:b71fa10a10fa596fe683dee83e9e89ca99f4d8af4787f6b0e73ccfc7bbed68ac

Observation 78b34cf3-9b9d-4b81-85d1-48c022f7a2b8 · outbound

This paper cites A graph placement methodology for fast chip design.Nature, 594(7862):207–212, 2021.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A graph placement methodology for fast chip design.Nature, 594(7862):207–212, 2021

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.584564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:56.693310Z digest=sha256:742afc7499fcd6ba6b50e92d1f9255b25faffe53c7829db7bae578fa392de2b5

Observation d14d986c-aa0b-4a88-a7b6-4179af0f8d06 · outbound

This paper cites Hierarchical reinforcement learning for scarce medical resource allocation with imperfect information.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Hierarchical reinforcement learning for scarce medical resource allocation with imperfect information

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.558171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:56.767959Z digest=sha256:e7b4789b1f3cc4882b5f78ed31bd726660b3488dbe19fa5c8190a0f7edb36dab

Observation 99e60500-87cd-4591-81ca-4282dfb822a5 · outbound

This paper cites Reinforcement learning enhances the experts: Large-scale covid-19 vaccine allocation with multi-factor contact network.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement learning enhances the experts: Large-scale covid-19 vaccine allocation with multi-factor contact network

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.538975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:56.855259Z digest=sha256:b00f85b45c6d39d2ae912ff47386afb70c7bd93fe9d45651f4880d2667159a2c

Observation f0e78dcd-1a14-4ab8-9db4-436492db1d43 · outbound

This paper cites Gat-mf: Graph attention mean field for very large scale multi-agent reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Gat-mf: Graph attention mean field for very large scale multi-agent reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.520522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:56.945786Z digest=sha256:0c97c9e108437c256918792fb0b415c3d4e12ca5e53cae19d4a9fbd6d5920c20

Observation 1638ec47-1810-47be-9ac0-132adbf5f3bb · outbound

This paper cites Spatial planning of urban communities via deep reinforcement learning.Nature Computational Science, 3(9):748– 762, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Spatial planning of urban communities via deep reinforcement learning.Nature Computational Science, 3(9):748– 762, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.047385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.047385Z digest=sha256:82f286580d259cd2e2bf48f11ff296cd68f8e0b18df867ea3da706048938224e

Observation 15e054f8-c016-49d9-a301-117d0711c46d · outbound

This paper cites A survey of machine learning for urban decision making: Applications in planning, transportation, and healthcare.ACM Computing Surveys, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey of machine learning for urban decision making: Applications in planning, transportation, and healthcare.ACM Computing Surveys, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.491251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:57.099196Z digest=sha256:4f5fa73337313720989344fc1c225760e89deeba174541a6324d4bd06d72e809

Observation 39339997-605b-4318-bbbf-bfec067c9de8 · outbound

This paper cites Dyps: Dynamic parameter sharing in multi-agent reinforcement learning for spatio-temporal resource allocation.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Dyps: Dynamic parameter sharing in multi-agent reinforcement learning for spatio-temporal resource allocation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.471478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:57.212195Z digest=sha256:4d409e5820d0e4f446aa3211456cc95a4a094359a8548188e808081f69ab7c25

Observation 3652024f-61f0-4ba1-b24b-bf63788475f8 · outbound

This paper cites Coopride: Cooperate all grids in city-scale ride-hailing dispatching with multi-agent reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Coopride: Cooperate all grids in city-scale ride-hailing dispatching with multi-agent reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.453447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:57.301068Z digest=sha256:2221a78180c63660a306ae413ceb71f31fa541faf8c74e56d840c8cefef6a405

Observation a86bb074-2a9b-47c2-90ba-95720d7c2bec · outbound

This paper cites A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.402374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.402374Z digest=sha256:116f4b6a3708444d5a7a75183754ca4ad4a5c02408bbc9c865c0e84cf9961fab

Observation 52045072-bd79-4cbd-a6a1-4dafc2302c9e · outbound

This paper cites How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.516895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.516895Z digest=sha256:77a910aa1ea1decf0489d90d64c1e3d878fbc43808db8b085761b583992f5eff

Observation 73c5c68b-f8d5-41cd-840a-450a3b80d005 · outbound

This paper cites Ensemble deep learning: A review.Engineering Applications of Artificial Intelligence, 115:105151, 2022.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble deep learning: A review.Engineering Applications of Artificial Intelligence, 115:105151, 2022

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.437638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:57.614737Z digest=sha256:f1bad238a4f4e43e74ce0dd4438f28f60a0eb3b0662fedd31dd3301037834802

Observation 85e2220f-f902-4ad9-a9af-b91f988f89e4 · outbound

This paper cites A survey on ensemble learning.Frontiers of Computer Science, 14:241–258, 2020.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey on ensemble learning.Frontiers of Computer Science, 14:241–258, 2020

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.421082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:57.701960Z digest=sha256:aa60b5dec64b563299bc712cc9b79155dd739907c3bb7390b57862f4ef09bbac

Observation 422ef0c2-f114-4985-bf1d-8e18ca44d06a · outbound

This paper cites Ensemble reinforcement learning: A survey.Applied Soft Computing, page 110975, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble reinforcement learning: A survey.Applied Soft Computing, page 110975, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.405209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:57.775550Z digest=sha256:db1e4685f0643bf9ac50079ae6abd16eb2de82d5c0109b24ed259b48396aa6a9

Observation 69160df6-5acd-4af3-b01f-1599a827a21e · outbound

This paper cites Neural network ensembles in reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Neural network ensembles in reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.389623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:57.856542Z digest=sha256:5ce49463bf48cb58ce1f516e0736a0d78df1d54457f6378ac5451136f7725bda

Observation 499cc08f-ef3e-4042-b788-fac959861c37 · outbound

This paper cites Reinforcement Learning based dynamic weighing of Ensemble Models for Time Series Forecasting.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement Learning based dynamic weighing of Ensemble Models for Time Series Forecasting

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.963590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.963590Z digest=sha256:9c61861489b2903eea8ba0f125f2f3f286c2225fa49b296ba5ec263c2c6df113

Observation f5a597ba-0a2c-4a3b-b84f-36281b50f99a · outbound

This paper cites Ensemble algorithms in reinforcement learning.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 38(4):930–936, 2008.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble algorithms in reinforcement learning.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 38(4):930–936, 2008

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.374081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:58.083139Z digest=sha256:dc7c633b24f756e4447da70ba9eceb764432829e79423c8e9d0336c4ffa9ce57

Observation cc9b4d82-1585-4b67-a68d-64dd57681529 · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.Journal of Artificial Intelligence Research, 47:253–279, 2013.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One The arcade learning environment: An evaluation platform for general agents.Journal of Artificial Intelligence Research, 47:253–279, 2013

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.358572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:58.191168Z digest=sha256:61351a835c156a66bf2701b0093a95b0047fc67892d6f0332627e5883795ae28

Observation 9e8f1398-449e-4f9f-a24d-73586e8337f3 · outbound

This paper cites Model-Based Reinforcement Learning for Atari.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Model-Based Reinforcement Learning for Atari

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.285479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.285479Z digest=sha256:71bfd968fc4c5b4acbce7c63b6bb692b4d87e3940cef48bc8b9af8544577488b

Observation 80b1eed5-579f-40ec-9f56-ebcb17ade291 · outbound

This paper cites an unresolved cited work.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.360163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.360163Z digest=sha256:a705167257d328a17cccb49def19d3803e1df1f63e0a4f093035cc486ac6be9d

Observation e4ac7fd4-1ae1-47eb-bc30-7a4d805c4141 · outbound

This paper cites A survey of gpt-3 family large language models including chatgpt and gpt-4.Natural Language Processing Journal, page 100048, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey of gpt-3 family large language models including chatgpt and gpt-4.Natural Language Processing Journal, page 100048, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.330048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:58.423192Z digest=sha256:dd164002133114c37b765c425ecaa8b1a1de49ff012ebc9fd46a3ffe623682ca

Observation 2cec5757-1406-40ba-8e63-a0a23a7a4af5 · outbound

This paper cites GPT-4 Technical Report.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.507475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.507475Z digest=sha256:71a566bb527eb1fb86e01e112034608a0b9d71a60585b57f6058cef22f2eb8aa

Observation 9ed98770-0f4d-4d91-a149-51d27d2c74ac · outbound

This paper cites Evaluation of openai o1: Opportunities and challenges of agi.arXiv preprint arXiv:2409.18486, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Evaluation of openai o1: Opportunities and challenges of agi.arXiv preprint arXiv:2409.18486, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.596746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.596746Z digest=sha256:2e4e2b8f05352ff8d6c653d939d55615039dbebf67cd0ff724f34b64f58634f9

Observation 0bd3492f-c5fb-4285-91ba-dc3f233a93e9 · outbound

This paper cites Early access for safety testing.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Early access for safety testing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.719213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.719213Z digest=sha256:228e127c3d8aef181a99b4f8b93866de9826582d12dd0419b4955242a4824cfd

Observation 7e419586-5f54-4b2f-97a6-d767d7a28875 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.773436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.773436Z digest=sha256:3e482b755216f2218a8e5f09afeae4474b98a4c80cd12e70d67926b3dac0cc4a

Observation 81c4a7cd-d620-47d9-aa37-efa2acf441b9 · outbound

This paper cites The Llama 3 Herd of Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.858777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.858777Z digest=sha256:a93c6f583dfdb0e4707c16ec520bbab3e9b9255b52dcfdba111dc407098480f0

Observation e50db835-66ed-4bc0-be9d-8dd885ffb078 · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.296489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:58.905319Z digest=sha256:cb8e5e81148d044e71f6ada0cf142460e1a7b0dc44eae0647de7a218aff23cc6

Observation 75835048-2848-449e-b468-48470fc8d4d8 · outbound

This paper cites HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.977041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.977041Z digest=sha256:10cd916de11406ad2ccb9e073916ac765265c0e593c363567eb2e2bef6b44c90

Observation 208f9e6d-4ddf-402e-904d-6970c1895359 · outbound

This paper cites Stance detection with collaborative role-infused llm-based agents.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Stance detection with collaborative role-infused llm-based agents

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.274908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:59.013226Z digest=sha256:cadbded637ceaabe8baf5ab7c82c9a329165e6d07da8a5ede2cdfce7a0ca85ed

Observation 9c59abd1-1c3a-417f-917f-fa74e16f8d15 · outbound

This paper cites A Survey of Large Language Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A Survey of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.120276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.120276Z digest=sha256:5e1a6e6f363a551b140d3fbefd6cd7841cc2d223c19cb72d43d977522fcfb98c

Observation 2b89717b-1b06-4ca5-a502-4431b39f1141 · outbound

This paper cites A survey on evaluation of large language models.ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey on evaluation of large language models.ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.187554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.187554Z digest=sha256:24a75cafa6fc12a99272b45a3de1784b1f97dfd57afdf9dc2fec60be39c2b27b

Observation 5e1fc835-6489-454b-9871-983b131ae05f · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.239623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.239623Z digest=sha256:cf9d7f7ec6dabd5bbb50e11282020bfe5543b16cc475c2c86fbf1c2b46c5c700

Observation 0feaeef7-b1aa-4336-a23b-c0835c105e9f · outbound

This paper cites Human-level control through deep reinforcement learning.Nature, 518(7540):529–533, 2015.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Human-level control through deep reinforcement learning.Nature, 518(7540):529–533, 2015

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.336489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.336489Z digest=sha256:661c8b33ba9b201ee852a0e174e081c3c414f42bfe6adf29dbd19f9187087c6c

Observation 64033e28-5714-48c8-ba06-23aa04f9f757 · outbound

This paper cites Deep reinforcement learning based ensemble model for rumor tracking.Information Systems, 103:101772, 2022.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Deep reinforcement learning based ensemble model for rumor tracking.Information Systems, 103:101772, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.226796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:59.406512Z digest=sha256:7e6244b0dafbd5d0936d7e7f6186c949677466620b81a58b5a44fd48db8cb9bc

Observation af3c0484-5644-417c-9949-b2722bfb93fe · outbound

This paper cites An oppositional-cauchy based gsk evolutionary algorithm with a novel deep ensemble reinforcement learning strategy for covid-19 diagnosis.Applied Soft Computing, 111:107675, 2021.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One An oppositional-cauchy based gsk evolutionary algorithm with a novel deep ensemble reinforcement learning strategy for covid-19 diagnosis.Applied Soft Computing, 111:107675, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.179323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:59.516768Z digest=sha256:1a88faf930bfe96b8a16087d42db9f837bbc30de4cba55ec9d823a6e34ffb50d

Observation fe683cb7-9a2a-4882-8a66-49e4a15e6b1d · outbound

This paper cites Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.577571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.577571Z digest=sha256:24022781815c23696161e8217b816c5ad7615e05cb887774e96ce07277bbb968

Observation d068ec1c-548d-4d1b-8f5e-9195ddd2e1f5 · outbound

This paper cites Augmenting autotelic agents with large language models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Augmenting autotelic agents with large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.080388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:59.675486Z digest=sha256:d77d121185dddd8052588f2555a3b70a818d8a2d86b0db10748b967cd78acbab

Observation 5674d13e-b4ba-43fb-b7cc-48b26df7d801 · outbound

This paper cites Read and reap the rewards: Learning to play atari with the help of instruction manuals.Advances in Neural Information Processing Systems, 36, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Read and reap the rewards: Learning to play atari with the help of instruction manuals.Advances in Neural Information Processing Systems, 36, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:00.823614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:24:59.806577Z digest=sha256:7752fd0347c4cdd696a1931dc6a5f0c09bdad5a026c57f02b48f7d947ff41224

Observation 1bfd96f7-721e-46bc-821e-f6528fd842e8 · outbound

This paper cites Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.961027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.961027Z digest=sha256:f8a589e0895f5fd55b3c7e528d7bbfc0c4133623a7b138fe52a0bfee08ec7354

Observation e6e1b00b-aa8c-4a42-a3a7-d870bb862034 · outbound

This paper cites Text2reward: Reward shaping with language models for reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Text2reward: Reward shaping with language models for reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:00.632744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:25:00.040054Z digest=sha256:75fd331ba280f27fd183dd2b456a9082adb79b3ef6a0dd2a884e6cf7d02884da

Observation 61197aeb-f728-4ac8-bfb1-da8bba5d30b2 · outbound

This paper cites LLM-Empowered State Representation for Reinforcement Learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One LLM-Empowered State Representation for Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:25:00.044677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:25:00.044677Z digest=sha256:41b86b39b73bd1c18c7f7ab825c06044eb9656b13b14559dce0fefa8beb52d0b

Observation dcc67497-61f2-4cda-bb79-433cc4f087b4 · outbound

This paper cites Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:25:00.049554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:25:00.049554Z digest=sha256:9471c874c1020fd9139628724724d51e2f86b633e73d93415f3466de0f58a32f

Observation b0b4ac66-6707-458b-a5e8-9cec8977cf8d · outbound

This paper cites Exploration Situation.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Exploration Situation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:00.520873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:25:00.054246Z digest=sha256:dfe641038d85b6405101261770a9ff95ad77506c1764b5221283e9f2caff2df5

Pith citing papers

No inbound Pith citation observations are available.