Pith. sign in

Paper Citation Record · LEDGER

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One

As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2505.15306.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15306 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:25:00.054246Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7ad5cac9-df3a-44a8-bfda-9bd22a7c4f8b · outbound

This paper cites Reinforcement learning: An introduction.A Bradford Book, 2018.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement learning: An introduction.A Bradford Book, 2018

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.681461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:56.234369Z digest=sha256:b402dced5ca4d2539ff8aecd3a2b32c5c8b4111ba49931814b1efa80a6d7dbb4

Observation b65d7d72-dd7a-4714-becf-e4a3fe7b93d8 · outbound

This paper cites An introduction to deep reinforcement learning.Foundations and Trends® in Machine Learning, 11(3-4):219–354, 2018.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One An introduction to deep reinforcement learning.Foundations and Trends® in Machine Learning, 11(3-4):219–354, 2018

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.660069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:56.288962Z digest=sha256:4af38890e937c6c2b56d2a3a5f4ae36f6fb89bd57456dc983feb366d03e49e19

Observation e379d43f-4d11-4265-8313-801cd0839fad · outbound

This paper cites Mastering the game of go without human knowledge.nature, 550(7676):354–359, 2017.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Mastering the game of go without human knowledge.nature, 550(7676):354–359, 2017

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:56.346937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:56.346937Z digest=sha256:87a2f5d2a35c65ef522e00e1d61dc3d36f3cf846bd7ca09cca2eaca31a56c24a

Observation 2db00582-0d4f-4bac-856b-d4dcff09eaab · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:56.431165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:56.431165Z digest=sha256:ebcd83df2e9f3309a5e66d5c862cc498e95124b184d08fc2f25447deb8f8bcd4

Observation b5c85c06-0d0d-4240-a7f2-d61777df09e3 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Dota 2 with Large Scale Deep Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:56.541701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:56.541701Z digest=sha256:38f52a78a389e001cf7a21d190450c98134b4f4ae17d079e2e70e9af8a685ad1

Observation eb94c72c-b50b-45d6-98d8-4ab6433b0bf5 · outbound

This paper cites Mastering atari games with limited data.Advances in neural information processing systems, 34:25476–25488, 2021.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Mastering atari games with limited data.Advances in neural information processing systems, 34:25476–25488, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.606430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:56.598174Z digest=sha256:c03e09fabf95dd57e77bfd2d8cbf1ba0808e91d7966921ff812e09a9268f816a

Observation 78b34cf3-9b9d-4b81-85d1-48c022f7a2b8 · outbound

This paper cites A graph placement methodology for fast chip design.Nature, 594(7862):207–212, 2021.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A graph placement methodology for fast chip design.Nature, 594(7862):207–212, 2021

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.584564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:56.693310Z digest=sha256:14ba465d958d79f1a94b45dd8330d87e0c249f2f35fba3ac36b9091473c1519b

Observation d14d986c-aa0b-4a88-a7b6-4179af0f8d06 · outbound

This paper cites Hierarchical reinforcement learning for scarce medical resource allocation with imperfect information.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Hierarchical reinforcement learning for scarce medical resource allocation with imperfect information

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.558171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:56.767959Z digest=sha256:22c9a434078a4254bcd703e10e31e48ebeea3ed97b024794fa5624286bf368bb

Observation 99e60500-87cd-4591-81ca-4282dfb822a5 · outbound

This paper cites Reinforcement learning enhances the experts: Large-scale covid-19 vaccine allocation with multi-factor contact network.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement learning enhances the experts: Large-scale covid-19 vaccine allocation with multi-factor contact network

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.538975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:56.855259Z digest=sha256:1afed79a7026ea5dcad418771b347f12175aa8711c846b2629d6d4bbcf5f85aa

Observation f0e78dcd-1a14-4ab8-9db4-436492db1d43 · outbound

This paper cites Gat-mf: Graph attention mean field for very large scale multi-agent reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Gat-mf: Graph attention mean field for very large scale multi-agent reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.520522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:56.945786Z digest=sha256:6066108627ebc8ced43d89c12391c4b82fb258a83cd1340243336af7c5102e6e

Observation 1638ec47-1810-47be-9ac0-132adbf5f3bb · outbound

This paper cites Spatial planning of urban communities via deep reinforcement learning.Nature Computational Science, 3(9):748– 762, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Spatial planning of urban communities via deep reinforcement learning.Nature Computational Science, 3(9):748– 762, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.047385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.047385Z digest=sha256:82f286580d259cd2e2bf48f11ff296cd68f8e0b18df867ea3da706048938224e

Observation 15e054f8-c016-49d9-a301-117d0711c46d · outbound

This paper cites A survey of machine learning for urban decision making: Applications in planning, transportation, and healthcare.ACM Computing Surveys, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey of machine learning for urban decision making: Applications in planning, transportation, and healthcare.ACM Computing Surveys, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.491251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:57.099196Z digest=sha256:1b94c7b3a751b7e5dea9b1586fdfc24715be4d874eddbd441fcf0ab597ff4727

Observation 39339997-605b-4318-bbbf-bfec067c9de8 · outbound

This paper cites Dyps: Dynamic parameter sharing in multi-agent reinforcement learning for spatio-temporal resource allocation.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Dyps: Dynamic parameter sharing in multi-agent reinforcement learning for spatio-temporal resource allocation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.471478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:57.212195Z digest=sha256:84ea1514d68687b7c8a070d83d7e634a3728cd7a9b7850dbf6fecc22016f6bc6

Observation 3652024f-61f0-4ba1-b24b-bf63788475f8 · outbound

This paper cites Coopride: Cooperate all grids in city-scale ride-hailing dispatching with multi-agent reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Coopride: Cooperate all grids in city-scale ride-hailing dispatching with multi-agent reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.453447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:57.301068Z digest=sha256:3a7eddeba3861d51ab04d1779e08202c8d13216762e16def71a67870b3656f19

Observation a86bb074-2a9b-47c2-90ba-95720d7c2bec · outbound

This paper cites A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.402374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.402374Z digest=sha256:116f4b6a3708444d5a7a75183754ca4ad4a5c02408bbc9c865c0e84cf9961fab

Observation 52045072-bd79-4cbd-a6a1-4dafc2302c9e · outbound

This paper cites How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.516895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.516895Z digest=sha256:77a910aa1ea1decf0489d90d64c1e3d878fbc43808db8b085761b583992f5eff

Observation 73c5c68b-f8d5-41cd-840a-450a3b80d005 · outbound

This paper cites Ensemble deep learning: A review.Engineering Applications of Artificial Intelligence, 115:105151, 2022.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble deep learning: A review.Engineering Applications of Artificial Intelligence, 115:105151, 2022

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.437638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:57.614737Z digest=sha256:2ceeab34d9867c234b1a1340c864110f797cd28e7e62cdb4db4b082441e1e4eb

Observation 85e2220f-f902-4ad9-a9af-b91f988f89e4 · outbound

This paper cites A survey on ensemble learning.Frontiers of Computer Science, 14:241–258, 2020.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey on ensemble learning.Frontiers of Computer Science, 14:241–258, 2020

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.421082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:57.701960Z digest=sha256:9697aa46a241fba2ee5dca773c903dd4fc9aebbb84c072a2aad3201df28b3f51

Observation 422ef0c2-f114-4985-bf1d-8e18ca44d06a · outbound

This paper cites Ensemble reinforcement learning: A survey.Applied Soft Computing, page 110975, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble reinforcement learning: A survey.Applied Soft Computing, page 110975, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.405209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:57.775550Z digest=sha256:795ac51d5595bd979a7ca728604615b591a6c8767af5a72661645e4715f9fb61

Observation 69160df6-5acd-4af3-b01f-1599a827a21e · outbound

This paper cites Neural network ensembles in reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Neural network ensembles in reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.389623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:57.856542Z digest=sha256:5cd141feb64c39bf7dc2a136a97cf882bff8e65f68cf760fc2165bd3d7f19fb4

Observation 499cc08f-ef3e-4042-b788-fac959861c37 · outbound

This paper cites Reinforcement Learning based dynamic weighing of Ensemble Models for Time Series Forecasting.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement Learning based dynamic weighing of Ensemble Models for Time Series Forecasting

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.963590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.963590Z digest=sha256:9c61861489b2903eea8ba0f125f2f3f286c2225fa49b296ba5ec263c2c6df113

Observation f5a597ba-0a2c-4a3b-b84f-36281b50f99a · outbound

This paper cites Ensemble algorithms in reinforcement learning.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 38(4):930–936, 2008.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble algorithms in reinforcement learning.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 38(4):930–936, 2008

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.374081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:58.083139Z digest=sha256:e9b1f13d06d1336d7104bfe3cafe94be91320640806d6d152295c19defe25d35

Observation cc9b4d82-1585-4b67-a68d-64dd57681529 · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.Journal of Artificial Intelligence Research, 47:253–279, 2013.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One The arcade learning environment: An evaluation platform for general agents.Journal of Artificial Intelligence Research, 47:253–279, 2013

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.358572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:58.191168Z digest=sha256:bcee2ccca484e95561f586a7c4b078bf002b41a84dd336431632e8f339255848

Observation 9e8f1398-449e-4f9f-a24d-73586e8337f3 · outbound

This paper cites Model-Based Reinforcement Learning for Atari.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Model-Based Reinforcement Learning for Atari

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.285479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.285479Z digest=sha256:71bfd968fc4c5b4acbce7c63b6bb692b4d87e3940cef48bc8b9af8544577488b

Observation 80b1eed5-579f-40ec-9f56-ebcb17ade291 · outbound

This paper cites an unresolved cited work.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.360163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.360163Z digest=sha256:a705167257d328a17cccb49def19d3803e1df1f63e0a4f093035cc486ac6be9d

Observation e4ac7fd4-1ae1-47eb-bc30-7a4d805c4141 · outbound

This paper cites A survey of gpt-3 family large language models including chatgpt and gpt-4.Natural Language Processing Journal, page 100048, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey of gpt-3 family large language models including chatgpt and gpt-4.Natural Language Processing Journal, page 100048, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.330048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:58.423192Z digest=sha256:7d29fc60c79918be8e33533afb35ee3fa60ab1b11b572aee2916429c4fadf6e2

Observation 2cec5757-1406-40ba-8e63-a0a23a7a4af5 · outbound

This paper cites GPT-4 Technical Report.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.507475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.507475Z digest=sha256:71a566bb527eb1fb86e01e112034608a0b9d71a60585b57f6058cef22f2eb8aa

Observation 9ed98770-0f4d-4d91-a149-51d27d2c74ac · outbound

This paper cites Evaluation of openai o1: Opportunities and challenges of agi.arXiv preprint arXiv:2409.18486, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Evaluation of openai o1: Opportunities and challenges of agi.arXiv preprint arXiv:2409.18486, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.596746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.596746Z digest=sha256:2e4e2b8f05352ff8d6c653d939d55615039dbebf67cd0ff724f34b64f58634f9

Observation 0bd3492f-c5fb-4285-91ba-dc3f233a93e9 · outbound

This paper cites Early access for safety testing.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Early access for safety testing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.719213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.719213Z digest=sha256:228e127c3d8aef181a99b4f8b93866de9826582d12dd0419b4955242a4824cfd

Observation 7e419586-5f54-4b2f-97a6-d767d7a28875 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.773436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.773436Z digest=sha256:3e482b755216f2218a8e5f09afeae4474b98a4c80cd12e70d67926b3dac0cc4a

Observation 81c4a7cd-d620-47d9-aa37-efa2acf441b9 · outbound

This paper cites The Llama 3 Herd of Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.858777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.858777Z digest=sha256:a93c6f583dfdb0e4707c16ec520bbab3e9b9255b52dcfdba111dc407098480f0

Observation e50db835-66ed-4bc0-be9d-8dd885ffb078 · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.296489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:58.905319Z digest=sha256:6d4c0dbda095f8b7dd05a8b25fe0f0780822aa2940d71681c96fc03d7b0a9582

Observation 75835048-2848-449e-b468-48470fc8d4d8 · outbound

This paper cites HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.977041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.977041Z digest=sha256:10cd916de11406ad2ccb9e073916ac765265c0e593c363567eb2e2bef6b44c90

Observation 208f9e6d-4ddf-402e-904d-6970c1895359 · outbound

This paper cites Stance detection with collaborative role-infused llm-based agents.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Stance detection with collaborative role-infused llm-based agents

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.274908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:59.013226Z digest=sha256:d84f011974cc58ec3a0958760ffb25dfdb62d75ac2be44f0c7fa3c60d6f4a5bf

Observation 9c59abd1-1c3a-417f-917f-fa74e16f8d15 · outbound

This paper cites A Survey of Large Language Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A Survey of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.120276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.120276Z digest=sha256:5e1a6e6f363a551b140d3fbefd6cd7841cc2d223c19cb72d43d977522fcfb98c

Observation 2b89717b-1b06-4ca5-a502-4431b39f1141 · outbound

This paper cites A survey on evaluation of large language models.ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey on evaluation of large language models.ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.187554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.187554Z digest=sha256:24a75cafa6fc12a99272b45a3de1784b1f97dfd57afdf9dc2fec60be39c2b27b

Observation 5e1fc835-6489-454b-9871-983b131ae05f · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.239623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.239623Z digest=sha256:cf9d7f7ec6dabd5bbb50e11282020bfe5543b16cc475c2c86fbf1c2b46c5c700

Observation 0feaeef7-b1aa-4336-a23b-c0835c105e9f · outbound

This paper cites Human-level control through deep reinforcement learning.Nature, 518(7540):529–533, 2015.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Human-level control through deep reinforcement learning.Nature, 518(7540):529–533, 2015

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.336489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.336489Z digest=sha256:661c8b33ba9b201ee852a0e174e081c3c414f42bfe6adf29dbd19f9187087c6c

Observation 64033e28-5714-48c8-ba06-23aa04f9f757 · outbound

This paper cites Deep reinforcement learning based ensemble model for rumor tracking.Information Systems, 103:101772, 2022.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Deep reinforcement learning based ensemble model for rumor tracking.Information Systems, 103:101772, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.226796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:59.406512Z digest=sha256:6275d6ed0efcf08a803bc79a36812b9e3adaddf85bb8f43fe4a669f7eb2d999d

Observation af3c0484-5644-417c-9949-b2722bfb93fe · outbound

This paper cites An oppositional-cauchy based gsk evolutionary algorithm with a novel deep ensemble reinforcement learning strategy for covid-19 diagnosis.Applied Soft Computing, 111:107675, 2021.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One An oppositional-cauchy based gsk evolutionary algorithm with a novel deep ensemble reinforcement learning strategy for covid-19 diagnosis.Applied Soft Computing, 111:107675, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.179323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:59.516768Z digest=sha256:b22fd9f828500bc9fbb566e6d3c85590f5c903fb6b20704c1ccbb078f25ca67b

Observation fe683cb7-9a2a-4882-8a66-49e4a15e6b1d · outbound

This paper cites Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.577571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.577571Z digest=sha256:24022781815c23696161e8217b816c5ad7615e05cb887774e96ce07277bbb968

Observation d068ec1c-548d-4d1b-8f5e-9195ddd2e1f5 · outbound

This paper cites Augmenting autotelic agents with large language models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Augmenting autotelic agents with large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.080388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:59.675486Z digest=sha256:e13d384dd42d18b62a1cef2cda0d9ee036f8b028a31e70285142f12c0887ffe8

Observation 5674d13e-b4ba-43fb-b7cc-48b26df7d801 · outbound

This paper cites Read and reap the rewards: Learning to play atari with the help of instruction manuals.Advances in Neural Information Processing Systems, 36, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Read and reap the rewards: Learning to play atari with the help of instruction manuals.Advances in Neural Information Processing Systems, 36, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:00.823614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:24:59.806577Z digest=sha256:2e1eb60b71e86d359911cd409cc77294e22087570f86494172012c82ee5adda7

Observation 1bfd96f7-721e-46bc-821e-f6528fd842e8 · outbound

This paper cites Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.961027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.961027Z digest=sha256:f8a589e0895f5fd55b3c7e528d7bbfc0c4133623a7b138fe52a0bfee08ec7354

Observation e6e1b00b-aa8c-4a42-a3a7-d870bb862034 · outbound

This paper cites Text2reward: Reward shaping with language models for reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Text2reward: Reward shaping with language models for reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:00.632744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:25:00.040054Z digest=sha256:344c5f5e243edd8622c728978b8e064397b906831c0968a27e5707dcbe5430fd

Observation 61197aeb-f728-4ac8-bfb1-da8bba5d30b2 · outbound

This paper cites LLM-Empowered State Representation for Reinforcement Learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One LLM-Empowered State Representation for Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:25:00.044677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:25:00.044677Z digest=sha256:41b86b39b73bd1c18c7f7ab825c06044eb9656b13b14559dce0fefa8beb52d0b

Observation dcc67497-61f2-4cda-bb79-433cc4f087b4 · outbound

This paper cites Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:25:00.049554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:25:00.049554Z digest=sha256:9471c874c1020fd9139628724724d51e2f86b633e73d93415f3466de0f58a32f

Observation b0b4ac66-6707-458b-a5e8-9cec8977cf8d · outbound

This paper cites Exploration Situation.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Exploration Situation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:00.520873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:25:00.054246Z digest=sha256:d75f01ffd282c07c1de3e0e72af8f606434115c29b3a3cbb195d70731da5b092

Pith citing papers

No inbound Pith citation observations are available.