Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:25:00.054246Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2505.15306.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:25:00.054246Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7ad5cac9-df3a-44a8-bfda-9bd22a7c4f8b · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement learning: An introduction.A Bradford Book, 2018
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b65d7d72-dd7a-4714-becf-e4a3fe7b93d8 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One An introduction to deep reinforcement learning.Foundations and Trends® in Machine Learning, 11(3-4):219–354, 2018
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e379d43f-4d11-4265-8313-801cd0839fad · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Mastering the game of go without human knowledge.nature, 550(7676):354–359, 2017
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2db00582-0d4f-4bac-856b-d4dcff09eaab · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c85c06-0d0d-4240-a7f2-d61777df09e3 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Dota 2 with Large Scale Deep Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb94c72c-b50b-45d6-98d8-4ab6433b0bf5 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Mastering atari games with limited data.Advances in neural information processing systems, 34:25476–25488, 2021
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78b34cf3-9b9d-4b81-85d1-48c022f7a2b8 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A graph placement methodology for fast chip design.Nature, 594(7862):207–212, 2021
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d14d986c-aa0b-4a88-a7b6-4179af0f8d06 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Hierarchical reinforcement learning for scarce medical resource allocation with imperfect information
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99e60500-87cd-4591-81ca-4282dfb822a5 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement learning enhances the experts: Large-scale covid-19 vaccine allocation with multi-factor contact network
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0e78dcd-1a14-4ab8-9db4-436492db1d43 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Gat-mf: Graph attention mean field for very large scale multi-agent reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1638ec47-1810-47be-9ac0-132adbf5f3bb · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Spatial planning of urban communities via deep reinforcement learning.Nature Computational Science, 3(9):748– 762, 2023
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15e054f8-c016-49d9-a301-117d0711c46d · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey of machine learning for urban decision making: Applications in planning, transportation, and healthcare.ACM Computing Surveys, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 39339997-605b-4318-bbbf-bfec067c9de8 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Dyps: Dynamic parameter sharing in multi-agent reinforcement learning for spatio-temporal resource allocation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3652024f-61f0-4ba1-b24b-bf63788475f8 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Coopride: Cooperate all grids in city-scale ride-hailing dispatching with multi-agent reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a86bb074-2a9b-47c2-90ba-95720d7c2bec · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52045072-bd79-4cbd-a6a1-4dafc2302c9e · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c5c68b-f8d5-41cd-840a-450a3b80d005 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble deep learning: A review.Engineering Applications of Artificial Intelligence, 115:105151, 2022
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85e2220f-f902-4ad9-a9af-b91f988f89e4 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey on ensemble learning.Frontiers of Computer Science, 14:241–258, 2020
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 422ef0c2-f114-4985-bf1d-8e18ca44d06a · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble reinforcement learning: A survey.Applied Soft Computing, page 110975, 2023
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69160df6-5acd-4af3-b01f-1599a827a21e · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Neural network ensembles in reinforcement learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 499cc08f-ef3e-4042-b788-fac959861c37 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement Learning based dynamic weighing of Ensemble Models for Time Series Forecasting
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5a597ba-0a2c-4a3b-b84f-36281b50f99a · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble algorithms in reinforcement learning.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 38(4):930–936, 2008
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc9b4d82-1585-4b67-a68d-64dd57681529 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One The arcade learning environment: An evaluation platform for general agents.Journal of Artificial Intelligence Research, 47:253–279, 2013
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e8f1398-449e-4f9f-a24d-73586e8337f3 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Model-Based Reinforcement Learning for Atari
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80b1eed5-579f-40ec-9f56-ebcb17ade291 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4ac7fd4-1ae1-47eb-bc30-7a4d805c4141 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey of gpt-3 family large language models including chatgpt and gpt-4.Natural Language Processing Journal, page 100048, 2023
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2cec5757-1406-40ba-8e63-a0a23a7a4af5 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One GPT-4 Technical Report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ed98770-0f4d-4d91-a149-51d27d2c74ac · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Evaluation of openai o1: Opportunities and challenges of agi.arXiv preprint arXiv:2409.18486, 2024
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bd3492f-c5fb-4285-91ba-dc3f233a93e9 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Early access for safety testing
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e419586-5f54-4b2f-97a6-d767d7a28875 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c4a7cd-d620-47d9-aa37-efa2acf441b9 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One The Llama 3 Herd of Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e50db835-66ed-4bc0-be9d-8dd885ffb078 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75835048-2848-449e-b468-48470fc8d4d8 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 208f9e6d-4ddf-402e-904d-6970c1895359 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Stance detection with collaborative role-infused llm-based agents
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c59abd1-1c3a-417f-917f-fa74e16f8d15 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A Survey of Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b89717b-1b06-4ca5-a502-4431b39f1141 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey on evaluation of large language models.ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e1fc835-6489-454b-9871-983b131ae05f · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0feaeef7-b1aa-4336-a23b-c0835c105e9f · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Human-level control through deep reinforcement learning.Nature, 518(7540):529–533, 2015
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64033e28-5714-48c8-ba06-23aa04f9f757 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Deep reinforcement learning based ensemble model for rumor tracking.Information Systems, 103:101772, 2022
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af3c0484-5644-417c-9949-b2722bfb93fe · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One An oppositional-cauchy based gsk evolutionary algorithm with a novel deep ensemble reinforcement learning strategy for covid-19 diagnosis.Applied Soft Computing, 111:107675, 2021
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe683cb7-9a2a-4882-8a66-49e4a15e6b1d · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d068ec1c-548d-4d1b-8f5e-9195ddd2e1f5 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Augmenting autotelic agents with large language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5674d13e-b4ba-43fb-b7cc-48b26df7d801 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Read and reap the rewards: Learning to play atari with the help of instruction manuals.Advances in Neural Information Processing Systems, 36, 2024
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1bfd96f7-721e-46bc-821e-f6528fd842e8 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6e1b00b-aa8c-4a42-a3a7-d870bb862034 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Text2reward: Reward shaping with language models for reinforcement learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61197aeb-f728-4ac8-bfb1-da8bba5d30b2 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One LLM-Empowered State Representation for Reinforcement Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcc67497-61f2-4cda-bb79-433cc4f087b4 · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0b4ac66-6707-458b-a5e8-9cec8977cf8d · outbound
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Exploration Situation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.