Pith. sign in

Paper Citation Record · LEDGER

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

As of 16 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2608.12253.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12253 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:15:04.780467Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact2
  • verified fuzzy12
  • unresolved59
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ea338534-eff5-4f60-b471-4449fd62e289 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.453312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.453312Z digest=sha256:155cd703fbda48317bdaeeae58f3614a6b77e0f79787585bc6c6fa124b14d934

Observation 8c737c19-cd18-448f-90d1-888e0e7f4e82 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Understanding R1-Zero-Like Training: A Critical Perspective

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.460262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.460262Z digest=sha256:e2f2025ac09a7650d217c09b1f549ec4c357a9b3b2520651c478bcfe6461fd83

Observation a79b4430-5332-4d40-8254-c46cd54e5875 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.465639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.465639Z digest=sha256:e65ea86648af1474ab1fd62cf214bfad316fb87f01a396de2917678a648b6ea2

Observation 9878a06f-50fd-4c68-a7f5-23ce74c7ed8b · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.471444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.471444Z digest=sha256:ef0d8b0daffc50ac8b46ec8c6a5d073865d1005a7569392aed63910246f63302

Observation bc2336a1-ff6b-40d7-9e78-707c7ae155b0 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.476116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.476116Z digest=sha256:6e733ca207aceb9d6417680ac6d1acf4ae82e49a55c7258d9bbfcbe27e8f4912

Observation 66688e70-7873-4b78-b699-10b0eabd0654 · outbound

This paper cites SWE-smith: Scaling Data for Software Engineering Agents.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL SWE-smith: Scaling Data for Software Engineering Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.480953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.480953Z digest=sha256:ef24ca8c58bfd635abc3167ac41d7523a8b6bb10fc3b89b4fb8b78cba109017e

Observation 757ed77b-732e-4bc0-baab-0d2158e02df5 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.486458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.486458Z digest=sha256:12eaa93286fb582129af5ecb2e62cd37cdc509011fac73542a4ce3fd7174de41

Observation a71afce6-d652-4621-aa52-428c2a44339f · outbound

This paper cites UserRL: Training Interactive User-Centric Agent via Reinforcement Learning, September.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL UserRL: Training Interactive User-Centric Agent via Reinforcement Learning, September

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:15:07.001260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.491426Z digest=sha256:cdf4cd3e292a2812b4c9d82143d209e192b4319bd1ee5cc192300eae6e104029

Observation 8fc7d44e-c8c1-40e2-8039-afce0e969ef4 · outbound

This paper cites TOM-SWE: User Mental Modeling For Software Engineering Agents, October 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL TOM-SWE: User Mental Modeling For Software Engineering Agents, October 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.500306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.500306Z digest=sha256:23a49c4877ba5dded54a8f3dda278b04eeae8cce8723b3b00fc922739949c445

Observation 56c06e39-cbb7-4a26-b23a-53b036d7c4f2 · outbound

This paper cites HumanLM: Simulating Users with State Alignment Beats Response Imitation, February 2026.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL HumanLM: Simulating Users with State Alignment Beats Response Imitation, February 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.504225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.504225Z digest=sha256:be82dd2223b4b09e73d7bf64ed59960eef7e66c9bc2835d7581a7421841622f9

Observation 558deb1f-850c-47cf-a278-ee375086da82 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.507966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.507966Z digest=sha256:34a1541e2ab907d5d6db1f78a415547d225ea4528cf14f631745d94d10c97f55

Observation 3e3bdbe8-e2ec-455f-8129-2d9b8a6e7f86 · outbound

This paper cites Lieberwirth, Xinkai Yu, Yicheng Fu, Michael J.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Lieberwirth, Xinkai Yu, Yicheng Fu, Michael J

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.511940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.511940Z digest=sha256:dd84c1afe553524e402cee408f3c12147370bde5c7e77025d814ada6969f338c

Observation a3ec8aee-1627-4f78-a2f2-0d710c2bc7d2 · outbound

This paper cites Position: Humans are missing from ai coding agent research.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Position: Humans are missing from ai coding agent research

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:15:06.986895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.516143Z digest=sha256:af3beb7c986725bd040ca156326f44c321425bd1769c4fb58ebc15c8f1699452

Observation 83b9cb85-b2d3-435b-86c3-92d328846d39 · outbound

This paper cites Persuasion for good: Towards a personalized persuasive dialogue system for social good.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Persuasion for good: Towards a personalized persuasive dialogue system for social good

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:15:06.972671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.520335Z digest=sha256:6752f521d8dae20fd17c8f6b22cd31525a941bccccf0d62067b9a4c8bf89ef88

Observation 73bb10d5-7452-4c18-91d9-21ef361ab848 · outbound

This paper cites Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning, October.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning, October

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.524125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.524125Z digest=sha256:0da59690cf8f762017737e28182dbcecb7b3556422be91b2faf09432f76b3a5b

Observation 354ddfcc-ccf8-446b-a58b-ad800dfc1ed8 · outbound

This paper cites Bernstein.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Bernstein

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.533329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.533329Z digest=sha256:55f414c12f3f10c3542a83a126070f44e77321f75209efa6f67fcd53123e825c

Observation 81d1c642-83c0-4b21-a540-5ea5a779176d · outbound

This paper cites arXiv:2511.00222 [cs].

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL arXiv:2511.00222 [cs]

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.528529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.528529Z digest=sha256:41f01cd48073d1e401dda0a5df1f34e179a5c20192dc79c389f36160865a7de2

Observation 1f1cec51-b85e-4999-b8a3-a21dde612dc8 · outbound

This paper cites Sotopia-RL: Reward Design for Social Intelligence, October 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Sotopia-RL: Reward Design for Social Intelligence, October 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.542316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.542316Z digest=sha256:61d1d70ca04dac17261b70fd975a73a3f41016c8a22bbbde3da96ab81280d482

Observation 53ce8f94-5f2f-409d-80d2-36873209c9da · outbound

This paper cites LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.537288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.537288Z digest=sha256:e47efe4b448ea089d76ad038536fb08ccfe068599ffc377251900da54ad8411f

Observation 4ece0c50-a8cc-47e5-b206-4a5bb9d7f39a · outbound

This paper cites Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond), October 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond), October 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.551128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.551128Z digest=sha256:fa40745bb2d4f2b506386a6bd0498ce7c5cd373bf9bb3c4fa2269cec7e04dcb9

Observation 0826ad6c-26d8-47a0-9f7d-84241ab7d5f9 · outbound

This paper cites LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.546670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.546670Z digest=sha256:54729942e483688c10c47a8084940da5c4a2461854278ac93290e750649d5c28

Observation 4c759508-dca8-4e29-9d72-620b09e13353 · outbound

This paper cites Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.558740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.558740Z digest=sha256:068ea982c30c8f631546a91b61ab1e923973d3c28a22c81d88d97c0016f6c986

Observation a82446fb-cdb2-44c9-94d3-738c206b6ebc · outbound

This paper cites KL-regularized reinforcement learning is designed to mode collapse, 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL KL-regularized reinforcement learning is designed to mode collapse, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.555118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.555118Z digest=sha256:93147a5f2532f5db4de15f2a7f7f1d2881b99600a9e1695234bba71420437725

Observation 5187323f-70e8-4871-b27b-b894c734c55f · outbound

This paper cites Chasing moving targets with online self-play reinforcement learning for safer language models, 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Chasing moving targets with online self-play reinforcement learning for safer language models, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.568741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.568741Z digest=sha256:ae727625936901c7a2b0c2e0d85ccb4e28ec0b82680aa5d8c79cc4ebe237aaff

Observation 1b5f4f66-123a-4944-8531-c3cb8d585c11 · outbound

This paper cites Natural emergent misalignment from reward hacking in production RL, 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Natural emergent misalignment from reward hacking in production RL, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.564081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.564081Z digest=sha256:2952f344522412898f16714823dfc973a1ee9c94f861ecd3d9a02a3d15af369d

Observation 8a2e3376-0928-46c9-9f6e-382e2fc0b6d0 · outbound

This paper cites Williams.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Williams

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.577287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.577287Z digest=sha256:3f2844f933050ab440d19f43e2f9f9851e0ea6f35bacf55645863423fa3511ff

Observation 78310470-81af-40b0-8f96-2253d485a236 · outbound

This paper cites SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning, July 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning, July 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.573268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.573268Z digest=sha256:1112f2f3bff7b17229784931b4f1f739a43e886a7026c577f70e861dd17a5c20

Observation 79b4edda-9754-47ec-bf58-3e997bd46958 · outbound

This paper cites Flipping the Dialogue: Training and Evaluating User Language Models, October 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Flipping the Dialogue: Training and Evaluating User Language Models, October 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.590001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.590001Z digest=sha256:06ade843297e1e6c74aaf81a44380b06df6d8506ef031be882b696b38b1b4cc4

Observation 66956b33-e5c5-4f58-b112-8a0619171687 · outbound

This paper cites Noveltybench: Evaluating language models for humanlike diversity,.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Noveltybench: Evaluating language models for humanlike diversity,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:15:06.942798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.581498Z digest=sha256:35297e51124c6025edf1ec5d8cbf238ac63217fcab1c3e5c03265555044fea50

Observation c524adc3-b61f-4a28-ad18-499a4ec70cbd · outbound

This paper cites NoveltyBench: Evaluating Language Models for Humanlike Diversity.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.585634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.585634Z digest=sha256:aeb034707028e3519181a15b6ee36ab5da6647a2a79a9d207ad2063d8c0fcfa6

Observation 500f716f-e128-4c32-a457-ba34fed773ea · outbound

This paper cites SPICE: Self-Play In Corpus Environments Improves Reasoning, October 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL SPICE: Self-Play In Corpus Environments Improves Reasoning, October 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.602450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.602450Z digest=sha256:edebf6105ed569e9f4d060f73459a3a209fdfed4035f530c78f2586399d0b9fe

Observation f747676c-b894-444b-9219-ced9c253466b · outbound

This paper cites Mind the Sim2Real Gap in User Simulation for Agentic Tasks.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Mind the Sim2Real Gap in User Simulation for Agentic Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.593555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.593555Z digest=sha256:5bd1978ca010a8f5f60fad9cb1bfb66056c5e136b76e7d0cef62a800f69bd4ec

Observation 29fce8ca-4b7b-4ce2-b59d-c7f8cc5b66d9 · outbound

This paper cites Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:15:05.782484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.598211Z digest=sha256:3d547196318e832419aef28d61c58ee03f9e5380633d0dd69856dcaed468b536

Observation e5a69584-fcaf-4d86-9026-f1b8ba864561 · outbound

This paper cites RAGEN-2: Reasoning Collapse in Agentic RL.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL RAGEN-2: Reasoning Collapse in Agentic RL

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.613847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.613847Z digest=sha256:9e506280bc0bce0ecba7a650f1a7f46b5c72016ef06fa8f5d9463935763d7965

Observation 35736ef2-fde6-4f85-8a06-f8422e014cf9 · outbound

This paper cites Hi- erarchical agenda reasoning for strategic multi-turn dialogue agents.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Hi- erarchical agenda reasoning for strategic multi-turn dialogue agents

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:15:06.927989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.606356Z digest=sha256:73afdd09eaefc0f1f464a42d2725563b621a81ec1de5dc80ed62e414ce0a5a8a

Observation 7a78fb76-ff17-4df9-9991-59e7de1a0857 · outbound

This paper cites Llm probability concentration: How alignment shrinks the generative horizon, 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Llm probability concentration: How alignment shrinks the generative horizon, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.610256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.610256Z digest=sha256:caf9e1daa50bf0112a16a2bf4fd9faf358e462bee223a8990515fcb32701ac64

Observation d46331dc-b1ae-4f63-97d0-e5f686348cc2 · outbound

This paper cites MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.627801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.627801Z digest=sha256:17355b28a25e3f752abbeedfa82e25c76a0eee592440097de64b481bc781fe8a

Observation ad985377-0b6b-436b-8bb0-52cfc35b826b · outbound

This paper cites The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.617997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.617997Z digest=sha256:d9bca827db440ed76873875a042c47c05d0ceed418e4f941d8e9f6a3203f8cf6

Observation f9dd7c70-72cd-4694-9e11-3fa92c1170f7 · outbound

This paper cites LLM Social Simulations Are a Promising Research Method.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL LLM Social Simulations Are a Promising Research Method

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.623422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.623422Z digest=sha256:580bf795c7cb4a0a29dc3cb437bfb7ec2ddc8be86d229f2de951fecdb399f34b

Observation fe626c99-dc7f-4f61-a654-bc4936757de8 · outbound

This paper cites Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.642064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.642064Z digest=sha256:8d079af6c5bd233c4fea4eaa8e516c3a5eedb07f9025218488d51a51069aff5b

Observation 6e90aac5-60d9-4210-830d-76adf774ed7e · outbound

This paper cites Training Proactive and Personalized LLM Agents, November 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Training Proactive and Personalized LLM Agents, November 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.632848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.632848Z digest=sha256:fb517133a1abf2bd30ad0a41fe5ba2c99347602b20f677962d71c264b2e9d630

Observation 2a0c0564-38f9-4f7b-be63-d0b313474b43 · outbound

This paper cites an unresolved cited work.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.637451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.637451Z digest=sha256:60b350d563dd3438f8bf2409d4620d490422f114ad97c8b396505464a5bdc1f1

Observation 4085fdb1-4365-4202-aa57-da42c1199097 · outbound

This paper cites Enhancing personalized multi-turn dialogue with curiosity reward, 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Enhancing personalized multi-turn dialogue with curiosity reward, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.654081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.654081Z digest=sha256:21a1a101fa156c6cde2eb7157ad519ac24cf37a5d1fe19f8f84bfe9ddc9fbec0

Observation 3ee0ed04-67ad-4600-82c0-53b2f33555f1 · outbound

This paper cites Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:15:05.239896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.646337Z digest=sha256:9ec8c84a2cb350879338665c91117898ac1baa555a7b6ceee22c901e2c8a12cb

Observation a8357c08-443e-43b1-b8ac-30337175e0d9 · outbound

This paper cites Non- collaborative user simulators for tool agents, September 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Non- collaborative user simulators for tool agents, September 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.650372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.650372Z digest=sha256:8bd88be5388f1465e24a96acce061b44ffc12199af140168f9b9aa7cad56c7e5

Observation ace82092-9a6c-4c0d-944a-246df5695d30 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Dota 2 with Large Scale Deep Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.669651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.669651Z digest=sha256:5ffafd8d7b3cc5815f97757d91f53a9e7b3c86d56746e2b741431faf00b45fae

Observation 070cc558-62ca-4e34-81d2-7ec53b558aab · outbound

This paper cites Multi-User Large Language Model Agents.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Multi-User Large Language Model Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.658345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.658345Z digest=sha256:54ff2b02f25141437ca9afcd399b2a8e408ef92bec13404e38670afacef39122

Observation f82917b2-2ab5-4a64-b603-54e0e999b232 · outbound

This paper cites an unresolved cited work.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.663523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.663523Z digest=sha256:4a447b22546837690492e6105796fbe32ca8ddf9c4028fb1479448b0c962ad38

Observation 93466362-ff19-47fc-bd89-79c4fcf393b6 · outbound

This paper cites Coevolving with the other you: Fine-tuning LLM with sequential cooperative multi-agent reinforcement learning.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Coevolving with the other you: Fine-tuning LLM with sequential cooperative multi-agent reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:15:06.914295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.682545Z digest=sha256:73bb304b71f955c828f469ab739d3d6cc7326c1128c622c9577c432c3e171560

Observation 47ad6414-e8a4-4302-8dd4-6dac1276fd00 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.674667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.674667Z digest=sha256:eeaad1836f9ec222662ba867795d543398763de0e08bf71216d0f5b3e53d82f6

Observation 0e0d651d-7a12-429a-b7ed-7ec8b4a1abfa · outbound

This paper cites Efficacy of Language Model Self-Play in Non-Zero-Sum Games.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Efficacy of Language Model Self-Play in Non-Zero-Sum Games

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.678569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.678569Z digest=sha256:0028c975bacc0cf1597dd250a46ff179f5eb5ba3cd45f71af7875e5c42aa9ec0

Observation 2c457f6f-b581-4af4-ac77-d68b18bbca7f · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.693596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.693596Z digest=sha256:f56f6de8a5f91c64710160173914e179676e597c54e10b57319d6fda2c809ca0

Observation 63c1f234-8256-4890-af30-af80c11a030d · outbound

This paper cites an unresolved cited work.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:15:06.900484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.686224Z digest=sha256:d810aba5942290aeac6cc882927fdb351317fced9bc2425093ffb2997b2bb779

Observation d64f14a8-bfb1-4203-b8e9-ada6c586c6b1 · outbound

This paper cites Tool-R0: Self-evolving LLM agents for tool-learning from zero data, 2026.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Tool-R0: Self-evolving LLM agents for tool-learning from zero data, 2026

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:15:06.887983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.689937Z digest=sha256:2853a5586f5b6729f14cae41d1c7094dad514d5bbc690c749839fdcabfaebedf

Observation d516b38c-3679-40e4-9459-4b727b4c1c38 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Manning, Stefano Ermon, and Chelsea Finn

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:15:06.873951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.705959Z digest=sha256:344acc247d6fcd2d0daa5d692924ae14b89e631f733aa0fa1b325e3bf1593a00

Observation a7a17a8e-a712-41c7-9b63-b7075dee5a45 · outbound

This paper cites Natural Language Actor- Critic: Scalable Off-Policy Learning in Language Space, December 2025.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Natural Language Actor- Critic: Scalable Off-Policy Learning in Language Space, December 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.697681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.697681Z digest=sha256:aaa0f20ece9af780c1c3e9987391c83639fdd41dacd49e5e4c388a291131f836

Observation c1cb2f5c-9429-4353-a794-7ef920329924 · outbound

This paper cites Ma, Seun Eisape, Ellie French, Tingting Du, Tianjiao Zhang, Alexander Koller, and Alane Suhr.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Ma, Seun Eisape, Ellie French, Tingting Du, Tianjiao Zhang, Alexander Koller, and Alane Suhr

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.701587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.701587Z digest=sha256:c62f262d9f4ff7b9061bef296a4849769e1a7077258def14e1bc9bf644242331

Observation a0b5a150-ffbf-49ad-a1e3-8cbaf1a3e4aa · outbound

This paper cites OpenAI Gym.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL OpenAI Gym

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.719186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.719186Z digest=sha256:84de6b46707d80a2e8529a9bf5740cb2afb3da8cc528ded0a1c9009ada247d68

Observation 904f6f92-78bf-45fc-aa6b-145d9552b2d0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.710147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.710147Z digest=sha256:a3c491dbb6efe9e3932b87e465ce3426c6ea9665faa6506005134d3e82fd625a

Observation 3adb7e51-8235-488b-b820-79e5e5a13341 · outbound

This paper cites Proximal Policy Optimization Algorithms.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Proximal Policy Optimization Algorithms

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.714785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.714785Z digest=sha256:6e65ebbd70006d1319079dac70dc97904bf63f8724f73751564e7f810863e8a5

Observation b3fcff93-69b1-487c-ba01-85caeb68a3b1 · outbound

This paper cites slime: An llm post-training framework for rl scaling.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL slime: An llm post-training framework for rl scaling

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.731708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.731708Z digest=sha256:2bbf3692ce3cd3c9baa918edb80fb272896d33a5be1d1a8264b51b1e49c48984

Observation e1252219-a1fd-4789-85e7-d96bcc40214c · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL SGLang: Efficient Execution of Structured Language Model Programs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.723393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.723393Z digest=sha256:9e481c24677019899bef7f62e5b890cf8414e342204e4e8575426ff814424344

Observation 83045f16-1c71-4538-b8b6-a9be9f71207b · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.727651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.727651Z digest=sha256:bab734bfffb21003b339e8a8c9eb137bc511cc5f69ea3add8e21f6f69589c38a

Observation bbbdf98a-1e51-4c62-8d99-6498a90f2f81 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Qwen3.5: Towards native multimodal agents, February 2026

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.743354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.743354Z digest=sha256:e3ebdc386816fe7c3a39aec596857a80f17e3a1dfb753eabea6c00d578704a12

Observation 7c53bbcc-df9b-4991-aaf0-e03e05a32f86 · outbound

This paper cites HybridFlow: A flexible and efficient RLHF framework.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL HybridFlow: A flexible and efficient RLHF framework

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:15:06.853370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.735504Z digest=sha256:1876069c47687e0a0fbebe0707be3cf10cf0bbd425fc90d668ad8e7314303fdf

Observation 14b14b77-d4f0-4ad5-ace9-8ac3a8523a23 · outbound

This paper cites AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:15:04.855445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.739486Z digest=sha256:bee8b981b5d1021db2d4e1b77341bdf883385a22d450bc29603c3d9242daa8d8

Observation 5ca3e580-9a2d-456c-8bdd-442861246c0c · outbound

This paper cites Olmo 3.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Olmo 3

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.747165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.747165Z digest=sha256:e79a240f42cca7eb8027c21dfbfbec447a2ccec76ec37b7e6ded0edfde835ae3

Observation a5d1f561-7b71-410b-b4d5-0a4db8808385 · outbound

This paper cites SPIRAL [ 25], SPICE [31], Absolute Zero [47]): one model serves both roles with role-specific loss masks.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL SPIRAL [ 25], SPICE [31], Absolute Zero [47]): one model serves both roles with role-specific loss masks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:15:06.833814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.752437Z digest=sha256:89e768d3de60f6db083eae1e5ccca72ea22e71aed04ef2eb3500a489e3dc1f3c

Observation 35a6214d-db4a-424a-9e49-e27de0c33ac4 · outbound

This paper cites an unresolved cited work.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:15:06.821870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.757120Z digest=sha256:2e34315ed290ab2c971cc6165a4c21ef21f37cfc957186ede65d9982f2a72920

Observation 638d7632-dfa5-486d-a1a7-6542aa1c841d · outbound

This paper cites Updated?.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Updated?

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:15:06.809353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.761430Z digest=sha256:2212f13e8fa3409d4b657591d475abab357b6ffdb9c84ce18d483627099d5f7e

Observation 46274806-3801-4e2e-85fa-08471e40c932 · outbound

This paper cites an unresolved cited work.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Unresolved cited work

Reference 72

Resolution
parse uncertain
raw_fallback, observed 2026-08-16T00:15:06.797513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.766075Z digest=sha256:532b805e6bea8392253ebbb649896c980ed9456e59f767493e54594cff21a929

Observation f161cbfc-21f5-4bc0-aa22-439f28c13bee · outbound

This paper cites an unresolved cited work.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:15:06.785421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.770435Z digest=sha256:6aa47ed3fdcdd951b9d6b6f0a42d0568477b8669b4e73f2423f3e32f1c633801

Observation 494a49c9-ada6-4e4a-9c7d-ef6d159a6d21 · outbound

This paper cites an unresolved cited work.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:15:06.772211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.775866Z digest=sha256:0d31052111cabc37b89ce971afa8820d44e830ecf40d4f0018b1aa15bc2875ab

Observation ba90d65d-a7c0-45aa-bcbe-c085f99c2fce · outbound

This paper cites responses.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL responses

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:15:06.757652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:15:04.780467Z digest=sha256:26b87f62a2a328a930dc9269720d025731dd2f7e239c38f2ce6e803a33fafbee

Observation 39f02e0f-dbb2-40b5-ae65-3044662a48e9 · outbound

This paper cites arXiv:2509.19736 [cs].

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL arXiv:2509.19736 [cs]

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.495639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.495639Z digest=sha256:75c4b5c5ee2b121fe52468b2da1dcabdffa001ea9a30c95de0ae2ba3b6692637

Pith citing papers

No inbound Pith citation observations are available.