Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:51.437970Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2505.20225.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:51.437970Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8f7cb68c-c85a-429a-b05b-0f20b46cd02f · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9399a12-234f-4d58-953b-785a7351a57b · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Emergent and predictable memorization in large language models.Advances in Neural Information Processing Systems, 36:28072–28090, 2023
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1066643-3fa7-4900-bf49-dd1df442b687 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Pythia: A suite for analyzing large language models across training and scaling
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19b5c372-8288-4763-b652-6a08202985e1 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models PIQA: Reasoning about physical commonsense in natural language
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4cc5e103-b3f9-402d-a872-6b5f11822cc6 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models On the representation collapse of sparse mixture of experts.Proc
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17867122-d6c0-4dd7-a013-08ec8d2efbcf · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Think you have solved question answering? Try ARC, the ai2 reasoning challenge.ArXiv preprint, 2018
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 979caedc-952a-44ab-8c32-168840215b12 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.JMLR, 2022
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce109ed3-19b0-4f5f-8f44-e4aa15571a9f · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models A framework for few-shot language model evaluation, 2023
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0260bff8-ab34-40aa-a0e4-839c568c536f · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Fastmoe: A fast mixture-of-expert training system.ArXiv preprint, 2021
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 837e542e-4add-4452-860f-13db535962b7 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Measuring massive multitask language understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a02d5d3-a089-44f8-a061-99e75156a5d4 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Rae, and Laurent Sifre
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ec5f9de-fb94-47fb-8161-ede4a9723c2c · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models MiniCPM: Unveiling the potential of small language models with scalable training strategies.ArXiv preprint, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d514ff9-0f8f-4d37-909c-3756780d7cc0 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Demystifying Verbatim Memorization in Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cf0d8ca-ff43-48f7-8858-1af889e32b07 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Tutel: Adaptive mixture-of-experts at scale.Proc
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7836f992-70e3-44e5-96aa-b364a4855eee · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Mixtral of experts.ArXiv preprint, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 600cf498-58d5-4668-a673-ceed6b75624c · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Kingma and Jimmy Ba
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91fb215e-94c9-4a4a-ba68-68fac46da7b2 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Gshard: Scaling giant models with condi- tional computation and automatic sharding.ArXiv preprint, 2020
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57be5bee-4181-4a08-a047-9b862dc2c2d6 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Base layers: Simplifying training of large, sparse models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32991eed-24c1-4f1d-8a80-38227fbea49e · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models DataComp-LM: In search of the next generation of training sets for language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 205aaada-f018-4f97-8661-88b85581ae3c · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language model.ArXiv preprint, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2488fdc3-327d-4a77-a490-b7155880a952 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Deepseek-v3 technical report.ArXiv preprint, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c893d78f-3ea9-4cd1-bfd2-30062ab246b4 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Tending Towards Stability: Convergence Challenges in Small Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 347390f4-ba5f-4991-9f2d-07015a526b31 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377c94ac-495e-4e96-84f2-1f3ca751b85a · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Can a suit of armor conduct electricity? A new dataset for open book question answering
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e3005a9-6025-4c94-977a-647799d1facc · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models OLMoE: Open mixture-of- experts language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42cb3bec-3ff6-43e2-9c9d-19449809682d · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Attributing mode collapse in the fine-tuning of large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59ccef2d-e88b-4f1f-b0ff-247c8dcb8098 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9049d2e6-f188-4158-a940-595c4b82c435 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Winogrande: An adversarial winograd schema challenge at scale
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee80a897-dbc9-4f7e-8163-6394dfc5ff60 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.ArXiv preprint, 2017
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 604aefab-5833-440f-b97d-1813b83ca0a6 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models JetMoE: Reaching llama2 performance with 0.1 m dollars.ArXiv preprint, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1fc7123e-646d-4393-8b86-6fab5e0cb37f · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Megatron-lm: Training multi-billion parameter language models using model parallelism.ArXiv preprint, 2019
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b62554e-adc6-4893-8488-a990ba9a1b9b · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Layer by Layer: Uncovering Hidden Representations in Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59facc6e-3fac-4b6c-90c8-1093272969e9 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Free dolly: Introducing the world’s first truly open instruction-tuned llm
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b53b6c84-44e8-4434-9e21-a2ac9b6741f3 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.ArXiv preprint, 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 675da49c-9505-4b43-86d1-2942a9db9fec · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models LLM Circuit Analyses Are Consistent Across Training and Scale
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7797fdd4-71a3-47c6-ab3e-5ba497646968 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models OpenMoE: An early effort on open mixture-of-experts language models.ArXiv preprint, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 236e214e-2813-4c4f-a6d9-7e262475f5f7 · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models M6-T: Exploring sparse expert models and beyond.arXiv preprint, 2021
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fb563fd-cd57-47fa-8787-b1992250163a · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models HellaSwag: Can a machine really finish your sentence? InProc
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7475c41a-9ede-430d-bf65-5afe0a0e104e · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models OPT: Open pre-trained transformer language models.ArXiv preprint, 2022
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 846f8683-413e-4ca0-842b-a049c33d399f · outbound
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models ST-MoE: Designing stable and transferable sparse expert models.ArXiv preprint, 2022
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.