Pith. sign in

Paper Citation Record · LEDGER

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2505.20225.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20225 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:51.437970Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f7cb68c-c85a-429a-b05b-0f20b46cd02f · outbound

This paper cites Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:44.098280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:44.098280Z digest=sha256:9ecc7a1481644e4349d09c71d4a1317df04e0f5a46d2faf5900a24599b1d78e8

Observation c9399a12-234f-4d58-953b-785a7351a57b · outbound

This paper cites Emergent and predictable memorization in large language models.Advances in Neural Information Processing Systems, 36:28072–28090, 2023.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Emergent and predictable memorization in large language models.Advances in Neural Information Processing Systems, 36:28072–28090, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:44.185028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:44.185028Z digest=sha256:cb7d5d370342124b58b9703da437532208f51d0081c95490299194b6d4744701

Observation b1066643-3fa7-4900-bf49-dd1df442b687 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Pythia: A suite for analyzing large language models across training and scaling

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.347994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:44.270403Z digest=sha256:7ad16a519d95b0b3bc89095206b9a90cc1549b3bbc773dc6a1ec7ae4392436d9

Observation 19b5c372-8288-4763-b652-6a08202985e1 · outbound

This paper cites PIQA: Reasoning about physical commonsense in natural language.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models PIQA: Reasoning about physical commonsense in natural language

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.131655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:44.386672Z digest=sha256:d78342bf73a338c156ef77c7a1eda3daaa3da9687dfaa3c073ddc00f8f03c0d9

Observation 4cc5e103-b3f9-402d-a872-6b5f11822cc6 · outbound

This paper cites On the representation collapse of sparse mixture of experts.Proc.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models On the representation collapse of sparse mixture of experts.Proc

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.991684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:44.465623Z digest=sha256:12f60b151777c7468e1bacbbe92162a13a058e70044b5a68e288f5a63b6c53cb

Observation 17867122-d6c0-4dd7-a013-08ec8d2efbcf · outbound

This paper cites Think you have solved question answering? Try ARC, the ai2 reasoning challenge.ArXiv preprint, 2018.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Think you have solved question answering? Try ARC, the ai2 reasoning challenge.ArXiv preprint, 2018

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.694982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:44.535555Z digest=sha256:6cdfef9857c08a6886e6d38e6fce932231045fd691b8530a6c7a27a9c89c4ede

Observation 979caedc-952a-44ab-8c32-168840215b12 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.JMLR, 2022.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.JMLR, 2022

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:44.870448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:44.870448Z digest=sha256:998b03431fc6a5bbabcd8f2769ed19bd344ba24135a4d7cf4837cf2b5c5e4113

Observation ce109ed3-19b0-4f5f-8f44-e4aa15571a9f · outbound

This paper cites A framework for few-shot language model evaluation, 2023.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models A framework for few-shot language model evaluation, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:46.280209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:46.280209Z digest=sha256:ed690b55ae2eb5b2809fbf0c3d3ae0c580573a06ebb8f328bee620e760875a57

Observation 0260bff8-ab34-40aa-a0e4-839c568c536f · outbound

This paper cites Fastmoe: A fast mixture-of-expert training system.ArXiv preprint, 2021.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Fastmoe: A fast mixture-of-expert training system.ArXiv preprint, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.459809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:46.500738Z digest=sha256:1a5f5623bcc8c8d76b888d20a9689e43c717b379e7b6bce688e3b92d7a953d7c

Observation 837e542e-4add-4452-860f-13db535962b7 · outbound

This paper cites Measuring massive multitask language understanding.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Measuring massive multitask language understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.294055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:47.956423Z digest=sha256:e8118c91b24f10c87e34569d00f8f1fd5d09f026137e95c1ac14246a1eb33a96

Observation 2a02d5d3-a089-44f8-a061-99e75156a5d4 · outbound

This paper cites Rae, and Laurent Sifre.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Rae, and Laurent Sifre

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.108386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:48.421407Z digest=sha256:df59bec0ac814eceeff5a8f0dcb534523823198cc7c13eb32a7bf1395014042f

Observation 1ec5f9de-fb94-47fb-8161-ede4a9723c2c · outbound

This paper cites MiniCPM: Unveiling the potential of small language models with scalable training strategies.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models MiniCPM: Unveiling the potential of small language models with scalable training strategies.ArXiv preprint, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:55.955220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:48.559340Z digest=sha256:aad0adbcfeeda458d82387a260be7c03634b0a48af5f6b13aab2e796bb0f50c0

Observation 4d514ff9-0f8f-4d37-909c-3756780d7cc0 · outbound

This paper cites Demystifying Verbatim Memorization in Large Language Models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Demystifying Verbatim Memorization in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:48.631328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:48.631328Z digest=sha256:0293321ef9cc6a441bef87cb6bfef9e06818d60dccd4c5f1c4ac02280c406f86

Observation 9cf0d8ca-ff43-48f7-8858-1af889e32b07 · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale.Proc.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Tutel: Adaptive mixture-of-experts at scale.Proc

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:55.778999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:48.696471Z digest=sha256:d4e4a1b56e28bd8eef2e7acce258790a0fab4738fe35efb787663feebb7134a0

Observation 7836f992-70e3-44e5-96aa-b364a4855eee · outbound

This paper cites Mixtral of experts.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Mixtral of experts.ArXiv preprint, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:55.645820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:48.774085Z digest=sha256:2992e422d21760031c91df11f40c6575068585157709ffc854c2cf2d600894ba

Observation 600cf498-58d5-4668-a673-ceed6b75624c · outbound

This paper cites Kingma and Jimmy Ba.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Kingma and Jimmy Ba

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:48.903883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:48.903883Z digest=sha256:1905c48a5feb7ed3a59c388d281a66ea24532c11730d848330b52f5c1e634c73

Observation 91fb215e-94c9-4a4a-ba68-68fac46da7b2 · outbound

This paper cites Gshard: Scaling giant models with condi- tional computation and automatic sharding.ArXiv preprint, 2020.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Gshard: Scaling giant models with condi- tional computation and automatic sharding.ArXiv preprint, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:55.494798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:48.998050Z digest=sha256:eb4b0b62c955c02c7c6bdad206df19b64bac4a8dae79cee95006099d768bf3e6

Observation 57be5bee-4181-4a08-a047-9b862dc2c2d6 · outbound

This paper cites Base layers: Simplifying training of large, sparse models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Base layers: Simplifying training of large, sparse models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:55.322288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:49.095709Z digest=sha256:6f38971c17ff71ed34f41b34b32cb32c3a75596aeca2b47705710663890511e5

Observation 32991eed-24c1-4f1d-8a80-38227fbea49e · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models DataComp-LM: In search of the next generation of training sets for language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:55.097841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:49.225017Z digest=sha256:d2b0278e937bc19fc749b5707b726842dbb987d26fd057f590615e49bca9b975

Observation 205aaada-f018-4f97-8661-88b85581ae3c · outbound

This paper cites DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language model.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language model.ArXiv preprint, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:54.915523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:49.299350Z digest=sha256:3087c2d18587af9e2c163a7c80c63bbc769f83fd91b25b4614fbb0940d5cb6dc

Observation 2488fdc3-327d-4a77-a490-b7155880a952 · outbound

This paper cites Deepseek-v3 technical report.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Deepseek-v3 technical report.ArXiv preprint, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:54.699463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:49.376132Z digest=sha256:7e229a04e2327fa84a67ff11de32fef919a589ce382f68743b5648068326d258

Observation c893d78f-3ea9-4cd1-bfd2-30062ab246b4 · outbound

This paper cites Tending Towards Stability: Convergence Challenges in Small Language Models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Tending Towards Stability: Convergence Challenges in Small Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:02:51.956017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:49.501554Z digest=sha256:87e7d3a767d301d8f41a1e1f3d3590353e12bcae8036d5cc6277c3d4ba42d17e

Observation 347390f4-ba5f-4991-9f2d-07015a526b31 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, 2025.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:49.622260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:49.622260Z digest=sha256:e051463bcd7088bde524941d35ca327250cfacc81e045452626baf11aa73e1dd

Observation 377c94ac-495e-4e96-84f2-1f3ca751b85a · outbound

This paper cites Can a suit of armor conduct electricity? A new dataset for open book question answering.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Can a suit of armor conduct electricity? A new dataset for open book question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:54.543552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:49.725393Z digest=sha256:61cabb354be50fe3199775be3bc8790f912dac1ef6fefbb4dc97d5b230957be1

Observation 2e3005a9-6025-4c94-977a-647799d1facc · outbound

This paper cites OLMoE: Open mixture-of- experts language models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models OLMoE: Open mixture-of- experts language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:54.352140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:49.831717Z digest=sha256:8bc1af15f13aa524488230d4924fef490f9e56c02c2dc05f7386c05c872779bd

Observation 42cb3bec-3ff6-43e2-9c9d-19449809682d · outbound

This paper cites Attributing mode collapse in the fine-tuning of large language models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Attributing mode collapse in the fine-tuning of large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:54.222040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:49.925737Z digest=sha256:86e32b31a4d8dfd97de9303b853e137b7917c1959bc80dc960af89ea6db97b4a

Observation 59ccef2d-e88b-4f1f-b0ff-247c8dcb8098 · outbound

This paper cites Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:54.005119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:50.016518Z digest=sha256:2d8f1f2c28cc5b91a0c4b4a1544cd0539af891086c642e0981e0262b21f0ac6c

Observation 9049d2e6-f188-4158-a940-595c4b82c435 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Winogrande: An adversarial winograd schema challenge at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:53.845348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:50.132350Z digest=sha256:0b144f2b6dcdc21feba12c0d564cf4683e5fcb49fea2cccc5fb815910b55794c

Observation ee80a897-dbc9-4f7e-8163-6394dfc5ff60 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.ArXiv preprint, 2017.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.ArXiv preprint, 2017

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:53.671642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:50.256613Z digest=sha256:ac806a226e4f78aa721d6a8a464cf0e8379b9d2645b6eab2b0bdd83f55d86167

Observation 604aefab-5833-440f-b97d-1813b83ca0a6 · outbound

This paper cites JetMoE: Reaching llama2 performance with 0.1 m dollars.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models JetMoE: Reaching llama2 performance with 0.1 m dollars.ArXiv preprint, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:53.496149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:50.353681Z digest=sha256:230c2c973228d2ab59f297056f4e52e5c193491e86f450459c15b20fa212f5c8

Observation 1fc7123e-646d-4393-8b86-6fab5e0cb37f · outbound

This paper cites Megatron-lm: Training multi-billion parameter language models using model parallelism.ArXiv preprint, 2019.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Megatron-lm: Training multi-billion parameter language models using model parallelism.ArXiv preprint, 2019

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:53.348965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:50.455008Z digest=sha256:6234f64819cf4b7663f7fb2b92ed37c3e1fe945ae399bcfcb48b164f3ed63295

Observation 3b62554e-adc6-4893-8488-a990ba9a1b9b · outbound

This paper cites Layer by Layer: Uncovering Hidden Representations in Language Models.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Layer by Layer: Uncovering Hidden Representations in Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:50.561305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:50.561305Z digest=sha256:5a51b95efdd7e37d976feb395ede6aded0f1563b9120225fde895781fafdc7b3

Observation 59facc6e-3fac-4b6c-90c8-1093272969e9 · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Free dolly: Introducing the world’s first truly open instruction-tuned llm

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:53.158449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:50.679920Z digest=sha256:fadc28fccec11bdcd3e2f2621ed27054564026665ae29dcbb84ef908fecc3df2

Observation b53b6c84-44e8-4434-9e21-a2ac9b6741f3 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.ArXiv preprint, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:52.961360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:50.778370Z digest=sha256:e7f7db31c58f89212a82fd00da012019fdab791a597eda935861ca3d6ae29596

Observation 675da49c-9505-4b43-86d1-2942a9db9fec · outbound

This paper cites LLM Circuit Analyses Are Consistent Across Training and Scale.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models LLM Circuit Analyses Are Consistent Across Training and Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:50.878275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:50.878275Z digest=sha256:349268b0f8474738cdb11bad8a7ae5c2a9c0aaecb8cf3abb179f8aafd1aad042

Observation 7797fdd4-71a3-47c6-ab3e-5ba497646968 · outbound

This paper cites OpenMoE: An early effort on open mixture-of-experts language models.ArXiv preprint, 2024.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models OpenMoE: An early effort on open mixture-of-experts language models.ArXiv preprint, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:52.797912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:51.005458Z digest=sha256:6b977590e02e72fc6e2fd1998f2967b2b7ca78b5a7c1ee8e3399f7c42d5cda73

Observation 236e214e-2813-4c4f-a6d9-7e262475f5f7 · outbound

This paper cites M6-T: Exploring sparse expert models and beyond.arXiv preprint, 2021.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models M6-T: Exploring sparse expert models and beyond.arXiv preprint, 2021

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:52.594448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:51.127391Z digest=sha256:05d7530f7fa70d4c06cace9f57318386a2688a16663adbb81aa5b5d4057e2e04

Observation 8fb563fd-cd57-47fa-8787-b1992250163a · outbound

This paper cites HellaSwag: Can a machine really finish your sentence? InProc.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models HellaSwag: Can a machine really finish your sentence? InProc

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:52.401440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:51.236526Z digest=sha256:bf0cbc1f8d5f6d7612f8567122fd4eeee1a0b4ddc92b5882dd85e0c4388da54e

Observation 7475c41a-9ede-430d-bf65-5afe0a0e104e · outbound

This paper cites OPT: Open pre-trained transformer language models.ArXiv preprint, 2022.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models OPT: Open pre-trained transformer language models.ArXiv preprint, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:52.220119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:51.365700Z digest=sha256:26c7c2cae5fc2ccf224b7f064b423dff7810932ef88c09453b7ca4ac481b6d0f

Observation 846f8683-413e-4ca0-842b-a049c33d399f · outbound

This paper cites ST-MoE: Designing stable and transferable sparse expert models.ArXiv preprint, 2022.

FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models ST-MoE: Designing stable and transferable sparse expert models.ArXiv preprint, 2022

Reference 40

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:02:51.728312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:02:51.437970Z digest=sha256:0e3fb12d922bee0fcabf1425a59730f3f86ca8a44e70c853ff98e089f116c205

Pith citing papers

No inbound Pith citation observations are available.