Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T05:36:26.207359Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 100 of 150 outbound references and 100 inbound Pith citation observations for arXiv:2405.04434.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T05:36:26.207359Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T20:10:00.206774Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 150 outbound references displayed
External citation measurements
5
pith, observed 2026-08-05T02:28:24.338817Z
Observation 944e2709-9885-415d-a193-0c6e364518da · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Llama 3 model card
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d24f057b-ebf7-44bd-92b2-c88a53efb6c3 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Introducing Claude
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0b4d7b31-7be9-435f-bf76-02761533340a · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Evaluating Large Language Models Trained on Code
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4086f8d8-ed30-4bb6-a04f-b87b6be6a561 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 344e9eae-3bdd-46b0-a244-ed55c48ddaa6 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 96908b73-98f0-492f-8f0c-6b23e1a08d19 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ef25b505-3aec-4e29-baae-758878d7d3e7 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3d03e31e-2a1c-4b25-9cb5-8ee700f0d844 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d5e4d8bf-ae36-4be6-a637-36245066b918 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Introducing gemini: our largest and most capable ai model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dc3f1c97-9b26-467f-84f1-2c56e4027b7d · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bdeae05e-a0f6-4ad5-b841-d56d8c48998c · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Hai-llm: 高效且轻量的大模型训练工具
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b6d19c84-6bc0-4145-8471-c17e03e40217 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d2c283fc-0a17-4ec6-9904-b77b004536b1 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a2cdc000-6e7d-4212-9592-b9a058011bda · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d3bd00e7-d85d-40f6-8ebf-5d92140d32d2 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Lepikhin, H
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 04265ffe-e7a5-4734-8d3f-de4c5c7f9c34 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model CMMLU: Measuring massive multitask language understanding in Chinese
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2d280c8b-1874-4aa5-b644-3ac0b629051c · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 53571ff3-e138-4594-b644-bb15a90a1bf4 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Cheaper, better, faster, stronger: Continuing to push the frontier of ai and making it accessible to all
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aeb3ac18-747d-46b9-b40e-790f1a9e076b · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Introducing ChatGPT
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 49e6a07a-77a9-4260-8f4b-9ff7c384fd00 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model GPT-4 Technical Report
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8f7baefc-2620-41dc-a4b0-bc34b81c9421 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Ouyang, J
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3d869ecb-9f7c-4d69-a960-5970ff095217 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Rajbhandari, J
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 479535a2-023f-4628-b1d3-0ff138afe6f2 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Riquelme, J
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c86154ed-fd88-420d-a570-976f0aad53f1 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Sakaguchi, R
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4e883403-25c9-4fc8-a2af-19cb1add8cfa · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Fast Transformer Decoding: One Write-Head is All You Need
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 55539873-c0a8-4081-a92e-cdfc5685b185 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Shazeer, A
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4881d738-2bb2-439f-8aea-ab08d670a168 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 478b12d8-2347-4e75-8b23-d8849a8419ab · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 88c45b90-055d-412a-ab31-cca9c7459dd7 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Vaswani, N
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 18a412c8-19c2-4fae-9ea7-b980700f4f28 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 14b49881-a09e-490f-b4a6-f518a0782182 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 41ebf1dd-bfee-4e1c-bd44-99e759ecc040 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Zheng, W.-L
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 682391b0-dffb-4868-9f4d-e9fca687ae1b · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 317e65dc-af55-41c5-8adf-0ae42109b8c3 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Zero Bubble Pipeline Parallelism
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8541ea1b-ba59-46a8-bb68-e671a41a5568 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model The Eleventh International Conference on Learning Representations
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b0c617b9-d3c9-4073-acef-6be678b541c4 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0618f57f-7314-4262-9980-4221ca783e0e · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Neurocomputing , volume=
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d816655a-0995-4b58-abc9-dc07541da7b8 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b290ead9-e843-4e39-98c3-86c0e2eac066 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model CoRR , volume =
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 40f89eac-28a2-4d20-ac97-dd7be0fff70a · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b88ca59e-5972-4a83-bbdc-f694e7b59ec4 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2fb7a04c-ed7e-4f65-b250-3d2f3cdec9f7 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0675cc30-f698-4f7b-a43f-2a581f81869a · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model International Conference on Machine Learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation db4ceee0-d641-48ae-a308-7ea2718ef87a · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Chi and Quoc V
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9abd2411-db74-420d-b70e-8a7cf37729b6 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Mishra, M
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 287bada5-97e4-42d6-b1d8-5923cb2fc32f · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71829804-8eb5-4b3b-9223-07cee86e6432 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7378c906-8c6d-4de9-af97-616a35bed4e6 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 70261566-b80e-4fef-9726-02fff60a1959 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2020 , eprint=
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d038485d-2b1e-4b96-bc55-abd63089c28a · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model OpenAI blog , volume=
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c72bc13f-b6c3-4325-971c-e141227ab149 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Introducing
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d65b098b-36cf-4be9-a699-d52f3c935b06 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cfb546d7-68fa-432f-9030-7f2b14d9efae · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0bf731f7-3398-4b15-ba9f-20fd34f092d2 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , pages=
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 50a8ca04-56a5-4bca-bf0a-2cce20ac9bf8 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Proceedings of Machine Learning and Systems , volume=
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 016667a2-f55c-4825-967d-6979cdeb5702 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model and Ermon, Stefano and Rudra, Atri and R
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f90584fe-b2e4-48c4-ad81-4c6fdeead012 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eb2b3c49-41cb-48ec-92c3-be4426768c99 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2061f0fa-9d4c-478f-a403-97d8d1191444 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Advances in neural information processing systems , volume=
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 66be332c-e772-4902-a743-b39e6230d4a1 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model SC20: International Conference for High Performance Computing, Networking, Storage and Analysis , pages=
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d5802070-3356-4ff9-b77c-aa4bed37ef0b · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2021 , eprint=
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 343fcf88-6488-4ebd-9047-1b3d4c35e1cd · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2018 , eprint=
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0d981eb2-ce8b-4065-a9ef-13549d49b813 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Introducing
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5884b14c-0964-4e7f-9605-29c619d00924 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model An important next step on our
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a59c411c-670a-43d2-a3de-b717121bc5a3 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ff9aa778-b593-4df1-99bd-699370e17c60 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2019 , eprint=
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5b33bbff-cd3b-4ecd-a311-da1d912aef9f · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model A Span-Extraction Dataset for C hinese Machine Reading Comprehension
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 36b982e0-225e-479a-a912-de07da21d81c · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2019 , eprint=
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 30f3c0be-83e5-4432-b9ca-c7e92ec8c983 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2023 , eprint=
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 714815ec-0d99-4943-97fc-56a4f3841319 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Measuring Massive Multitask Language Understanding
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 63d22337-1b03-41f1-8275-a38837e26b52 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 560c8b0a-cc5e-4f39-9449-66be3d97a4e8 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Program Synthesis with Large Language Models
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 83b24957-e578-4956-a479-88c2cc0f0c24 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Proceedings of the 28th International Conference on Computational Linguistics
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6de569d8-e0d9-4d1f-b3f2-75b1ad34495d · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b12aa55f-1b87-41fc-aaf1-263fcb54acb7 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Zheng, M
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e71b0ce6-83a9-4840-b899-ab8700b35df1 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model doi: 10.18653/v1/D17-1082
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 173666b5-53b4-4acc-a5ff-3bc3667c1f9f · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Dua, D., Wang, Y ., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4dde6ff6-bee7-4187-acba-68c9901b5dd0 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1f294b04-f991-41c0-959f-36482aea156c · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model LLaMA: Open and Efficient Foundation Language Models
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2cbea312-c172-4b03-acfb-ae8e2a6284ac · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 41f51c55-e9b0-453c-819f-70556e88f170 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Code Llama: Open Foundation Models for Code
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8bfeaf28-b990-40db-8638-a9b485640951 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Training Verifiers to Solve Math Word Problems
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c009aa65-c77a-48bb-963c-b7f2bdba031e · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 82f29955-ba9e-46ba-a460-84dd8ebe4ddd · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c6e68b6-0da3-40fc-941c-a756565d1259 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Evaluating Large Language Models Trained on Code , journal =
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5cee350f-6357-46eb-a740-da3cc5f9225e · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2024 , eprint=
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b9e18150-3dbc-467a-b6f7-b3f8566afedb · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2022 , eprint=
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2ee2c534-a257-4a97-a595-0a26587602f2 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model URL https:// doi.org/10.18653/v1/p19-1472
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5ebd7204-1768-462b-9731-a755fba98c26 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model PIQA: Reasoning about physical commonsense in natural language
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4540ca4b-9026-40f7-9589-4477fc2bde87 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Measuring Mathematical Problem Solving With the MATH Dataset
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9d1bf0d5-a516-43b2-8646-bdaa6c670afa · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model SocialIQA: Commonsense Reasoning about Social Interactions
Reference 120
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0ec704b1-b80c-474d-926c-be559c777325 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 57181a17-8dd6-48ba-9a94-0796bc8b29b7 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3e712cad-c768-4d53-b35f-b976c2983262 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f79f2252-dc25-436a-a8b5-d51dd22c6cd7 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a5fac8ce-eebf-46ca-abde-509c7ab266a4 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 580745d5-4f8f-4a21-96e6-2086f9addc52 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Emergent Abilities of Large Language Models
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 55e99925-3917-4306-b191-06afaafa6fc8 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model B ool Q : Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51399515-0086-42fc-9421-8fb0809ba3f8 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Instruction-Following Evaluation for Large Language Models
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ccdd5d03-3ad5-4a11-a816-c58e3d0188b8 · outbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 69f45e0c-e93f-4ca0-aed9-cb4c9eee727c · inbound
A Comprehensive Overview of Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 139
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 30c646c2-6f98-4096-9bad-108a9a0e0ec9 · inbound
A Survey on Large Language Models for Code Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 162
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7f6a0545-5e4a-4332-8917-2727ccc89313 · inbound
LiveBench: A Challenging, Contamination-Limited LLM Benchmark DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f4246301-8160-4800-bc2b-65b7431f475e · inbound
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 005f53ae-eaf7-4265-a08f-0735f7ea5834 · inbound
Large Language Monkeys: Scaling Inference Compute with Repeated Sampling DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 819a5a45-b23b-49e7-bcb3-582b54e075c8 · inbound
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 29fc9b2b-188c-4ba0-b630-2c8f205a7ad6 · inbound
Optimization Hyper-parameter Laws for Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7f798d6e-bf4f-47ca-b80d-96dc7b51556f · inbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 476a60c7-8422-4d96-af61-7dabfc07b7ce · inbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64169dcd-7543-40cd-b53a-7a7859339e53 · inbound
Efficient Transfer Learning for Video-language Foundation Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d66f7588-e108-4fb3-babe-5fc137180918 · inbound
CKGFuzzer: LLM-Based Fuzz Driver Generation Enhanced By Code Knowledge Graph DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83969801-c87c-4964-b61e-f202f0979f00 · inbound
From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b815b85-e9ab-4701-9624-96aa1d4dbaf3 · inbound
Ultra-Sparse Memory Network DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e4f6b5-10d6-4f12-8071-c0bb3b1d3877 · inbound
InstCache: A Predictive Cache for LLM Serving DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e2aa33b-0680-4a83-b178-c13104739de5 · inbound
Masala-CHAI: A Large-Scale SPICE Netlist Dataset for Analog Circuits by Harnessing AI DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1830ac7e-5fec-42f3-aa80-44318e9a7f4e · inbound
Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc9b9737-2fcb-4270-85f3-f079faa7112e · inbound
Traditional Chinese Medicine Case Analysis System for High-Level Semantic Abstraction: Optimized with Prompt and RAG DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcda7819-e101-42bc-a016-a0a466e36dce · inbound
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be961e9a-07a7-40ec-8c96-0e20a6df2852 · inbound
Hiding Communication Cost in Distributed LLM Training via Micro-batch Co-execution DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 936a3f44-63da-45ab-8f05-074f56abeed3 · inbound
From Laws to Motivation: Guiding Exploration through Law-Based Reasoning and Rewards DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eb31719-7604-4112-aecb-66e81e67e924 · inbound
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fe4c1fb-76cf-435f-b423-bab0dc224dc5 · inbound
From CISC to RISC: language-model guided assembly transpilation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5de252b6-756c-4638-89d1-0ed458182e27 · inbound
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced3d7ac-354d-4ae5-8102-0cd4742e08dd · inbound
CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed97d7b7-c440-456b-8930-945c2dd856c0 · inbound
CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4e34c8c-0d60-4e6d-be5b-a520734ca693 · inbound
CheckMate: LLM-Powered Approximate Intermittent Computing DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 593b19d8-e4eb-4572-9463-5f4c6b4afb31 · inbound
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs? DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f38e194c-d62f-41c8-a5a4-215fa8e5eac0 · inbound
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae8baa3f-74bd-410d-84d2-f59578610929 · inbound
Towards Adaptive Mechanism Activation in Language Agent DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 836485a9-f86c-4bb0-894b-ede0812430d5 · inbound
Yi-Lightning Technical Report DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97adbfeb-aef9-4d28-973e-33d255b53012 · inbound
FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f33968a4-7f52-4a65-8c53-7783935bb369 · inbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 839629ed-f3af-48db-93f0-345442fd0851 · inbound
DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e92aa4d9-474d-4f0c-9bc9-6f4ec3430183 · inbound
Weighted-Reward Preference Optimization for Implicit Model Fusion DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 116d738e-abe0-4088-a56b-ba41f9d8362d · inbound
Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e65f076-5237-4001-b76d-4f6aa739ba71 · inbound
Can Large Language Models Effectively Process and Execute Financial Trading Instructions? DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5d7a32c-1d43-4e5a-bc5c-31496e5db9f8 · inbound
Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfd63bb6-096a-47a6-903f-f137a0032516 · inbound
RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df3854e-ffd8-4a08-8f13-9f1bd4153b3c · inbound
Breaking the Stage Barrier: A Novel Single-Stage Approach to Long Context Extension for Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aecb592-5ef2-4213-ad46-45b5c9a348e5 · inbound
Federated In-Context LLM Agent Learning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 238a6e9d-fc4a-49c8-a83d-c40f645908ab · inbound
DialogAgent: An Auto-engagement Agent for Code Question Answering Data Production DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a72d15b6-fd72-471e-a51d-cf82ec857e32 · inbound
Towards Wireless Native Big AI Model: The Mission and Approach Differ From Large Language Model DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e555fa4-3e94-4e0a-aedf-7c0794b6d997 · inbound
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c225df22-e0e7-48de-b586-3e2de20edeba · inbound
SCBench: A KV Cache-Centric Analysis of Long-Context Methods DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4ea3613-51e5-455d-a2ae-f5df95c7c7f1 · inbound
Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a1cdf2d-9a12-4467-a7eb-c2723f1a61f5 · inbound
RecSys Arena: Pair-wise Recommender System Evaluation with Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13349e22-1b1b-41c2-bf16-26644b466d81 · inbound
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers? DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47f656c3-4bbf-4ead-b092-3adb7a000df4 · inbound
More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b4ac8c-239c-47d9-9d40-2a17a8fa8cc8 · inbound
UITrans: Seamless UI Translation from Android to HarmonyOS DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 136fcb09-72cc-46bb-addc-de82990e8d10 · inbound
BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4a18125-189f-4c47-ae75-7b88649711c7 · inbound
Language Models as Continuous Self-Evolving Data Engineers DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 172557b3-2085-4393-afdb-a81242135a36 · inbound
CodeV: Issue Resolving with Visual Data DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b76e7ab-593f-46ea-a41d-0730763377ac · inbound
A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc5c1223-b083-41f4-94d7-3c48f1089703 · inbound
YuLan-Mini: An Open Data-efficient Language Model DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c3a3da0-9b7e-444b-9ea7-9e5134f53ad8 · inbound
In Case You Missed It: ARC 'Challenge' Is Not That Challenging DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db60e678-aaca-40c2-8c99-d91446af1d55 · inbound
Multi-matrix Factorization Attention DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73a20849-83dc-4539-b7f4-56290923b57d · inbound
A Survey on Large Language Model Acceleration based on KV Cache Management DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8202b31-bbbc-4914-9ae9-4795852d7478 · inbound
BaiJia: A Large-Scale Role-Playing Agent Corpus of Chinese Historical Characters DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a489ebb9-8328-4c56-8ea7-5a33c9091de7 · inbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86ee5653-5959-45a3-898e-d4776ea430f0 · inbound
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b30ab26-4207-4482-920c-93f59be5f71e · inbound
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c5be030-5340-4c47-aee8-273f73006960 · inbound
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ab6639-1f2e-4950-8bb6-cc8afdd71ba8 · inbound
GroverGPT: A Large Language Model with 8 Billion Parameters for Quantum Searching DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bcd55fe-8caf-4238-a58d-eb4da5729f18 · inbound
SPDZCoder: Combining Expert Knowledge with LLMs for Generating Privacy-Computing Code DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23b738d-afc2-495b-97df-a64d1b3f2a41 · inbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac5e5624-377a-4491-b372-b1577c3b36bb · inbound
AgentRefine: Enhancing Agent Generalization through Refinement Tuning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e05187-17ee-4e58-b815-c572ab96726e · inbound
MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df2668ec-012f-4a7c-a354-8112ef86a33f · inbound
Long Context vs. RAG for LLMs: An Evaluation and Revisits DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c109257c-7a89-4c89-a680-499c5e8e0782 · inbound
Instruction-Following Pruning for Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af31485e-0ef6-4a8c-a61f-2093a3400e0c · inbound
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37b81e57-e86d-4adc-b0ed-fe200c8079b9 · inbound
Ultrasound-QBench: Can LLMs Aid in Quality Assessment of Ultrasound Imaging? DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb8a9ca5-6116-48e3-a3c7-ce21d45a457a · inbound
Beyond Pass or Fail: Multi-Dimensional Benchmarking of Foundation Models for Goal-based Mobile UI Navigation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cfe15f4-f9c0-470e-b5f0-f969146b7861 · inbound
MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbc66737-a108-421f-a3be-43bd22820999 · inbound
Powerful Design of Small Vision Transformer on CIFAR10 DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6119b63-645e-4fc5-aba1-bc75ed39954d · inbound
LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d8f7dc2-706e-425d-b3af-3c2f4f3db144 · inbound
OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74bc134c-f68a-4b82-82ee-60932b2f7f89 · inbound
Sympathy over Polarization: A Computational Discourse Analysis of Social Media Posts about the July 2024 Trump Assassination Attempt DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3775be2-59b5-40c0-89c5-ac32b53d0c83 · inbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6a6b5af-4d9b-4ec7-8c9c-90751895dddf · inbound
SR-FoT: A Syllogistic-Reasoning Framework of Thought for Large Language Models Tackling Knowledge-based Reasoning Tasks DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc3be2fc-d092-457e-a576-76e2976d1c2c · inbound
Panoramic Interests: Stylistic-Content Aware Personalized Headline Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950a6b8e-890d-4465-a941-3e9a104100a9 · inbound
Pseudocode-Injection Magic: Enabling LLMs to Tackle Graph Computational Tasks DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a76af209-837b-4033-9eb1-9254dcff67e5 · inbound
Mixture of Experts (MoE): A Big Data Perspective DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dffec26-ad04-42ff-8152-7cac27223013 · inbound
Scaling Inference-Efficient Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2f18128-890d-4c80-8908-13c4a6f87e8f · inbound
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7071e68a-a6a1-400c-b327-0df162bf3b31 · inbound
OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377ae90a-b793-43db-9794-e3700844ad21 · inbound
Position: AI Scaling: From Up to Down and Out DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2558f0b6-7fca-402f-a2b0-af24642dc7cf · inbound
COFFE: A Code Efficiency Benchmark for Code Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d1ce5d2-9612-4ef2-98cb-35b4a89334a5 · inbound
GistVis: Automatic Generation of Word-scale Visualizations from Data-rich Documents DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e836b4a3-3e60-4d92-ba62-c69efb74df8a · inbound
Steel-LLM:From Scratch to Open Source -- A Personal Journey in Building a Chinese-Centric LLM DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1b0cd1-8d68-490a-884a-4d81a6bd0aa5 · inbound
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae3c64e5-41b6-4ea1-b54c-9c533c2c8c2c · inbound
CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c0eb8d3-4657-462d-9d52-f98af83a59b4 · inbound
Memory Analysis on the Training Course of DeepSeek Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8363b8c-7845-4cd3-a8ea-02c502c0c98e · inbound
TransMLA: Multi-Head Latent Attention Is All You Need DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80d2a90f-0307-4119-bded-4c4af08d5431 · inbound
Explanation based In-Context Demonstrations Retrieval for Multilingual Grammatical Error Correction DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f315a9ca-f5f2-485d-a0d1-0ab920c721ea · inbound
Universal Model Routing for Efficient LLM Inference DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3545b50a-5b7b-4782-8228-be7689c494c1 · inbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00e310b6-3e27-4074-b52c-906ae8674f8a · inbound
You Do Not Fully Utilize Transformer's Representation Capacity DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fea2f38a-a783-4d78-96f7-373c093114c6 · inbound
Climber: Toward Efficient Scaling Laws for Large Recommendation Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8f2a7ec-b82b-440f-8e1d-3fed18c87108 · inbound
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 85c2f796-d36f-44bb-8a1a-21c61db8f5e2 · inbound
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.