Pith. sign in

Paper Citation Record · LEDGER

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

As of 13 August 2026, this Paper Citation Record lists 100 of 150 outbound references and 100 inbound Pith citation observations for arXiv:2405.04434.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.04434 v5

Coverage vector

measured 100 of 150 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T05:36:26.207359Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 100 of 431 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:10:00.206774Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 150 outbound references displayed

  • verified exact27
  • verified fuzzy39
  • unresolved13
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch20

External citation measurements

5
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 944e2709-9885-415d-a193-0c6e364518da · outbound

This paper cites Llama 3 model card.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Llama 3 model card

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.279750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:449c2eed424f9d27b40f7e6b8c6a544ca9fba77f38860180911e572a46966415

Observation d24f057b-ebf7-44bd-92b2-c88a53efb6c3 · outbound

This paper cites Introducing Claude.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Introducing Claude

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.292283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:3a12c0d8a4f0da99d3268bc3c72b731783e08b17a228796be90efcd7fa88c41b

Observation 0b4d7b31-7be9-435f-bf76-02761533340a · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Evaluating Large Language Models Trained on Code

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:36:26.993658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:8aa47f68b81faa9b6e9af9c2ab3b4bb3cd6dfdc9aa871dbb6fa1b5c2dad792be

Observation 4086f8d8-ed30-4bb6-a04f-b87b6be6a561 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:50:20.439308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:9cb9021e10e7f9aa0b50b4dee5af74530b640ac75fd7e38a5ddbdefa74ae281d

Observation 344e9eae-3bdd-46b0-a244-ed55c48ddaa6 · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.300498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:ad02830208f7d2fa23b0470c70e35228f1aec795dde8ff93e19e433862fd8a26

Observation 96908b73-98f0-492f-8f0c-6b23e1a08d19 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:08:06.759074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:e4cb51b094395fa7da5df281b0288ef9905f894929b412db62fc532c425c3f22

Observation ef25b505-3aec-4e29-baae-758878d7d3e7 · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:57:11.134962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:9234e97298017a959f5ddaeeb24a9c1dd7aee5839cf0ce6adb1f139dc56ceb45

Observation 3d03e31e-2a1c-4b25-9cb5-8ee700f0d844 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:36:27.244210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:f43057cc6da94222204a556d947c175428749921121190b09c19516f75ebadf3

Observation d5e4d8bf-ae36-4be6-a637-36245066b918 · outbound

This paper cites Introducing gemini: our largest and most capable ai model.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Introducing gemini: our largest and most capable ai model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.304949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:8e46ded80a82090625d9523dfefc7c7eff6342bde034f390c1a1c1d629d6559d

Observation dc3f1c97-9b26-467f-84f1-2c56e4027b7d · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.309184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:89f53a4c0f5510949f72056297937d04ad8d68b3a64d3f0699dc59ef3ce596dd

Observation bdeae05e-a0f6-4ad5-b841-d56d8c48998c · outbound

This paper cites Hai-llm: 高效且轻量的大模型训练工具.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Hai-llm: 高效且轻量的大模型训练工具

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.313353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:aaf17be611c99163102cee9320a606ba2436ce7dc105b872917fd8da4ceb1d63

Observation b6d19c84-6bc0-4145-8471-c17e03e40217 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:26.480168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:99314c6b3d0c76506200424ed98689e5a84e5bdd9f39f87ca7ca6e5e7ba71e8e

Observation d2c283fc-0a17-4ec6-9904-b77b004536b1 · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:27.180015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:ecddb928f068390f89d1238c47fc6e24fa42093223261890cdf8c06dd5daa4b7

Observation a2cdc000-6e7d-4212-9592-b9a058011bda · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.323385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:3644aa58e9cdb5b9327aa24e4c7195ab67b64835cfd6b1c92f642fd5ae603a9a

Observation d3bd00e7-d85d-40f6-8ebf-5d92140d32d2 · outbound

This paper cites Lepikhin, H.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Lepikhin, H

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.330832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:853956a659b3977388428cfc33e7b640107ab0114b2353417dc36eab3ad3d5cb

Observation 04265ffe-e7a5-4734-8d3f-de4c5c7f9c34 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model CMMLU: Measuring massive multitask language understanding in Chinese

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:01:11.049549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:4d45fed8ca3a691e62b727c29f2a70e714aaef252c882a680b7b7c02905bfa7b

Observation 2d280c8b-1874-4aa5-b644-3ac0b629051c · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.338391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:424f79a53a4140d3304335a30518dabdfb2d775d1d1f7d7d7d497d40fd72c1db

Observation 53571ff3-e138-4594-b644-bb15a90a1bf4 · outbound

This paper cites Cheaper, better, faster, stronger: Continuing to push the frontier of ai and making it accessible to all.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Cheaper, better, faster, stronger: Continuing to push the frontier of ai and making it accessible to all

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.343169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:200527f10f547d9e39a7d4412c950d8def2d905230460d0354681c45240d85fc

Observation aeb3ac18-747d-46b9-b40e-790f1a9e076b · outbound

This paper cites Introducing ChatGPT.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Introducing ChatGPT

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.347040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:ff5d9b331dac98e4e5112d3c5ab0746350f86635c043d06ea6ed471bec23d05c

Observation 49e6a07a-77a9-4260-8f4b-9ff7c384fd00 · outbound

This paper cites GPT-4 Technical Report.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model GPT-4 Technical Report

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:36:27.001721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:4dc764bf0ca82f43aad350cc1e63ff69b4128c86da104a31f7d09b8c9bc9708b

Observation 8f7baefc-2620-41dc-a4b0-bc34b81c9421 · outbound

This paper cites Ouyang, J.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Ouyang, J

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.351179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:890c13a73e2ace8f82771ba02b71a86fec05d481b4aebd0688aedbfd12c81760

Observation 3d869ecb-9f7c-4d69-a960-5970ff095217 · outbound

This paper cites Rajbhandari, J.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Rajbhandari, J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.358281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:7d1359be22a719cf7f6184dd44a0200044b1dde413c6ef8d991cef6359e23fff

Observation 479535a2-023f-4628-b1d3-0ff138afe6f2 · outbound

This paper cites Riquelme, J.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Riquelme, J

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.362583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:139597b67aa1f896c9dbdbdc1091ab0be5eb0dd461638f24464d6b6aeaffffef

Observation c86154ed-fd88-420d-a570-976f0aad53f1 · outbound

This paper cites Sakaguchi, R.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Sakaguchi, R

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.366755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:6793ba3975e23655bf53748b3387840eb0218afba2aa7422de9821eccda77459

Observation 4e883403-25c9-4fc8-a2af-19cb1add8cfa · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Fast Transformer Decoding: One Write-Head is All You Need

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:36:26.710392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:dde6b3172c9a959caa4784250073277e5406d9861a6d942b3f695d33f7ac70d1

Observation 55539873-c0a8-4081-a92e-cdfc5685b185 · outbound

This paper cites Shazeer, A.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Shazeer, A

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.370849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:4503cefdcde20d81ecbbe4a4bcd84b843a7e4c1be7d2a4de0297ff95b1d5036e

Observation 4881d738-2bb2-439f-8aea-ab08d670a168 · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.374627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:0c58646c18ff8c444c6dc4d7cfcc556fedaa77b389e840da1d24a1be10043d29

Observation 478b12d8-2347-4e75-8b23-d8849a8419ab · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.380630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:92f95e20d758fc0dc884531f0eca491d4e2571e70867838bfe628e5fdba5e381

Observation 88c45b90-055d-412a-ab31-cca9c7459dd7 · outbound

This paper cites Vaswani, N.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Vaswani, N

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.384763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:22cb4cfcaa192fc40b2170ecc7c39e1a444fa9d08e8efde9c22ab9a6181d99fb

Observation 18a412c8-19c2-4fae-9ea7-b980700f4f28 · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.393266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:6dc7b9e6a6b4595a80d60dc8109e08dc2fded70848efcd869d593d01d3d0218d

Observation 14b49881-a09e-490f-b4a6-f518a0782182 · outbound

This paper cites Atom: Low-bit Quantization for Efficient and Accurate LLM Serving.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:26.468374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:bf99f3d236cc50bd1072e459d40afb97886c45a8b68baba7cba8349ee4167a85

Observation 41ebf1dd-bfee-4e1c-bd44-99e759ecc040 · outbound

This paper cites Zheng, W.-L.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Zheng, W.-L

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.397892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:5c39489d8df95fe57a20bbdfce085e6166b4e7e1256951c69f7f8fb79ef64b1e

Observation 682391b0-dffb-4868-9f4d-e9fca687ae1b · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.405272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:32e7e5e6037a80de3f89285c5b4174a51f54b2f7195fb65854acd59461716ea8

Observation 317e65dc-af55-41c5-8adf-0ae42109b8c3 · outbound

This paper cites Zero Bubble Pipeline Parallelism.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Zero Bubble Pipeline Parallelism

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:27.188875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:d6cc329a2be34ddfad51de01978c2f2d6488eb329f6c003b01c090beda06f218

Observation 8541ea1b-ba59-46a8-bb68-e671a41a5568 · outbound

This paper cites The Eleventh International Conference on Learning Representations.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model The Eleventh International Conference on Learning Representations

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.409369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:edc6b38c13d3eea84b387ba812659a642f0ed7aec90980e12691568a6dc2830d

Observation b0c617b9-d3c9-4073-acef-6be678b541c4 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:27.229249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:54b438972dfe1ec326d89fba06148981b5d16cda80dfc5109bf0f1b56500fbd6

Observation 0618f57f-7314-4262-9980-4221ca783e0e · outbound

This paper cites Neurocomputing , volume=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Neurocomputing , volume=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.412948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:532b1acbfbbd6541a9c04af3956d496334c58964a76ba42439b1d3a336c4fa74

Observation d816655a-0995-4b58-abc9-dc07541da7b8 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:53:59.644169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:377a35e6004d87fedd717c1253a402856bd8de1ff7be818cc26d05436f678504

Observation b290ead9-e843-4e39-98c3-86c0e2eac066 · outbound

This paper cites CoRR , volume =.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model CoRR , volume =

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.416901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:705c0ceb0f0b58a87537ba79f83f85758c16ae9103503b433e5e5500c6eabfe0

Observation 40f89eac-28a2-4d20-ac97-dd7be0fff70a · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:00:08.752459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:2d2e7cbf7aa0ad2e4198b2f9f45b2fe0df99d3acc4e69d2dfcc80a30ef6d232c

Observation b88ca59e-5972-4a83-bbdc-f694e7b59ec4 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:19:36.299641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:6e062a79278f2fa62386253bc19a35772ca7f7af3c09d911b8d5b07c03db372e

Observation 2fb7a04c-ed7e-4f65-b250-3d2f3cdec9f7 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:48:28.215201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:fd3683166cdbea0b3c067b0b848adcf02f2e39bc97fd829f9057dd35642f81f7

Observation 0675cc30-f698-4f7b-a43f-2a581f81869a · outbound

This paper cites International Conference on Machine Learning.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model International Conference on Machine Learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.420971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:20198edce45ffab5d51ad6c753fa8b3ff9cb5e2b4e89e5895b9e68a5996910c0

Observation db4ceee0-d641-48ae-a308-7ea2718ef87a · outbound

This paper cites Chi and Quoc V.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Chi and Quoc V

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.425082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:198e56160a834d612408483f27be0d9e5cce3d3731cf532f0d594d9830f3b2eb

Observation 9abd2411-db74-420d-b70e-8a7cf37729b6 · outbound

This paper cites Mishra, M.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Mishra, M

Reference 74

Resolution
metadata mismatch
doi, observed 2026-05-11T05:36:26.552637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:84cd12a0ed64e1e056bce2d1c3ace281bd73ebed2debf3774bd4b1102d77c107

Observation 287bada5-97e4-42d6-b1d8-5923cb2fc32f · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:46:39.773471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:dc1e6b8510e00b507d76a27fd6216c6b428ff6976e54f78f9d28c4d2d8acbf8b

Observation 71829804-8eb5-4b3b-9223-07cee86e6432 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:07:53.977132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:ab93d754e573d1df401f793f759901e32c6800df5650c9dc605ce6cf64d07644

Observation 7378c906-8c6d-4de9-af97-616a35bed4e6 · outbound

This paper cites TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension

Reference 77

Resolution
verified exact
doi, observed 2026-05-11T05:36:26.608683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:8010ef3e0839e6cb072f11bcfcf59e940869a5a89be1b1c873d849a7e7394e89

Observation 70261566-b80e-4fef-9726-02fff60a1959 · outbound

This paper cites 2020 , eprint=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2020 , eprint=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.431126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:14f3c98219626e3ddc554b4a8718dc46c084d571033abb3b5eb3944b3e0834f3

Observation d038485d-2b1e-4b96-bc55-abd63089c28a · outbound

This paper cites OpenAI blog , volume=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model OpenAI blog , volume=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.437350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:cce43a4271f008c88ea9269441e822203c637bfbaaf974e794883c3dca15e75e

Observation c72bc13f-b6c3-4325-971c-e141227ab149 · outbound

This paper cites Introducing.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Introducing

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.442304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:70a41d47d4ff1b47dee32a463f9f43636ccdde085acd5fe3bbf99781bed02b9f

Observation d65b098b-36cf-4be9-a699-d52f3c935b06 · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.459278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:da87c57d8c83c35982d92d5e7ea908642ace892cd0a9c14a4a547d8adce695e4

Observation cfb546d7-68fa-432f-9030-7f2b14d9efae · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 82

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:27.124128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:66486588dbc64074bdc0f555f5e4ecbc60a79b8de4c5c0a910116a1a8d1bc2e1

Observation 0bf731f7-3398-4b15-ba9f-20fd34f092d2 · outbound

This paper cites Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , pages=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , pages=

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.464047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:fea840c7bf661c6030fb9a5ef410b79b51fa720617e9f5a03d0cb8f4127bdf65

Observation 50a8ca04-56a5-4bca-bf0a-2cce20ac9bf8 · outbound

This paper cites Proceedings of Machine Learning and Systems , volume=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Proceedings of Machine Learning and Systems , volume=

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.468292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:0a7a7d0979f7e88fb1b0603332f0dc2053be2c318ab35b8ab5ddd8f3c1b0e762

Observation 016667a2-f55c-4825-967d-6979cdeb5702 · outbound

This paper cites and Ermon, Stefano and Rudra, Atri and R.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model and Ermon, Stefano and Rudra, Atri and R

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.478514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:594571b07a8c0f2d776d45108b8f0241fe254c1fa857b1e500d592348bdb01d7

Observation f90584fe-b2e4-48c4-ad81-4c6fdeead012 · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 86

Resolution
parse uncertain
raw_fallback, observed 2026-05-11T05:36:27.483727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:54b26427d76703309baa7df337f6f9683a3b3ca5c11c49932d2bd0e532da3fcd

Observation eb2b3c49-41cb-48ec-92c3-be4426768c99 · outbound

This paper cites Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.487859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:d22cbd18b9bd45ae89bea9d1c2f390817e08da69109a0664591f362ddeddb33b

Observation 2061f0fa-9d4c-478f-a403-97d8d1191444 · outbound

This paper cites Advances in neural information processing systems , volume=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Advances in neural information processing systems , volume=

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.492048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:f44207c0c13706be9db3b67cc8acc99b582f7043f90bbda1a34a34df40e2e056

Observation 66be332c-e772-4902-a743-b39e6230d4a1 · outbound

This paper cites SC20: International Conference for High Performance Computing, Networking, Storage and Analysis , pages=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model SC20: International Conference for High Performance Computing, Networking, Storage and Analysis , pages=

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.496480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:736b9c68d81779973891bdf0c01a26ffcc07738886d6f6a407357fd77347b9f8

Observation d5802070-3356-4ff9-b77c-aa4bed37ef0b · outbound

This paper cites 2021 , eprint=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2021 , eprint=

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.500684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:5d10c830bd6b47b006e78ff6b7feff9dd237054be9cbd749ae4dc324c6810317

Observation 343fcf88-6488-4ebd-9047-1b3d4c35e1cd · outbound

This paper cites 2018 , eprint=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2018 , eprint=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.508794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:d70925a2a43be74e4f30ecb7930438a5843874a86e5749760dc2974cbaec6c2c

Observation 0d981eb2-ce8b-4065-a9ef-13549d49b813 · outbound

This paper cites Introducing.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Introducing

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.513065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:7953912e559dab693d2bbdc321e8b328bc10b5e87a4f09b25d071ae29dfdab55

Observation 5884b14c-0964-4e7f-9605-29c619d00924 · outbound

This paper cites An important next step on our.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model An important next step on our

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.517786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:ee2cd11d715c0ae3de183078d4c5c1d62527aa69f05e7e1e11170fbc9c8cdb01

Observation a59c411c-670a-43d2-a3de-b717121bc5a3 · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.525441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:4962fb5825fb2db9e22388ac50e0653b523795821d7125b1e2702549dd19a34f

Observation ff9aa778-b593-4df1-99bd-699370e17c60 · outbound

This paper cites 2019 , eprint=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2019 , eprint=

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.534658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:13e4199ecd2dd6a335ef0a88629d84077f2b269690926af1137d47e91e256f8d

Observation 5b33bbff-cd3b-4ecd-a311-da1d912aef9f · outbound

This paper cites A Span-Extraction Dataset for C hinese Machine Reading Comprehension.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model A Span-Extraction Dataset for C hinese Machine Reading Comprehension

Reference 96

Resolution
verified exact
doi, observed 2026-05-11T05:36:26.545356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:e9e3e6db26128ca3eee39dbba4504d753827c488ca98ffe936209109a9ad1552

Observation 36b982e0-225e-479a-a912-de07da21d81c · outbound

This paper cites 2019 , eprint=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2019 , eprint=

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.542354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:3c4af1351af3e3a0a2fbe206652328da899641f588aa82df3307b21912abfd51

Observation 30f3c0be-83e5-4432-b9ca-c7e92ec8c983 · outbound

This paper cites 2023 , eprint=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2023 , eprint=

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.550893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:de0ed933436c57666f72c6015f9b0673ae50735bf2d56662197d585113ebe56f

Observation 714815ec-0d99-4943-97fc-56a4f3841319 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Measuring Massive Multitask Language Understanding

Reference 99

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:26.890165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:0814784605f7b58641f6f815ab1e4335a19eacfb3d040f3b6b757e9db292d412

Observation 63d22337-1b03-41f1-8275-a38837e26b52 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:15:24.133279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:ac86998d1226f89122963a514d236d121342fe88147781a9e2b330d998e3c13b

Observation 560c8b0a-cc5e-4f39-9449-66be3d97a4e8 · outbound

This paper cites Program Synthesis with Large Language Models.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Program Synthesis with Large Language Models

Reference 101

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:26.902671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:f84c4c203285ddc0cd0f7e6ff1b5e5920f0ac81ea6c8a36d1cce928771634d4b

Observation 83b24957-e578-4956-a479-88c2cc0f0c24 · outbound

This paper cites Proceedings of the 28th International Conference on Computational Linguistics.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Proceedings of the 28th International Conference on Computational Linguistics

Reference 102

Resolution
verified exact
doi, observed 2026-05-11T05:36:26.650969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:3b72eb2362661779490f04f670ee92c39c7001b259ef86af1ee2b0740fff9788

Observation 6de569d8-e0d9-4d1f-b3f2-75b1ad34495d · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 103

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.555618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:156872ee40c22ad67049212b89a4ed935af0e35635de94ed9d3818cc654dccee

Observation b12aa55f-1b87-41fc-aaf1-263fcb54acb7 · outbound

This paper cites Zheng, M.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Zheng, M

Reference 104

Resolution
verified exact
doi, observed 2026-05-11T05:36:26.677247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:2e326703d9cbb1009637b541d16ba387131943a5c0fb6d0017289d041f03aebf

Observation e71b0ce6-83a9-4840-b899-ab8700b35df1 · outbound

This paper cites doi: 10.18653/v1/D17-1082.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model doi: 10.18653/v1/D17-1082

Reference 105

Resolution
verified exact
doi, observed 2026-05-11T05:36:26.685567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:dd69ad56ad8b2622abd8e57edc285c068d025da469fcfbca64f16638a65c9bb0

Observation 173666b5-53b4-4acc-a5ff-3bc3667c1f9f · outbound

This paper cites Dua, D., Wang, Y ., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Dua, D., Wang, Y ., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M

Reference 106

Resolution
metadata mismatch
doi, observed 2026-05-11T05:36:26.691153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:95fab50ee6bae18f79a9188867854995b033b435b082070a3f1e0d7d86752d04

Observation 4dde6ff6-bee7-4187-acba-68c9901b5dd0 · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 107

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.559634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:a73c82b4d6cb7565b12ce6ad3c6e5c91b59ccb79df7ed378e9f3b091849e3a46

Observation 1f294b04-f991-41c0-959f-36482aea156c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model LLaMA: Open and Efficient Foundation Language Models

Reference 108

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:26.967637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:ac3c371f3beb1bb5252496365d13f29007f0bd95c5574430dd1e24e4a6520fef

Observation 2cbea312-c172-4b03-acfb-ae8e2a6284ac · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 109

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:26.460929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:843c56fb2683ef7c96c90d803825aff394c311fe87a75e9df8664e0da19ab28d

Observation 41f51c55-e9b0-453c-819f-70556e88f170 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Code Llama: Open Foundation Models for Code

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:36:26.977652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:1cd57dde47c700e64e970058e42e5ed68a417f2b085a98a98380b71df69c436a

Observation 8bfeaf28-b990-40db-8638-a9b485640951 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Training Verifiers to Solve Math Word Problems

Reference 111

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:26.986981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:2ca1d5f9763b93f253752aaa9b36ac6c9566d2fa055b5b5424d7f411b30a2dc1

Observation c009aa65-c77a-48bb-963c-b7f2bdba031e · outbound

This paper cites an unresolved cited work.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Unresolved cited work

Reference 112

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:36:27.564098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:9cf151c0607c13afdf14d83d345e909b48f99b85ab996f6f9f3a6e23ccb390b0

Observation 82f29955-ba9e-46ba-a460-84dd8ebe4ddd · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:c82baae4875811dcfaa3cca415dd485188460c02c2f87a642828747863d8394a

Observation 8c6e68b6-0da3-40fc-941c-a756565d1259 · outbound

This paper cites Evaluating Large Language Models Trained on Code , journal =.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Evaluating Large Language Models Trained on Code , journal =

Reference 114

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.570080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:552ece554138302b9d5045595866b64ed782e33a2dd46fe0afb7e3ed4ac200ca

Observation 5cee350f-6357-46eb-a740-da3cc5f9225e · outbound

This paper cites 2024 , eprint=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2024 , eprint=

Reference 115

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.574281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:cd7d8b1c270f894d904e939e6531e3a6e68620bb3096cd4bf809331d8f7d44ab

Observation b9e18150-3dbc-467a-b6f7-b3f8566afedb · outbound

This paper cites 2022 , eprint=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model 2022 , eprint=

Reference 116

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.585467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:d64ebbc44ae9d6f96781db98995eab0888182a22c06d078530df9ce1c5866838

Observation 2ee2c534-a257-4a97-a595-0a26587602f2 · outbound

This paper cites URL https:// doi.org/10.18653/v1/p19-1472.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model URL https:// doi.org/10.18653/v1/p19-1472

Reference 117

Resolution
metadata mismatch
doi, observed 2026-05-11T05:36:26.514140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:760029eae1b1cd2857589544ff6c59c4d5f21d2e08f5ce125f237bfd6dce0623

Observation 5ebd7204-1768-462b-9731-a755fba98c26 · outbound

This paper cites PIQA: Reasoning about physical commonsense in natural language.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model PIQA: Reasoning about physical commonsense in natural language

Reference 118

Resolution
verified exact
doi, observed 2026-05-11T05:36:26.597352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:6ad1f051690403720e78a611e0098f346fa308c341e88e8a0e2e166184bc44d2

Observation 4540ca4b-9026-40f7-9589-4477fc2bde87 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Measuring Mathematical Problem Solving With the MATH Dataset

Reference 119

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:27.022721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:f7e6db1c47adf293b74eb986ab7a431f1ef7eb26ba3f5e825418cbffcdd8ce8d

Observation 9d1bf0d5-a516-43b2-8646-bdaa6c670afa · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model SocialIQA: Commonsense Reasoning about Social Interactions

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:22:26.192007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:961e430f812cb7d7f3015ebc4df8aeab46c5350bb75cc29848256785dc8071b0

Observation 0ec704b1-b80c-474d-926c-be559c777325 · outbound

This paper cites Evaluating the Performance of Large Language Models on GAOKAO Benchmark.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:28:32.509177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:a4dcc1ede41bbb58b973e9383c1177641592e8d262c58cd6f8c19655e92c8259

Observation 57181a17-8dd6-48ba-9a94-0796bc8b29b7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 123

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:27.038705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:63199c9feae70930fc1ad062ded547cbc9ebccc07865bb891a7cee10b3496bf3

Observation 3e712cad-c768-4d53-b35f-b976c2983262 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 124

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:00:53.610624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:4a1efb1408302127b453c4ca56cbb718e7bd36e629086f65fc6049d08007ecd3

Observation f79f2252-dc25-436a-a8b5-d51dd22c6cd7 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 125

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:36:27.063417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:9727fd1754c0a9b1ef0759b4d74bd6fac8d12c4b3b82e718445b4fb0185b3e60

Observation a5fac8ce-eebf-46ca-abde-509c7ab266a4 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 126

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T11:13:05.519270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:c81bd6bdd203c2975a20f29890276793c6ba95697453ad0783f02dc548bc0b40

Observation 580745d5-4f8f-4a21-96e6-2086f9addc52 · outbound

This paper cites Emergent Abilities of Large Language Models.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Emergent Abilities of Large Language Models

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:38:38.557341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:c81bf1d1ae240133e7273308298e5cd517bbdcaced1bce8031452b49197d385d

Observation 55e99925-3917-4306-b191-06afaafa6fc8 · outbound

This paper cites B ool Q : Exploring the Surprising Difficulty of Natural Yes/No Questions.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model B ool Q : Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 128

Resolution
verified exact
doi, observed 2026-05-11T05:36:26.628157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:195116621c28a339469ee68b4384f1bfe50e1be2cb40e399c2f1cc5ff8fa7841

Observation 51399515-0086-42fc-9421-8fb0809ba3f8 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Instruction-Following Evaluation for Large Language Models

Reference 129

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:36:27.106762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:86a5b3f8e051592d659b4878955f4dcc16b868f277512ae16fe10001cd7852bd

Observation ccdd5d03-3ad5-4a11-a816-c58e3d0188b8 · outbound

This paper cites Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 130

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:36:27.590264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:942c875aac91e918e69b11d3c847f0f1bee92ad567d76f22a520c2281009224f

Pith citing papers

Observation 69f45e0c-e93f-4ca0-aed9-cb4c9eee727c · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 139

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:28:39.459444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:58fdc14556856bde824c37cb66ee40efa069da433cd7d355fb01c56c24a96922

Observation 30c646c2-6f98-4096-9bad-108a9a0e0ec9 · inbound

A Survey on Large Language Models for Code Generation cites this paper.

A Survey on Large Language Models for Code Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 162

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:18:06.399220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T20:18:06.304134Z digest=sha256:0b404890f909120504a01fe17116995eeb3c9351ac50d3a1db8a4c26caf8363c

Observation 7f6a0545-5e4a-4332-8917-2727ccc89313 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T04:48:26.465022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:dd3373240e3253f4dca2cd9855faea96d5e1a48f9168aff5687c4028e5cee637

Observation f4246301-8160-4800-bc2b-65b7431f475e · inbound

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision cites this paper.

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:45:36.386458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T19:45:36.337956Z digest=sha256:d2b0a8ec4ed780097b2f0f5f6f77778940ffe5757feef30c99ee849c21aca255

Observation 005f53ae-eaf7-4265-a08f-0735f7ea5834 · inbound

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling cites this paper.

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T04:42:23.429350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T04:42:23.297389Z digest=sha256:92630d96093127f364527e4d272134f53f6c32137e112148b572e60a60c51728

Observation 819a5a45-b23b-49e7-bcb3-582b54e075c8 · inbound

Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts cites this paper.

Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:55:26.068326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-15T12:55:26.013002Z digest=sha256:2362bf46ed9925ea5efbf0ac412225e8f62cfb0f0f988524b26031e15722591c

Observation 29fc9b2b-188c-4ba0-b630-2c8f205a7ad6 · inbound

Optimization Hyper-parameter Laws for Large Language Models cites this paper.

Optimization Hyper-parameter Laws for Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:45:48.854973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T20:45:31.427677Z digest=sha256:7b4fefece94c2cc3a7897eb0970bacc57577d982cf4b0b95b13b35cac8adc826

Observation 7f798d6e-bf4f-47ca-b80d-96dc7b51556f · inbound

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models cites this paper.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.206774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.206774Z digest=sha256:96d5e179891203a10ac42c7f8b07c21b90b1e1218f2428f3cac3b7a664609e00

Observation 476a60c7-8422-4d96-af61-7dabfc07b7ce · inbound

MARS: Unleashing the Power of Variance Reduction for Training Large Models cites this paper.

MARS: Unleashing the Power of Variance Reduction for Training Large Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:43:01.833269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:43:01.833269Z digest=sha256:3a749490e8284c0d41f1ea17c119d886cd355a867f6d7738cc766c5622b3de14

Observation 64169dcd-7543-40cd-b53a-7a7859339e53 · inbound

Efficient Transfer Learning for Video-language Foundation Models cites this paper.

Efficient Transfer Learning for Video-language Foundation Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.185399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.185399Z digest=sha256:cda9a4637202bf64518bb36ff6008a6ee6c083b9caaced34589de07a8021e837

Observation d66f7588-e108-4fb3-babe-5fc137180918 · inbound

CKGFuzzer: LLM-Based Fuzz Driver Generation Enhanced By Code Knowledge Graph cites this paper.

CKGFuzzer: LLM-Based Fuzz Driver Generation Enhanced By Code Knowledge Graph DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T18:29:12.676620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:29:12.676620Z digest=sha256:96e2366104c1d318a7616a1b25239b7a1b313e5aeb8b566766cd9fa08d27f814

Observation 83969801-c87c-4964-b61e-f202f0979f00 · inbound

From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing cites this paper.

From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T18:48:16.372571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:48:16.372571Z digest=sha256:e74b0ded2fc03f2a3fc6baccb74ef8658620af62068ab3792515a8581a6abbdb

Observation 7b815b85-e9ab-4701-9624-96aa1d4dbaf3 · inbound

Ultra-Sparse Memory Network cites this paper.

Ultra-Sparse Memory Network DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T17:40:11.807604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:40:11.807604Z digest=sha256:fd6abf39b5331af0e535e7bcfe61c361bc113599860ae9c0b43a018032cc69b8

Observation c1e4f6b5-10d6-4f12-8071-c0bb3b1d3877 · inbound

InstCache: A Predictive Cache for LLM Serving cites this paper.

InstCache: A Predictive Cache for LLM Serving DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:48.071378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:48.071378Z digest=sha256:a71316deec01bcccb3ffbb7d237fcd07fcd9a89e2517331033f17ac3d52b6003

Observation 5e2aa33b-0680-4a83-b178-c13104739de5 · inbound

Masala-CHAI: A Large-Scale SPICE Netlist Dataset for Analog Circuits by Harnessing AI cites this paper.

Masala-CHAI: A Large-Scale SPICE Netlist Dataset for Analog Circuits by Harnessing AI DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:23:21.613312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:23:21.613312Z digest=sha256:98abdc7827d0d70df6dddfd846984bd5accac437d19c5f988738a05286a4fd78

Observation 1830ac7e-5fec-42f3-aa80-44318e9a7f4e · inbound

Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning cites this paper.

Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:24:14.314647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:24:14.314647Z digest=sha256:937e398c73e160910282c5154a9018779fcd1b0a5b616a03ddfb619057203a77

Observation fc9b9737-2fcb-4270-85f3-f079faa7112e · inbound

Traditional Chinese Medicine Case Analysis System for High-Level Semantic Abstraction: Optimized with Prompt and RAG cites this paper.

Traditional Chinese Medicine Case Analysis System for High-Level Semantic Abstraction: Optimized with Prompt and RAG DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:16:47.407167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:16:47.407167Z digest=sha256:63a37b11ee43c55b5e5ea51f304776003638bd0513c0032ecd6ee94b8ad5bb8a

Observation dcda7819-e101-42bc-a016-a0a466e36dce · inbound

LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training cites this paper.

LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:38.854212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:02:38.854212Z digest=sha256:8ca9b16242eb0446d8c4496ba87b728f27ae6d977029f9c980565e172ba45068

Observation be961e9a-07a7-40ec-8c96-0e20a6df2852 · inbound

Hiding Communication Cost in Distributed LLM Training via Micro-batch Co-execution cites this paper.

Hiding Communication Cost in Distributed LLM Training via Micro-batch Co-execution DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:07.360731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:07.360731Z digest=sha256:aa0ce76dfe99569e2c8443ea15cf7ab637ad9a7ca3e2eee2f4de0e648f65e8d5

Observation 936a3f44-63da-45ab-8f05-074f56abeed3 · inbound

From Laws to Motivation: Guiding Exploration through Law-Based Reasoning and Rewards cites this paper.

From Laws to Motivation: Guiding Exploration through Law-Based Reasoning and Rewards DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T13:49:22.443566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:49:22.443566Z digest=sha256:c67401f37c9e580fe491a697f2c48a1e761b53929a5a00b11662989d5de34b09

Observation 8eb31719-7604-4112-aecb-66e81e67e924 · inbound

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching cites this paper.

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:43.906039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:43.906039Z digest=sha256:945c0f3da386b0f779bf3899dd7b421edbc2259bc7e034b482f664ce3ffbd4f0

Observation 9fe4c1fb-76cf-435f-b423-bab0dc224dc5 · inbound

From CISC to RISC: language-model guided assembly transpilation cites this paper.

From CISC to RISC: language-model guided assembly transpilation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T13:19:49.548111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:19:49.548111Z digest=sha256:23f3f15af53e62cf8347e4f45e13547af28e81a40d6a9f4c406de65b18320979

Observation 5de252b6-756c-4638-89d1-0ed458182e27 · inbound

Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines cites this paper.

Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T13:16:38.284845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:16:38.284845Z digest=sha256:d7d05500c03433a58efe9587eaa53e5c44ebf22e10faa2cf55cc40f665fdb3b7

Observation ced3d7ac-354d-4ae5-8102-0cd4742e08dd · inbound

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning cites this paper.

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:55.179499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:55.179499Z digest=sha256:6d506ea304a6832cbf4fa3f896703e094bcffed635b3d1f448b9ea50283404a7

Observation ed97d7b7-c440-456b-8930-945c2dd856c0 · inbound

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning cites this paper.

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:55.243241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:55.243241Z digest=sha256:ec518f59b9edb965be6d5f529fd26e76e8b501907f5ec18c1c3a9482a77cfc60

Observation c4e34c8c-0d60-4e6d-be5b-a520734ca693 · inbound

CheckMate: LLM-Powered Approximate Intermittent Computing cites this paper.

CheckMate: LLM-Powered Approximate Intermittent Computing DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:53:26.080230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:53:26.080230Z digest=sha256:804a4fb0acc5c2fe9435e98afde84f59491094c1288875f3968188f8da98b8d1

Observation 593b19d8-e4eb-4572-9463-5f4c6b4afb31 · inbound

SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs? cites this paper.

SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs? DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T10:57:22.845501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:57:22.845501Z digest=sha256:14d010c455f36af80870bda8224e67078c9bfdcebb64190a52537c181d51255e

Observation f38e194c-d62f-41c8-a5a4-215fa8e5eac0 · inbound

Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference cites this paper.

Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T11:05:46.868872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:05:46.868872Z digest=sha256:e3eb5baf8dacc1d805f73bc2f6cca8fef362f03c14c56225a2a4062688f3795b

Observation ae8baa3f-74bd-410d-84d2-f59578610929 · inbound

Towards Adaptive Mechanism Activation in Language Agent cites this paper.

Towards Adaptive Mechanism Activation in Language Agent DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T05:09:47.733693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:09:47.733693Z digest=sha256:1aded75bb2c64b75d431a16762f43e2f8aaf1dacccd24deee57d9fed92d36777

Observation 836485a9-f86c-4bb0-894b-ede0812430d5 · inbound

Yi-Lightning Technical Report cites this paper.

Yi-Lightning Technical Report DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:34:12.075538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:34:12.075538Z digest=sha256:3c42563d0bab7622af4f7eb8203222f339689833087f48a7069c7e12954fabc6

Observation 97adbfeb-aef9-4d28-973e-33d255b53012 · inbound

FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism cites this paper.

FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:24:48.807586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:24:48.807586Z digest=sha256:41befd99fe05475863a01a59b3b03e070c0960b8045429c3a63dc8c0f244c4c2

Observation f33968a4-7f52-4a65-8c53-7783935bb369 · inbound

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity cites this paper.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.050262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.050262Z digest=sha256:5b2b0d48b5ebc145b5e2effcdf2c6eb11c99170e726bf3bd8889b372de0f5580

Observation 839629ed-f3af-48db-93f0-345442fd0851 · inbound

DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction cites this paper.

DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:12.138949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:12.138949Z digest=sha256:636251835c2424817095c1a4401bd828954d1602ca82948558edbdcffe02f82c

Observation e92aa4d9-474d-4f0c-9bc9-6f4ec3430183 · inbound

Weighted-Reward Preference Optimization for Implicit Model Fusion cites this paper.

Weighted-Reward Preference Optimization for Implicit Model Fusion DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:48:22.404833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:48:22.404833Z digest=sha256:f127a510a170e6e70e1cc45723089e412ea6ac797cdfb268d222b4ca9284f523

Observation 116d738e-abe0-4088-a56b-ba41f9d8362d · inbound

Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement cites this paper.

Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T21:56:12.064452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:56:12.064452Z digest=sha256:8a9673e72a1be6b690cbb7ae4119b852a808e041444fc1a6555e5424b48357c1

Observation 9e65f076-5237-4001-b76d-4f6aa739ba71 · inbound

Can Large Language Models Effectively Process and Execute Financial Trading Instructions? cites this paper.

Can Large Language Models Effectively Process and Execute Financial Trading Instructions? DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:16:05.172128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:16:05.172128Z digest=sha256:f0e7283a314652d103c0a6475a8d0050f8e9ebe7443f78c98f55113f4b0c76e7

Observation b5d7a32c-1d43-4e5a-bc5c-31496e5db9f8 · inbound

Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference cites this paper.

Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:11:39.158881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:11:39.158881Z digest=sha256:64592813a68ebaafebb2fb7d3971517564cce61696f5f2ac526f048a36a36fe5

Observation cfd63bb6-096a-47a6-903f-f137a0032516 · inbound

RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts cites this paper.

RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:32:31.609654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:32:31.609654Z digest=sha256:a18dfc72321960b87a74587f812713407131fabc1542701229d03bf2b6659981

Observation 0df3854e-ffd8-4a08-8f13-9f1bd4153b3c · inbound

Breaking the Stage Barrier: A Novel Single-Stage Approach to Long Context Extension for Large Language Models cites this paper.

Breaking the Stage Barrier: A Novel Single-Stage Approach to Long Context Extension for Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T19:11:06.149976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:11:06.149976Z digest=sha256:62a1003826a6437f9815c62be1a2bbe5b4ef83d3d1b34a20900f66a942d87f08

Observation 6aecb592-5ef2-4213-ad46-45b5c9a348e5 · inbound

Federated In-Context LLM Agent Learning cites this paper.

Federated In-Context LLM Agent Learning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T18:20:55.793098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:20:55.793098Z digest=sha256:0fb4c9c8283d1ddd4be6b6745959a32bdad73fe331dfd999bd59ae926c7eb9aa

Observation 238a6e9d-fc4a-49c8-a83d-c40f645908ab · inbound

DialogAgent: An Auto-engagement Agent for Code Question Answering Data Production cites this paper.

DialogAgent: An Auto-engagement Agent for Code Question Answering Data Production DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T18:19:43.718312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:19:43.718312Z digest=sha256:1299aeeb6277aeafa1e4bc85eca0ff5931c073a864fd1927a8e7d9d422c6299b

Observation a72d15b6-fd72-471e-a51d-cf82ec857e32 · inbound

Towards Wireless Native Big AI Model: The Mission and Approach Differ From Large Language Model cites this paper.

Towards Wireless Native Big AI Model: The Mission and Approach Differ From Large Language Model DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.416649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.416649Z digest=sha256:70f9148b31a3d220892a5897593e4c0039d057ffbf09140eeaec39128751cdba

Observation 4e555fa4-3e94-4e0a-aedf-7c0794b6d997 · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:09:24.100402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:14d245d4b14f5b8ce1da2a74285810ab62cade5bc41b234e20307ebaa85d7c51

Observation c225df22-e0e7-48de-b586-3e2de20edeba · inbound

SCBench: A KV Cache-Centric Analysis of Long-Context Methods cites this paper.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.889712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.889712Z digest=sha256:211dac6047fc132f09fdb8fa6216f8980254f86ba8ce55e0823d81b9a951cf71

Observation e4ea3613-51e5-455d-a2ae-f5df95c7c7f1 · inbound

Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation cites this paper.

Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:47:21.261517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:47:21.261517Z digest=sha256:4e36efa710f1d13424e6031115e900f32775425f9d6979b57d7f6cf425ac3b5d

Observation 4a1cdf2d-9a12-4467-a7eb-c2723f1a61f5 · inbound

RecSys Arena: Pair-wise Recommender System Evaluation with Large Language Models cites this paper.

RecSys Arena: Pair-wise Recommender System Evaluation with Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:22:54.486451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:22:54.486451Z digest=sha256:e9ff402ea5b00bb9dee1c3632ccecea480fadfccbffeb8ecf081769f38af914b

Observation 13349e22-1b1b-41c2-bf16-26644b466d81 · inbound

NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers? cites this paper.

NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers? DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:03:53.674982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:03:53.674982Z digest=sha256:0920c06fd14b0a440bbdc17ac07cf7a040374ba9ecadc6306661374b2000f266

Observation 47f656c3-4bbf-4ead-b092-3adb7a000df4 · inbound

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression cites this paper.

More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:48.727805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:48.727805Z digest=sha256:946ea90bc2ee1b9d15d0f55022c1603404b79f208d9bb9aaebf4d09134e161f4

Observation d3b4ac8c-239c-47d9-9d40-2a17a8fa8cc8 · inbound

UITrans: Seamless UI Translation from Android to HarmonyOS cites this paper.

UITrans: Seamless UI Translation from Android to HarmonyOS DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:56:04.296341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:56:04.296341Z digest=sha256:9dc3196ef38ea3c8589427c59259e2999fe3e0a6d2642ecbbb425db8ea1e3057

Observation 136fcb09-72cc-46bb-addc-de82990e8d10 · inbound

BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement cites this paper.

BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 2

Resolution
malformed identifier
no resolver link, observed 2026-08-11T14:36:49.580244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:36:49.580244Z digest=sha256:395afcf1a8d76eab457ecda76e2f3d5ae536b7830876f2f0b0a801b972093d58

Observation e4a18125-189f-4c47-ae75-7b88649711c7 · inbound

Language Models as Continuous Self-Evolving Data Engineers cites this paper.

Language Models as Continuous Self-Evolving Data Engineers DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:40:52.519401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:40:52.519401Z digest=sha256:2ab1bd12e310e5ae522978160664f91226c7ce2e6787866ae737d0c079984449

Observation 172557b3-2085-4393-afdb-a81242135a36 · inbound

CodeV: Issue Resolving with Visual Data cites this paper.

CodeV: Issue Resolving with Visual Data DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:39:29.897006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:39:29.897006Z digest=sha256:6f9fe02b162ab4c9d01c628b628f60d809bd0ebee5c1f9ce8398f0d06ea1fad3

Observation 6b76e7ab-593f-46ea-a41d-0730763377ac · inbound

A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression cites this paper.

A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:29.407130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:29.407130Z digest=sha256:dbea7de25b40629b36d48e3717694a9d31d77a7c72f6e463ecee87f1d113e163

Observation bc5c1223-b083-41f4-94d7-3c48f1089703 · inbound

YuLan-Mini: An Open Data-efficient Language Model cites this paper.

YuLan-Mini: An Open Data-efficient Language Model DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:17:55.563036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:17:55.563036Z digest=sha256:0c5096d54a4f8142fdf45616554576952d8b2bac340cc40853a56616a4d0b6cf

Observation 7c3a3da0-9b7e-444b-9ea7-9e5134f53ad8 · inbound

In Case You Missed It: ARC 'Challenge' Is Not That Challenging cites this paper.

In Case You Missed It: ARC 'Challenge' Is Not That Challenging DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:13:07.193694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:13:07.193694Z digest=sha256:49191735caf33625d4dbf2fda487c86bef134d5f626feac38c1f27df668f4a3f

Observation db60e678-aaca-40c2-8c99-d91446af1d55 · inbound

Multi-matrix Factorization Attention cites this paper.

Multi-matrix Factorization Attention DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.802359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.802359Z digest=sha256:9a5bc80e76c803a1000481e4824e9aecfed1853beb554fabd419a876547aba80

Observation 73a20849-83dc-4539-b7f4-56290923b57d · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.341796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.341796Z digest=sha256:8c21ae577c313f356fbcc73aabae446bfd4a8c4f6344ee05c6559c392b50e7e4

Observation f8202b31-bbbc-4914-9ae9-4795852d7478 · inbound

BaiJia: A Large-Scale Role-Playing Agent Corpus of Chinese Historical Characters cites this paper.

BaiJia: A Large-Scale Role-Playing Agent Corpus of Chinese Historical Characters DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:42:30.572122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:42:30.572122Z digest=sha256:5166dba759ada920ac39a3ed255177c3de93b1d6f82c915e91c55c1574c61dfb

Observation a489ebb9-8328-4c56-8ea7-5a33c9091de7 · inbound

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA cites this paper.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.028493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.028493Z digest=sha256:00a519db7afb2ccd1b314b0480e70d840f24e0a3f6b6ff3896902244194f20ec

Observation 86ee5653-5959-45a3-898e-d4776ea430f0 · inbound

SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity cites this paper.

SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:13:15.096585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:13:15.096585Z digest=sha256:e86446e81da85d5888a2d80fee1a3caa88eac14da3006ef5bcd3443b794d6ec3

Observation 4b30ab26-4207-4482-920c-93f59be5f71e · inbound

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism cites this paper.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:26.514375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:26.514375Z digest=sha256:1b327220936099617ccd2b2a24831ea14b0b78648d367471d748c0f36b6911a4

Observation 2c5be030-5340-4c47-aee8-273f73006960 · inbound

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation cites this paper.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.704255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.704255Z digest=sha256:f94aca732d94cbff40c1b2783729ce4ee9f3836c507c32169fe99d4730bebb6d

Observation 99ab6639-1f2e-4950-8bb6-cc8afdd71ba8 · inbound

GroverGPT: A Large Language Model with 8 Billion Parameters for Quantum Searching cites this paper.

GroverGPT: A Large Language Model with 8 Billion Parameters for Quantum Searching DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:04:43.035562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:04:43.035562Z digest=sha256:d9ffa804484156e423da3349e20a0792e3eed6c92662369e4350858191467ec2

Observation 5bcd55fe-8caf-4238-a58d-eb4da5729f18 · inbound

SPDZCoder: Combining Expert Knowledge with LLMs for Generating Privacy-Computing Code cites this paper.

SPDZCoder: Combining Expert Knowledge with LLMs for Generating Privacy-Computing Code DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:16.886242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:16.886242Z digest=sha256:389d50a5c88c73d4251a1fe72da9633bbd5c56307999de75832ae880b83b7200

Observation c23b738d-afc2-495b-97df-a64d1b3f2a41 · inbound

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings cites this paper.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.458249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.458249Z digest=sha256:0491f0f1178bcbbc4be3bb1c24a990431a196e651bee16adaa1d9bc53495f7ba

Observation ac5e5624-377a-4491-b372-b1577c3b36bb · inbound

AgentRefine: Enhancing Agent Generalization through Refinement Tuning cites this paper.

AgentRefine: Enhancing Agent Generalization through Refinement Tuning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:24:43.084676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:24:43.084676Z digest=sha256:5a944e62d9b38dcbe16a16eed74d0be97adddda39b3979b8b822c98bd6377918

Observation 82e05187-17ee-4e58-b815-c572ab96726e · inbound

MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning cites this paper.

MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-10T22:23:27.589782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:23:27.589782Z digest=sha256:7422b225fbf74d913f6c691e166007b8590087aadb28b9a5650b451457fdf6f5

Observation df2668ec-012f-4a7c-a354-8112ef86a33f · inbound

Long Context vs. RAG for LLMs: An Evaluation and Revisits cites this paper.

Long Context vs. RAG for LLMs: An Evaluation and Revisits DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:08:55.809067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:08:55.809067Z digest=sha256:38740165d483937e28e8b6568e6e92c42ff5e227643643a970bf6191cc366219

Observation c109257c-7a89-4c89-a680-499c5e8e0782 · inbound

Instruction-Following Pruning for Large Language Models cites this paper.

Instruction-Following Pruning for Large Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:29.141016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:29.141016Z digest=sha256:c426b0f97d3fe9fa52ec968858f81fe7bdae13636c3c073a2037ed7907eccdb9

Observation af31485e-0ef6-4a8c-a61f-2093a3400e0c · inbound

AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference cites this paper.

AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:51.167526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:51.167526Z digest=sha256:31f70f3b0e1d2c92a20a1a3128ece25b5057da01f11cf217eeb15b700b0b3675

Observation 37b81e57-e86d-4adc-b0ed-fe200c8079b9 · inbound

Ultrasound-QBench: Can LLMs Aid in Quality Assessment of Ultrasound Imaging? cites this paper.

Ultrasound-QBench: Can LLMs Aid in Quality Assessment of Ultrasound Imaging? DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:10:38.414598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:10:38.414598Z digest=sha256:6f01660e5054d7d0828a307042e82222d5a744485049b7dd686b2495ec2cc771

Observation bb8a9ca5-6116-48e3-a3c7-ce21d45a457a · inbound

Beyond Pass or Fail: Multi-Dimensional Benchmarking of Foundation Models for Goal-based Mobile UI Navigation cites this paper.

Beyond Pass or Fail: Multi-Dimensional Benchmarking of Foundation Models for Goal-based Mobile UI Navigation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:05:23.849608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:05:23.849608Z digest=sha256:59d6cb0538895c30ab0c8a81152c803cff19afe9a75736dc8fb7de75526d93a5

Observation 7cfe15f4-f9c0-470e-b5f0-f969146b7861 · inbound

MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training cites this paper.

MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:18.983386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:49:18.983386Z digest=sha256:02d842cfe9b33fff3771914d634c82c452bafb9d173dd4de9317995e22f15a9e

Observation cbc66737-a108-421f-a3be-43bd22820999 · inbound

Powerful Design of Small Vision Transformer on CIFAR10 cites this paper.

Powerful Design of Small Vision Transformer on CIFAR10 DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:58:02.272146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:58:02.272146Z digest=sha256:da51bb78379f7341fbc1ee2e437936c73478419524e6cffb4dfbaceec671effd

Observation e6119b63-645e-4fc5-aba1-bc75ed39954d · inbound

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch cites this paper.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.240581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.240581Z digest=sha256:1f33fbd6237a4ac91a4240f80d1e1a7746d9f47eb9d0a914db39797a2dcdd218

Observation 7d8f7dc2-706e-425d-b3af-3c2f4f3db144 · inbound

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training cites this paper.

OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:20.559243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:20.559243Z digest=sha256:fa833374b38e47cf5749a6baa0ff5db94be4c2767920187977ffeeb471e0e4a8

Observation 74bc134c-f68a-4b82-82ee-60932b2f7f89 · inbound

Sympathy over Polarization: A Computational Discourse Analysis of Social Media Posts about the July 2024 Trump Assassination Attempt cites this paper.

Sympathy over Polarization: A Computational Discourse Analysis of Social Media Posts about the July 2024 Trump Assassination Attempt DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T19:35:21.002504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:35:21.002504Z digest=sha256:5b12988804b1527cd7798a8af8ad645aab205b4e82eff1337e4d3d9394782c8b

Observation d3775be2-59b5-40c0-89c5-ac32b53d0c83 · inbound

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training cites this paper.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.539124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.539124Z digest=sha256:6b99ea7695ea0153cc6c1a0a02d8662668ffd7641da6b2e6bf1f977d16af2563

Observation d6a6b5af-4d9b-4ec7-8c9c-90751895dddf · inbound

SR-FoT: A Syllogistic-Reasoning Framework of Thought for Large Language Models Tackling Knowledge-based Reasoning Tasks cites this paper.

SR-FoT: A Syllogistic-Reasoning Framework of Thought for Large Language Models Tackling Knowledge-based Reasoning Tasks DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T18:06:04.373098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:06:04.373098Z digest=sha256:bee4beefd5c186ac501817ae69a4decc110a8504830dbbd05b00cf3831b8b133

Observation cc3be2fc-d092-457e-a576-76e2976d1c2c · inbound

Panoramic Interests: Stylistic-Content Aware Personalized Headline Generation cites this paper.

Panoramic Interests: Stylistic-Content Aware Personalized Headline Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T17:48:42.555717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:48:42.555717Z digest=sha256:c621877b3519cc1a7a0d5f9814d11fc8bdead1abd535417df51178c652127b80

Observation 950a6b8e-890d-4465-a941-3e9a104100a9 · inbound

Pseudocode-Injection Magic: Enabling LLMs to Tackle Graph Computational Tasks cites this paper.

Pseudocode-Injection Magic: Enabling LLMs to Tackle Graph Computational Tasks DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:42:43.981050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:42:43.981050Z digest=sha256:dd9a7cd0ba8959a70fb0e1747bf3c87d408f4a933da06aba55697a2d46ddf2e9

Observation a76af209-837b-4033-9eb1-9254dcff67e5 · inbound

Mixture of Experts (MoE): A Big Data Perspective cites this paper.

Mixture of Experts (MoE): A Big Data Perspective DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-10T18:56:37.575262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:56:37.575262Z digest=sha256:e2fddb01b834abd637de9e3d1a9c260b92dbec4e78b194d905e4fe9ff21aa3be

Observation 9dffec26-ad04-42ff-8152-7cac27223013 · inbound

Scaling Inference-Efficient Language Models cites this paper.

Scaling Inference-Efficient Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.964120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.964120Z digest=sha256:1b2edf06203045ab826a87bdaeb31103d39e7c18bd0ec70f96841df1ae7058b6

Observation d2f18128-890d-4c80-8908-13c4a6f87e8f · inbound

PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration cites this paper.

PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T18:46:51.047168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:46:51.047168Z digest=sha256:58c11c1f8103d452ce083420d099a56648866ee06ca7783a031e2b6c7cae5e19

Observation 7071e68a-a6a1-400c-b327-0df162bf3b31 · inbound

OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology cites this paper.

OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T16:10:24.931071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:10:24.931071Z digest=sha256:ec9cd693bc1b0e0bb655904fa61870f2bdffd952db33cde9ef8ce89a7ca3923d

Observation 377ae90a-b793-43db-9794-e3700844ad21 · inbound

Position: AI Scaling: From Up to Down and Out cites this paper.

Position: AI Scaling: From Up to Down and Out DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T18:19:21.016100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:19:21.016100Z digest=sha256:add2a6eae1823e039f4daed34bb3ab4561f3c29c32ff3ea12a81f1d30b3260ec

Observation 2558f0b6-7fca-402f-a2b0-af24642dc7cf · inbound

COFFE: A Code Efficiency Benchmark for Code Generation cites this paper.

COFFE: A Code Efficiency Benchmark for Code Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:01:27.638350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:01:27.638350Z digest=sha256:42d1a8f53f2d9eee35f188a0f7f8b07630766fba8771d18a3666854592e8684f

Observation 8d1ce5d2-9612-4ef2-98cb-35b4a89334a5 · inbound

GistVis: Automatic Generation of Word-scale Visualizations from Data-rich Documents cites this paper.

GistVis: Automatic Generation of Word-scale Visualizations from Data-rich Documents DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T00:48:08.500927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:48:08.500927Z digest=sha256:e2cc8f1222edde0c6fd3d153487df6e6af34c7705bdd746fc14a4c8cb3beb3a3

Observation e836b4a3-3e60-4d92-ba62-c69efb74df8a · inbound

Steel-LLM:From Scratch to Open Source -- A Personal Journey in Building a Chinese-Centric LLM cites this paper.

Steel-LLM:From Scratch to Open Source -- A Personal Journey in Building a Chinese-Centric LLM DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:50:58.756577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:50:58.756577Z digest=sha256:7d4770a21244cbe2a124b93b09b91d95bea37d116c4dae0520a707fba191a4a5

Observation 3f1b0cd1-8d68-490a-884a-4d81a6bd0aa5 · inbound

Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline cites this paper.

Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T17:55:53.994050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:55:53.994050Z digest=sha256:5603841f8783187993c4b541a0b3ab6d9b81fc84194341cf8814b2e1fdcafc19

Observation ae3c64e5-41b6-4ea1-b54c-9c533c2c8c2c · inbound

CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction cites this paper.

CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:11:51.555174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:11:51.555174Z digest=sha256:170019eba2bde595dc5e4e7614452ba4d277cce900473cb909d055b9bbc64d7c

Observation 4c0eb8d3-4657-462d-9d52-f98af83a59b4 · inbound

Memory Analysis on the Training Course of DeepSeek Models cites this paper.

Memory Analysis on the Training Course of DeepSeek Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:59:38.687594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:59:38.687594Z digest=sha256:4c5bfea706230299a59ec70579c2a18889ca689b775e2d1b0a52bb0d087d52fe

Observation d8363b8c-7845-4cd3-a8ea-02c502c0c98e · inbound

TransMLA: Multi-Head Latent Attention Is All You Need cites this paper.

TransMLA: Multi-Head Latent Attention Is All You Need DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T11:48:01.162895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:48:01.162895Z digest=sha256:87e7a982c56a83ddeebd292518466f7f706efe6599a127fd5cb72a9acec55e63

Observation 80d2a90f-0307-4119-bded-4c4af08d5431 · inbound

Explanation based In-Context Demonstrations Retrieval for Multilingual Grammatical Error Correction cites this paper.

Explanation based In-Context Demonstrations Retrieval for Multilingual Grammatical Error Correction DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:01.830167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:54:01.830167Z digest=sha256:773ae81cb25e5531cfe08e9557a4fb2b1136d14d6629a007c218e2b9dd428461

Observation f315a9ca-f5f2-485d-a0d1-0ab920c721ea · inbound

Universal Model Routing for Efficient LLM Inference cites this paper.

Universal Model Routing for Efficient LLM Inference DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:48:00.855639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:48:00.855639Z digest=sha256:423a4f51e02ece0f3a0505ce0edc3bfc6aced672b45c1f4e714c2ded1a06cfff

Observation 3545b50a-5b7b-4782-8228-be7689c494c1 · inbound

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU cites this paper.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.811005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.811005Z digest=sha256:fe27a84d7d12935ea777241209cd5929f0f103bab2168effdc2baf37a44e6c98

Observation 00e310b6-3e27-4074-b52c-906ae8674f8a · inbound

You Do Not Fully Utilize Transformer's Representation Capacity cites this paper.

You Do Not Fully Utilize Transformer's Representation Capacity DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T22:19:19.351787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:19:19.351787Z digest=sha256:1018427a8183e0b5df852480b4a1392221590b03fdf11cee0c6902458d1e5da1

Observation fea2f38a-a783-4d78-96f7-373c093114c6 · inbound

Climber: Toward Efficient Scaling Laws for Large Recommendation Models cites this paper.

Climber: Toward Efficient Scaling Laws for Large Recommendation Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T20:15:14.070781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:15:14.070781Z digest=sha256:24c093a47b549f26f3fd2707b397d8b05b42db6853b43e6dedd47b1e7404444e

Observation b8f2a7ec-b82b-440f-8e1d-3fed18c87108 · inbound

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention cites this paper.

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:46:30.032182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-16T23:46:29.975858Z digest=sha256:eee96168b8e25fb52f059425820f00705eb96629eeacbc8fcd213f43ad9946fc

Observation 85c2f796-d36f-44bb-8a1a-21c61db8f5e2 · inbound

MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections cites this paper.

MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T22:34:08.058138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:34:08.058138Z digest=sha256:f862ddc3cd8b4f2bfe6ff548b3b123a407e8e4eefc57977f52f57087369e7c40