Pith. sign in

Paper Citation Record · LEDGER

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models

As of 7 August 2026, this Paper Citation Record lists 100 of 142 outbound references and 0 inbound Pith citation observations for arXiv:2507.09955.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09955 v1

Coverage vector

measured 100 of 142 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:47:40.258067Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 142 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92074b83-7e64-41e9-867f-3f7a37855124 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.866977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.866977Z digest=sha256:337259f178b832a51a0a51635b368b7c216d9fb954dae01beb028b8e27e5b28f

Observation f1dc8827-3bf3-4305-b6db-a38dfb845eae · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.918113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.918113Z digest=sha256:bbe3290bba0240b20be8b5e4d8d02e823ff8fc8b7317e8ed335a7646567a515d

Observation ddb35eec-f457-4a72-8535-cdc1633a5159 · outbound

This paper cites China’s cheap, open AI model DeepSeek thrills scientists,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models China’s cheap, open AI model DeepSeek thrills scientists,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.968566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.968566Z digest=sha256:aa913f6da2ff2b50c6277ac5ef4b8360ca51abe03c129ba5859a9ca6b7b36dd8

Observation 1ca8a6d7-63cb-47bc-8eae-b139caf3a9f8 · outbound

This paper cites What to know about DeepSeek and how it is up- ending A.I. - The New York Times,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models What to know about DeepSeek and how it is up- ending A.I. - The New York Times,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.971862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.971862Z digest=sha256:6c138722e17a2104e72b61306a7f79fbca88900a8bfec6a312cf8cc4c40aee2d

Observation ec9b6116-5b32-452f-8788-10c257fde08f · outbound

This paper cites What is DeepSeek, and why is it causing Nvidia and other stocks to slump? - CBS News,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models What is DeepSeek, and why is it causing Nvidia and other stocks to slump? - CBS News,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.975132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.975132Z digest=sha256:48283c487859d52250f03592a3f53107a9d3b670f2b3afb63f72378794202871

Observation 6bf67c6e-1566-4bc2-9c5e-3db0bfb0f2a6 · outbound

This paper cites A brief overview of ChatGPT: The history, status quo and potential future development,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models A brief overview of ChatGPT: The history, status quo and potential future development,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.978543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.978543Z digest=sha256:7a98ef3a7029c0e75ba9bee54239a5a23ffeebb3b9098b2483c6bc5aa4b24273

Observation eba4530e-881a-41f6-a531-b1552c9ea905 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.981924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.981924Z digest=sha256:fe5047f5982a7eabdaf660204a2946490193579c3279b8dbf476706e4a9e8b3a

Observation 2b61b4bd-a087-43a4-945e-ec269b2b1b24 · outbound

This paper cites Training language models to follow instructions with human feedback,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Training language models to follow instructions with human feedback,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.985035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.985035Z digest=sha256:715b3757b7db65bc5c44757fda86c6d0127864101506b7d2328319c41b24739d

Observation 508a59c2-21f9-4c20-a3fc-0d2b40398938 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Learning transferable visual models from natural language supervision,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.987912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.987912Z digest=sha256:4d14ee513db4e2ee265dd7ae1ed6e2aa5fce7e4a474e51b60de3d235f7362121

Observation 02ce712d-d8bf-4ceb-92d2-05a162d081c5 · outbound

This paper cites Pre-trained language models for text generation: A survey,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Pre-trained language models for text generation: A survey,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.990600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.990600Z digest=sha256:e9b5b1a0a7a03a35798411600406a296969db5d4411faa7e6a41c7d063be4ab3

Observation 83795e88-4990-4172-96bd-a827f183ac0c · outbound

This paper cites Recent advances in natural language processing via large pre-trained language models: A survey,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Recent advances in natural language processing via large pre-trained language models: A survey,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.993254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.993254Z digest=sha256:296b3aaeba33ac3ace971a760a5061af3cbb0b3055da6a54899a2318f13147ed

Observation 3a6f7e12-ce04-448d-8bcb-ff9005580b30 · outbound

This paper cites Large language models versus natural language under- standing and generation,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Large language models versus natural language under- standing and generation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.996069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.996069Z digest=sha256:8f378bb34712800af58a85333abc9314ada0b967355db0ed7a2b830150340ac9

Observation 4c7cfa2c-1b45-4230-8687-4602466361a5 · outbound

This paper cites Application of deep belief networks for natural language understanding,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Application of deep belief networks for natural language understanding,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:39.999322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:39.999322Z digest=sha256:ee64d6554545353a2e76132f7ce9323343e4a497c703e41c2e06bc79f645c88d

Observation 5fdae1a1-a2f1-450e-91af-3713ab1fccc5 · outbound

This paper cites PaLM: Scal- ing language modeling with pathways,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models PaLM: Scal- ing language modeling with pathways,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.002306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.002306Z digest=sha256:ce2bb70469426c21ff95da31e3186a011e119b3d8e70ab967efbcc636155c600

Observation 519e6f24-fb65-40e1-a2c4-2f61018178a5 · outbound

This paper cites Vision-enabled large language and deep learning models for image-based emotion recognition,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Vision-enabled large language and deep learning models for image-based emotion recognition,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.005510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.005510Z digest=sha256:9f0bb2ccdf7efb3d4ebcbb28c08e39e72aa71475ff1f6fb17cc98fb54afc86d2

Observation d017a179-eb73-4246-886b-fdd7e7dd74e9 · outbound

This paper cites A survey on deep multimodal learning for computer vision: advances, trends, applications, and datasets,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models A survey on deep multimodal learning for computer vision: advances, trends, applications, and datasets,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.008318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.008318Z digest=sha256:ca4078648341ed20b7a00f1b4b435774c8f87ca3f2be0171a5109003b1573ace

Observation deda322d-2d2c-4541-8983-a349d4bf0ecc · outbound

This paper cites Deep learning models for digital image processing: A review,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Deep learning models for digital image processing: A review,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.011232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.011232Z digest=sha256:2696b729fcb836c46ae6e6b74b7c1f2413e82aedaa5b1c90d0d1a751293c8ae1

Observation 2fc1e6c6-0333-4274-b3c4-e08a81bec49c · outbound

This paper cites AI Action Summit (10 and 11 february 2025),.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models AI Action Summit (10 and 11 february 2025),

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.014079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.014079Z digest=sha256:c2e839756c5977e8e9cf6ca3c40de19234ef0a139ded6fc20ebe67c08187e427

Observation 24109014-e54a-44fe-bf46-cc24a7a7ec99 · outbound

This paper cites Can ChatGPT replace traditional KBQA models? An in-depth analysis of the question answering performance of the GPT LLM family,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Can ChatGPT replace traditional KBQA models? An in-depth analysis of the question answering performance of the GPT LLM family,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.017311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.017311Z digest=sha256:5a6f51c2fd4de868560eedfc6364b4a0a36bfd79ed7205a98fd32f86b82d900d

Observation 1f0f5a32-f24b-4004-95da-f8f49d7e4160 · outbound

This paper cites Reasoning with large language models for medical question answering,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Reasoning with large language models for medical question answering,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.020675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.020675Z digest=sha256:f534869ca81c279acad683a25843c28347499b2ab60995643da08b62c036b596

Observation 3dd5671d-8234-4873-9d1c-963414a8d4f3 · outbound

This paper cites Proactive conversational agents in the post-ChatGPT world,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Proactive conversational agents in the post-ChatGPT world,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.024306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.024306Z digest=sha256:965b0cf67c519ed0082cfd7606888c7b4389ce51a007d0f12459adfb82f1e6ee

Observation 7a795c53-071b-4131-864c-d5ba0da1af46 · outbound

This paper cites Unlock life with a chat GPT: Integrating conversational AI with large language models into everyday lives of autistic individuals,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Unlock life with a chat GPT: Integrating conversational AI with large language models into everyday lives of autistic individuals,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.027453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.027453Z digest=sha256:2c2ae641c0c521da08aa896e0f8416d879f6405db00ce3d3fdc6afdbb084406c

Observation a81e4371-5062-4a89-9e24-9881f4d81f18 · outbound

This paper cites A contemporary review on chatbots, AI-powered virtual conversational agents, ChatGPT: Applications, open challenges and future research directions,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models A contemporary review on chatbots, AI-powered virtual conversational agents, ChatGPT: Applications, open challenges and future research directions,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.030681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.030681Z digest=sha256:e4baa68eca55e831eec71cbaaaa9aab4e518d002148e1ddb0e86af351aa4fbdf

Observation 9f2fcb8d-761d-405e-b986-e18e73fc5598 · outbound

This paper cites Self-collaboration code gener- ation via ChatGPT,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Self-collaboration code gener- ation via ChatGPT,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.033602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.033602Z digest=sha256:8110a27d2e3872f3d807b7526d3c6a0eac4198dbf22b770958f8469e4b563d25

Observation a49b61a1-00f1-4f2d-838c-2dd4b6ac04c5 · outbound

This paper cites Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.036683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.036683Z digest=sha256:e6292255c1e8d4ad0a94315a18f90b0e790ce7ed50fed34b834ef3eece74e1c0

Observation 9aff0331-9273-4b15-96b8-cf5ad8a720cd · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Gemini: A Family of Highly Capable Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.039690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.039690Z digest=sha256:e6faa1dd63b93c8a94436d27e0ea9d502a40007c63b0142dd706286c8556736b

Observation bd157fda-7fea-47c5-acbb-68166b010aaa · outbound

This paper cites Introducing Claude 2.1,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Introducing Claude 2.1,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.042967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.042967Z digest=sha256:90622e11bd890cb64e3f58b73f1474db672c439c5178d0aa16ed31cf303fca98

Observation ecb943fa-ca36-4da2-a099-bf97c0efb468 · outbound

This paper cites Introducing llama 3.1: Our most capable models to date,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Introducing llama 3.1: Our most capable models to date,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.046496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.046496Z digest=sha256:302ac202477b8dd1a8a6b2ff20a834c9caa3edd96c9c0e75df7220689393d795

Observation 3133a507-f0d3-4e44-b195-d0c8dcd64725 · outbound

This paper cites Mistral 7B.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Mistral 7B

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.049924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.049924Z digest=sha256:40776685ea61375492248631cac7b14a5caa0c04edcb6fdf6a55cdbeb91b0613

Observation a54af932-c4a9-4673-bdd3-d6824d564078 · outbound

This paper cites A comprehensive survey on transfer learning,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models A comprehensive survey on transfer learning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.053192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.053192Z digest=sha256:4edb07032ecf0a923701d10de01a9475eb7dc42cd9407274871366ebdb058cd2

Observation 00d1e4ec-8184-42e7-b560-b368e6e37033 · outbound

This paper cites Multimodal learning with transform- ers: A survey,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Multimodal learning with transform- ers: A survey,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.056373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.056373Z digest=sha256:26bfb45df0da474305c931b1b3dde56605889c178a8db32e03f248b65ee55ab9

Observation f2dcf991-e4c8-47ed-adaa-c9243f38109f · outbound

This paper cites GPT-4 Technical Report.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models GPT-4 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.059461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.059461Z digest=sha256:c029144106a06c437e816422c3d421904fba6c1b4f7349ff0cb7c7dc50c3ae59

Observation 7e67bb5d-6465-4742-a8b4-71d3a9b07a48 · outbound

This paper cites DALL·E 2,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models DALL·E 2,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.063226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.063226Z digest=sha256:a511bd3694362c9b25d5e86cdb569072a3ede1eb09aefce7f211d5c8b237332b

Observation aca591a8-f1e8-48e4-9bba-79439d258f36 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.066287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.066287Z digest=sha256:e2dc8b2f9338879fef106e6c6a663d18a911d4c8f34023c29f27a2a530c04f57

Observation 9ce2e707-82b4-46d2-aa04-6ab53b787977 · outbound

This paper cites Introducing OpenAI o1,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Introducing OpenAI o1,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.069629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.069629Z digest=sha256:3337bea51462acd3e021112849fdcb33420dfe82bb9f25d0061e373a6ba8e5a2

Observation 34f17dc2-7197-4a22-a9b8-fcc9569466b8 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Math-shepherd: Verify and reinforce llms step-by-step without human annotations,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.073209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.073209Z digest=sha256:96e3ee0059b652a5fb72b17863c80709eec361f4ad27b653d6203788d6cc3976

Observation 8f557111-f919-4bf4-9116-2a49950fd7e4 · outbound

This paper cites Grok 3 beta — the age of reasoning agents,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Grok 3 beta — the age of reasoning agents,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.076107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.076107Z digest=sha256:1ad509986cabbc0b9f5337d17469230049c3d5bbf7246d50004ade5232c2c9c6

Observation b86e704d-e0c2-469c-adcc-37dedcda7723 · outbound

This paper cites DeepSeek-V3 Technical Report.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models DeepSeek-V3 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.078990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.078990Z digest=sha256:c09771c79aa89e9cd4d601fc7089396b4a0b4e1a675349db4d902dfea3c48413

Observation 9490a70d-5e07-44a4-ba95-d1bcb6814d45 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.082238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.082238Z digest=sha256:c2484dae40bcecd15416a1d91eea4dc6566b3078672d2a1d10d4ee958a3df66b

Observation 8a9a38f1-681d-44a5-aa31-c58902be3ce9 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.085404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.085404Z digest=sha256:92bbd520ba0e4da3127bb6edefa5ae0b0a763bf52a5adacd7f900ac31302f983

Observation 3437fb8f-f996-46c8-be8c-ebfba3597bad · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.088517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.088517Z digest=sha256:2ba1223b539ad4cf2af9e9cdad6447183c580731915ea9d1313dc94678038d0c

Observation b8baf641-84e4-4f4e-914a-b8451b77ace5 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.091431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.091431Z digest=sha256:cb8d42b885da1dd274c56519b6585750903bd8ab0d2d9427d3d658f71ceba83c

Observation 7dd3a60f-51a4-4f0c-85d8-83b0ef82e23c · outbound

This paper cites Hidden markov models,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Hidden markov models,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.094795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.094795Z digest=sha256:f3706221c9134442347b7ddc9cb210559358107a1f82f5ef42043ffbc964a729

Observation ce9febf7-f751-4191-bf72-4d7855790936 · outbound

This paper cites Large language models in machine translation,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Large language models in machine translation,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.097597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.097597Z digest=sha256:41b5ddccd2d52c9b885a70b5e7e088f32eb237f4a8d1bb7cea87d81da0f2f495

Observation adba33ed-a799-4310-91d1-7bb725631250 · outbound

This paper cites Dis- tributed representations of words and phrases and their composition- ality,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Dis- tributed representations of words and phrases and their composition- ality,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.100318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.100318Z digest=sha256:bdb57746f33b46ada83fe1287ad18a5990fff5a41e2119539dcf07cc6f2314f4

Observation 7058d49c-513f-49dc-81db-d5a65068f3a9 · outbound

This paper cites Glove: Global vectors for word representation,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Glove: Global vectors for word representation,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.103188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.103188Z digest=sha256:0b64e392389f7bd85316f1f391b84710ad69f6494b80151c373082ff15d9708d

Observation 8fee7647-9e6a-4695-aa7d-28f3e2eeb5cb · outbound

This paper cites Fundamentals of recurrent neural network (RNN) and long short-term memory (LSTM) network,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Fundamentals of recurrent neural network (RNN) and long short-term memory (LSTM) network,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.105871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.105871Z digest=sha256:c81d919de7317d22da7f792fba4e5f6b9a8cf85d887b8fe2d08a358c915da32a

Observation 1e5f51cd-e2fd-4d0c-a7fe-f888b12f7b99 · outbound

This paper cites Long short-term memory network for learning sentences similarity using deep contextual embeddings,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Long short-term memory network for learning sentences similarity using deep contextual embeddings,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.108903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.108903Z digest=sha256:d09a484850b8736a0326e4e4203df003d66bbd33cb4999f9e165e821b732b99d

Observation f84e484a-0dce-4131-bab1-4678a40553af · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.111522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.111522Z digest=sha256:9ab23ce2ccd3c1ef96bb0b23bc87b60671d7de09d6d0d5edf4159befe0ca03cc

Observation 66b9917d-d3e5-4c0d-8f3c-5288bb9da120 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.114788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.114788Z digest=sha256:191d7d41f7f7ad1f888706a31bf625debd36fb6d1d4ef0e41d1217cc361cfcfd

Observation db441385-dba8-4447-9b2c-a49a203eb58e · outbound

This paper cites BioBERT: a pre-trained biomedical language representation model for biomedical text mining,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models BioBERT: a pre-trained biomedical language representation model for biomedical text mining,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.117740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.117740Z digest=sha256:dae9ab580b5a533bb1421880f5abe104aa2d0711c83f840d923340e74ff2572b

Observation 7fee2620-5985-46c4-be11-03b89ae19137 · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.120391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.120391Z digest=sha256:2f2409d2ed3d2116c7f770da00df8bc60396bf9a3f10e20a144c93f9f0449812

Observation b8082bd0-aa27-4449-a25a-26f29579803c · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models A General Language Assistant as a Laboratory for Alignment

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.123293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.123293Z digest=sha256:37e10e79688bcb5a6f6f127274d93f833857d64d939989fb6465939f00a41429

Observation 63014fb7-1378-4262-90f6-05ccfd027a62 · outbound

This paper cites ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.126152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.126152Z digest=sha256:e884b6fd0e6b7255fea902c84837d41a3b114246a87d75c68a6aabe2b929c993

Observation e52f074a-28cd-46c2-8629-c421bc329d4f · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.129359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.129359Z digest=sha256:fd858fc5302be0d6ac7ab9d3cc7ef0e38ba14cd3d5b4acc82779cd0d44623cf0

Observation 6b3552a3-4045-4668-8fa4-655c3fc1d823 · outbound

This paper cites Language models are unsupervised multitask learners,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Language models are unsupervised multitask learners,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.132643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.132643Z digest=sha256:d1fc56dd12733f2176c01c58cb859a9c3786bcf0255f7550bbd2f6b7ef97689d

Observation 1f730cbd-470c-4678-a7a7-227779c2a3f1 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Evaluating Large Language Models Trained on Code

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.135490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.135490Z digest=sha256:e7908b7fa21547100595e41266169bb40d7a15067c0ec81fb08914f874309e6c

Observation f4f5d346-a886-4313-a195-0061b974b6d9 · outbound

This paper cites CodeGeeX: A pre-trained model for code generation with multilingual benchmarking on HumanEval-X,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models CodeGeeX: A pre-trained model for code generation with multilingual benchmarking on HumanEval-X,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.138212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.138212Z digest=sha256:2d24a3427fa586d0bacb08533d8500e1c9fa833590b44268bb16c85fb5931209

Observation 6258d46d-85ec-4677-a4bb-953cb8ea637e · outbound

This paper cites Llama: Open and efficient foundation language models,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Llama: Open and efficient foundation language models,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.140738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.140738Z digest=sha256:712c4d1652ddc1e744087744c2e19f8a2b1ae1fe7eb48dab7e96095709fd213b

Observation b1ab39a1-a3df-475d-bb31-ef1909d9025a · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Pythia: A suite for analyzing large language models across training and scaling,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.143391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.143391Z digest=sha256:a505866c9e53cc5c3d01c99856c0a2a0e2bf039e01ad5ae21fd74e3c26465fca

Observation 5eb97d9f-b6e5-4389-b8c8-40da58628730 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.146017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.146017Z digest=sha256:083c17cea919799e5fa251d8205e27c65dbfd0cea14315e1c8ff6ce4ab24d1c5

Observation d38b566d-63f4-456f-b389-6a858cc0e2fa · outbound

This paper cites Language models are few-shot learners,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Language models are few-shot learners,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.148792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.148792Z digest=sha256:b0ee85ad12d375d08a36e3e4776aba5cf52a987dbe9d5f0ed9fb6b7c1daefd43

Observation 3fd0c0a3-47aa-403d-955c-eb5d81ec8ae7 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.152032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.152032Z digest=sha256:e8c448709f5bb659ac0b3f1d4ab0b0d849048772be38964bb3180103dff631ba

Observation 0e49895d-b10f-489c-9cd3-e4eb4ba82d1c · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.154938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.154938Z digest=sha256:13e45c0e3aef95c0148bccb6105c0f8804fc9a5ce1b68a10024b4c8a071eb650

Observation f2786884-7275-4457-8eec-dd3161cefb43 · outbound

This paper cites Training Compute-Optimal Large Language Models.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Training Compute-Optimal Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.157903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.157903Z digest=sha256:d3e44662a8d0bd83c01ba732f3792d14b0f45f742af35b6c9a5ceea57f21981b

Observation 671319ef-14ac-4f6f-b4be-2e9148ced7ff · outbound

This paper cites GLaM: Efficient scaling of lan- guage models with mixture-of-experts,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models GLaM: Efficient scaling of lan- guage models with mixture-of-experts,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.160712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.160712Z digest=sha256:1a1b020ca01df09fa9793366c16ed8655ffaa2079767bf8abd67659f9c115b6b

Observation 5ab8cee8-0eec-4f59-87b7-36e36fa06d30 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models LaMDA: Language Models for Dialog Applications

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.163443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.163443Z digest=sha256:7070b8261632abf2c5dd1e2abf128dbdd7e119a8b3a2bc7992fcd0985adfee24

Observation 3dfbd059-797e-47a7-b746-58c67c178a86 · outbound

This paper cites Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.166161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.166161Z digest=sha256:be1911424a3b941a0aff43509d76decfd59c258a2f66936893c28f6fbd970286

Observation 4a830d00-7af0-4933-a88f-f9346d932b0b · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models OPT: Open Pre-trained Transformer Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.169086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.169086Z digest=sha256:e6b303f687dd21ba9a62e7123f95d1e6aa8f7c5911fd9b140df75fdef61c14a6

Observation f8f6dc85-ec5c-42d2-8e92-ed43b30c1bb3 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models BloombergGPT: A Large Language Model for Finance

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.172271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.172271Z digest=sha256:5d2f0b78022d394ef702daec705ed0cbafe4acea4b7e14ba6ca27ecc54d2b102

Observation 4b73b2ee-4a4f-44e4-b24c-edbc1ae19fc3 · outbound

This paper cites The Llama 3 Herd of Models.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models The Llama 3 Herd of Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.175437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.175437Z digest=sha256:6e78673844b0b47cb3dea3f1200a402c595755ba0a0fd875935d5b00141ed280

Observation ee7a3292-c3d1-4a62-b7f8-9e5a2048d886 · outbound

This paper cites Qwen Technical Report.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Qwen Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.178384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.178384Z digest=sha256:e050c618b2b4b55b7b4efcbce5ba141f92f7b5537465674145ae2ebb58372f36

Observation ba90cc22-65af-40c1-9b83-523597597de8 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.180958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.180958Z digest=sha256:8e89fbbe9dd043d8b6ca39117b36b8ff399f52fa31478fecc2bbe727c87152d1

Observation 44c00772-270d-4177-8086-3de63693540f · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.183987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.183987Z digest=sha256:9a9a4df0da6147a019a51d78a833b6341c7e7617130627cf4f248bea65c6ccc3

Observation e7603c62-fa59-4b2e-90e8-18c29d0c25b3 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Gemma: Open Models Based on Gemini Research and Technology

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.186821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.186821Z digest=sha256:5f10ac1cbf57b160675fd7980c42a349fb50e46359f04f8e4ad8cebf83f53bf3

Observation 44726ed4-049b-4ad1-9a96-def5fc077e9a · outbound

This paper cites Grok-2 beta release,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Grok-2 beta release,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.189816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.189816Z digest=sha256:c87a8573e478f8925af7774ff3f79025f73c9238f33df64f7fd83880f78e42f3

Observation 9379c27f-8ebb-419c-9e8f-8dc731e1414a · outbound

This paper cites Llama 3.1: An in-depth analysis of the next-generation large language model,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Llama 3.1: An in-depth analysis of the next-generation large language model,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.192786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.192786Z digest=sha256:720ddf381e236936b1cf01314ec7dd0431591df64455ec26aa93b600eb326d52

Observation 862aef8b-caa3-41cf-a7fc-12cb91dbe380 · outbound

This paper cites PredRNN: A recurrent neural network for spatiotemporal predictive learning,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models PredRNN: A recurrent neural network for spatiotemporal predictive learning,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.195263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.195263Z digest=sha256:053aee1f631a8af91eeb48d5476680ba5140fe06560dd3c238f5752f6acbf0d1

Observation 933fd5ff-37ab-463c-911c-e2a4a16dc593 · outbound

This paper cites Efficient and effective training of sparse recurrent neural networks,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Efficient and effective training of sparse recurrent neural networks,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.198436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.198436Z digest=sha256:bced7b74b839dae42521e9054ebce684db886860d2c478c0ee77ca0563f492e6

Observation f254d36e-4046-4813-a601-19801a87718e · outbound

This paper cites Attention is all you need,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Attention is all you need,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.201057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.201057Z digest=sha256:5139d33527b6213f44de0984b65e5cf502874e00097542b8a03ea088b42de26a

Observation 2db19066-3c93-40d1-9c73-2bf8d7a14b65 · outbound

This paper cites A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.204182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.204182Z digest=sha256:8a13433aad114193d683d5ccbf14704757594db55b6d3bac7f0330f3f78f1316

Observation 5dbc411f-0794-47ed-89ec-fbcfafba1714 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Star: Bootstrapping reasoning with reasoning,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.206881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.206881Z digest=sha256:b153f9ad800049fa7c144ee2c2e23799af9036fc381e4b1a9b26db67654729cd

Observation 60d6e2b2-8357-4380-9b77-2c4b9d9fc6b8 · outbound

This paper cites Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.209380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.209380Z digest=sha256:922ab2f9a2d62b17ba9165e53d2127e3a80bde9f1ac95b7b3e81447b25b36a1f

Observation c45363bb-1725-427c-95da-0c9a3ea2b716 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.212052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.212052Z digest=sha256:0435f4aa85dc2ea1304ac18dbf18e245bca59dac67e3f345c5b80df58fd78ce1

Observation 1352d2a4-96dc-40f8-b34c-4790bfb18015 · outbound

This paper cites Making Large Language Models Better Reasoners with Step-Aware Verifier.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Making Large Language Models Better Reasoners with Step-Aware Verifier

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.214785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.214785Z digest=sha256:dc9fa45419c3ccd18e3cbd2097f1975dd716f8059f5bdf4f7dba73f22c262122

Observation 02fd27ff-8c44-4c2b-ba6e-5bb016739e1d · outbound

This paper cites An empirical analysis of compute-optimal inference for problem-solving with language models,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models An empirical analysis of compute-optimal inference for problem-solving with language models,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.217670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.217670Z digest=sha256:fb344107544e2ae472aea340a32e570068ca61f8e856a08b0c49a0d3b4afacc1

Observation e4b5ef6b-ffaa-40e9-8de9-4f232cb26193 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.220308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.220308Z digest=sha256:8ee6e2f9a7d7e83d4af59ede75be18a67b55232169f1e8b9cd24527cefd448ce

Observation 2cae8d5f-4135-4923-bc17-5d42cee0e4d3 · outbound

This paper cites Dynamic programming and stochastic control processes,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Dynamic programming and stochastic control processes,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.223250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.223250Z digest=sha256:7472cd38f9bd78abb614e43dabbdc46815b209395dc5ee321db26b87574bf4e7

Observation f89b80f9-c662-4759-ae6b-f9a7c713db12 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.225928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.225928Z digest=sha256:7398fc1a67779f043ccef2b4f6230914040e2aef2be2704854903bdb2fa3d1ea

Observation 39672797-a393-4fc9-a70a-0f83732d6f1e · outbound

This paper cites Pangu-Agent: A Fine-Tunable Generalist Agent with Structured Reasoning.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Pangu-Agent: A Fine-Tunable Generalist Agent with Structured Reasoning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.228486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.228486Z digest=sha256:afc6000ee4f9cc4d972a1bb7b64c57749df6c848c56e6269d5d3f42f3235db4d

Observation a2ca418b-5482-445a-a33d-b10ff5ea8de7 · outbound

This paper cites Qwen2.5: A party of foundation models,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Qwen2.5: A party of foundation models,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.231791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.231791Z digest=sha256:f4b2c6d717029e5038d69243fe3b056f7d5cd7f928021fde2a444da42f3182ed

Observation ac06e6c2-2968-4c6c-9a9b-bb0660368604 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.234718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.234718Z digest=sha256:0c20b435e5788512252714ad24a0b66e45ea62f004b66fa8ecf9fb255cf8d260

Observation 531becb3-1767-4013-aa74-131cf861ce33 · outbound

This paper cites Visual instruction tuning,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Visual instruction tuning,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.237901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.237901Z digest=sha256:55ac09d84b199077b79ce347de15c211777aae94629ab76f3f442d4e3d0fafc5

Observation 6aa9439b-c1e3-48fc-a7f6-7d03c5e71df0 · outbound

This paper cites Sigmoid loss for language image pre-training,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Sigmoid loss for language image pre-training,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.240685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.240685Z digest=sha256:793c98f59044d697a222ab4f2b80f13c70e54cc2a358c4d6865ac08eb0801b2b

Observation 2f52e35c-5f34-4937-bc79-1a0115ca072e · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.243919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.243919Z digest=sha256:9074352beea2546dfa3b2ba1b3c8d85b4a19aac17b7eb1c1af7bd64052ed37cc

Observation eb68c668-028c-4847-b115-9b70ec0b2552 · outbound

This paper cites Roformer: En- hanced transformer with rotary position embedding,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Roformer: En- hanced transformer with rotary position embedding,

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.246864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.246864Z digest=sha256:9bf6685df1c6434224a49cb22fda3c5ec91ccba3320aebe83bf918be14968dd6

Observation 18f656ea-0aaf-4e33-936c-c30e23dae5ff · outbound

This paper cites MHA-Net: Multipath hybrid attention network for building footprint extraction from high-resolution remote sensing im- agery,.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models MHA-Net: Multipath hybrid attention network for building footprint extraction from high-resolution remote sensing im- agery,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.249549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.249549Z digest=sha256:a494e7a5a0a6ac891d2af2de5b525821f87eb6cef59b950ad3bcc8e2573dd9a4

Observation a1c15df1-44df-4760-8fd7-c7bc65460b6e · outbound

This paper cites Improving Transformers with Dynamically Composable Multi-Head Attention.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Improving Transformers with Dynamically Composable Multi-Head Attention

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.252282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.252282Z digest=sha256:079e4c42c2dc541050be416f0a6f493e33a9c1de2bac18a1b1db055f0ab7af62

Observation f37ec339-0b23-4a03-bf5b-9a5ccd011380 · outbound

This paper cites Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.255004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.255004Z digest=sha256:186287396eb9e4d1c106584cb1a298fc3b6ba54c485145313b07ac052dd1dc71

Observation a585dbb0-4ad6-4bb8-8d12-ddad2ac716d3 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Fast Transformer Decoding: One Write-Head is All You Need

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.258067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.258067Z digest=sha256:21c8e3164440d4cb8acb026faea59d02074ea37dd50bccaa03dbbd1de1d0dc7e

Pith citing papers

No inbound Pith citation observations are available.