Pith. sign in

Paper Citation Record · LEDGER

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

As of 29 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2501.12948.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12948 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-28T06:31:03.373048+00:00

measured 100 of 2227 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:17:51.467877Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9d11db15-e491-4f63-8eca-9369f2313f81 · inbound

Scaling and renormalization in high-dimensional regression cites this paper.

Scaling and renormalization in high-dimensional regression DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T01:55:55.085863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-24T01:54:48.781227Z digest=sha256:6532b0b65988ccaf9f5a1df118f9daa802de63651615c12c67775904ad0eb620

Observation 92948f84-05c2-4d2e-9fcd-845885754947 · inbound

OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework cites this paper.

OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:28:57.146429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-15T03:28:57.008431Z digest=sha256:0d40a227a572e27f92162c6d790066cc52ec8a5dcc8f52fd3108723947977d16

Observation 762eac57-ccd7-48ff-bcfe-646e64a5986e · inbound

Retrieval-Augmented Generation for Natural Language Processing: A Survey cites this paper.

Retrieval-Augmented Generation for Natural Language Processing: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-23T23:08:35.723552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T23:06:41.081461Z digest=sha256:f61a0c90602ec0712fddc98afb499caa85b57c75c74d6031f84e300b5b6b34fd

Observation 27a7bd9b-26cd-48c0-bf15-ddcedfb63719 · inbound

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities cites this paper.

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:16:04.675720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-17T22:16:04.386706Z digest=sha256:9f16007535b3567849942bbdbfd95d68ddebb8e95815a76d174d26ea3c8e14dd

Observation 74586e5a-1166-476f-ab26-1e2418675231 · inbound

Enhancing Clinical Trial Patient Matching through Knowledge Augmentation and Reasoning with Multi-Agent cites this paper.

Enhancing Clinical Trial Patient Matching through Knowledge Augmentation and Reasoning with Multi-Agent DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T17:03:12.477184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T17:01:30.638046Z digest=sha256:553e6fd0dae6403e052d7cbea403fdfe22e6c1b874d5afbd3b9c5d89d793f2a9

Observation 6a332db2-c5c0-4cf7-943e-6ebdfb87ea26 · inbound

Training Large Language Models to Reason in a Continuous Latent Space cites this paper.

Training Large Language Models to Reason in a Continuous Latent Space DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:29:05.677264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-11T10:29:05.384381Z digest=sha256:2bf1c5d9f835c1d1ad21e6ef05c0c2ca5915a553e6753ad2dfc4ab749ea3dcba

Observation 206a5d9f-1c6e-4102-880a-1b846521b70b · inbound

SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation cites this paper.

SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:22:43.149402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T07:20:57.953573Z digest=sha256:22c18ff8ef3153ee8f951c8a1ffae9fe2cf2f0a12c11d4804746b22ca81b4f4b

Observation 1776ca05-fe61-4453-90f9-f98e4894e39e · inbound

Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation cites this paper.

Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:15:33.975137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-05-25T08:12:15.133695Z digest=sha256:f1d70df3fd56fa7ba213db65026649c0d93e2f158937fa38102c3766727c87f3

Observation cfe1a6af-393e-4e2e-b8f2-dee6f9727076 · inbound

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs cites this paper.

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T15:51:29.231425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-05-13T15:51:29.022336Z digest=sha256:5269936b4d9abe26926146c9941a7ffa13c81634cc5fa873bae9907ccbff08de

Observation b707e1b1-845e-4c85-8edb-415edac1f881 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:02:37.448157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:e0c6e402fe726e3a888826d008481f7aae794008cdfb4a8c04e459da065aecbf

Observation 8767a48a-0425-4028-a7f3-803aa0e0d932 · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:23:30.917683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:dffe6a64beb2754c3e1e8db36dab88f05149e841df16c264ef4cd262bd677e39

Observation efd852f4-5519-4f3b-86e5-9a19c3a810d1 · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:17:31.228380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:5a27b7269e2f15ff64acb2cae2b4eb434fd0b4824cb73275273ecd6d9ccddccb

Observation 32bd76dd-b6f7-4929-bc4c-c2f67ca14f06 · inbound

Distributional Statistics Restore Training Data Auditability in One-step Distilled Diffusion Models cites this paper.

Distributional Statistics Restore Training Data Auditability in One-step Distilled Diffusion Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:32:34.073158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T04:27:59.317818Z digest=sha256:02041ec46dfc05d9ecd50f36b0c560c44cbb27102d332e7ffd9ec3021761779f

Observation 3cf708f7-52c2-48eb-b279-124bee20a0d3 · inbound

LIMO: Less is More for Reasoning cites this paper.

LIMO: Less is More for Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 130

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:11:37.552410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-05-17T02:11:36.932541Z digest=sha256:7509159f931f559354c5b62957aa828de850c0e31e574f2cbe3bf5ec4b4ca255

Observation 9a1ab93c-6bbf-442c-80fa-e03deec7a1a1 · inbound

Large Language Models for Multi-Robot Systems: A Survey cites this paper.

Large Language Models for Multi-Robot Systems: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:32:32.071591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T04:32:05.138744Z digest=sha256:7298dbed975f07a6e90ffddb33d7f500cdb9d4dfea9f9975a7bbbf09b0567299

Observation 38cfc498-da59-42ea-9fa1-cd41c6c5a4cc · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T15:39:41.085481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:1e0030e8fe7bc7d54bbf9c97231a96051ef6795e59d634132747803309b4be48

Observation cbcc6279-61f2-4c9a-bc04-55a63117a991 · inbound

Large Language Diffusion Models cites this paper.

Large Language Diffusion Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:42:54.763954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-11T01:42:54.279353Z digest=sha256:4e7f11fdd542dd66a54fa742510da136216311296de489f756a37e32f6262ae2

Observation f6b0c620-dc27-4df5-b667-8f25f95a4d3e · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.544061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:d9d63b05c3ab2997a075a9b6c0b7ac661e1a4e1e73d2ba2da5b183d74cee551c

Observation 4c454568-6b64-4975-afca-cc83155628b1 · inbound

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention cites this paper.

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:46:30.038088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-05-16T23:46:29.975858Z digest=sha256:61c45d854b80c1aad4d02f5e2d206b0b65da911417b8148ac1dc4b968398705e

Observation 7ca610c1-2904-4f2a-b3e5-93500c54d47b · inbound

A-MEM: Agentic Memory for LLM Agents cites this paper.

A-MEM: Agentic Memory for LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:47:28.994970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-11T00:47:28.898112Z digest=sha256:9e06e5e9153ca2669b499ee814f0dc725f8b49d4a809f5e1fb585a0658a83aec

Observation 11497c5e-639e-4a54-b725-c48e6ac33320 · inbound

Hallucinations are inevitable but can be made statistically negligible cites this paper.

Hallucinations are inevitable but can be made statistically negligible DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:25.843335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T02:41:59.640979Z digest=sha256:3b6d86dcd1170b1e7dff647917b422551af28cda4e5beb23c4f4bae70a80d324

Observation 4f7f7a4c-d5f1-4184-8a17-8f69d175f469 · inbound

Learning to Reason at the Frontier of Learnability cites this paper.

Learning to Reason at the Frontier of Learnability DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:42:26.069190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T02:41:21.571824Z digest=sha256:fe017a466f5c9d3231ed638f9a153be9b26ae5ef3bf68df0467178f47634be33

Observation 44924d88-cb1c-4f33-938c-4cb66f2ec26c · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:15:46.162426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:32ab80b38bdc2cc8a6e1e7bd1af8fd98cffb1a42a2b64d3f76914d91d1c44a4b

Observation 7f780ee7-29de-460e-9c07-66b590fa0526 · inbound

Supervising the search process produces reliable and generalizable information-seeking agents cites this paper.

Supervising the search process produces reliable and generalizable information-seeking agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:22:25.441515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T02:18:27.204122Z digest=sha256:ea8abd2b294dd24e14af548c2845c4f1e8f15b54bdf988f26970a2a99cfb6c1b

Observation 84b52af0-bc9d-4b88-8ac3-3f9fd90962fa · inbound

Tokenizing Single-Channel EEG with Time-Frequency Motif Learning cites this paper.

Tokenizing Single-Channel EEG with Time-Frequency Motif Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:52:22.898183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T01:51:12.410547Z digest=sha256:36931c066a80b983d6386f9c02a1bf94f33915869a8c7805697e61e772ebaa16

Observation 8555a31e-9eef-4501-81e5-0bff8f3eb4fe · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:36:24.570662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:b27521d361d5d9438e5768420be5a7eb226a0fc37740804807942c81748ad251

Observation 8d8df19b-750f-4734-9816-e2ac6d4dd1b3 · inbound

LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification cites this paper.

LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:57:22.756536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T01:57:19.398245Z digest=sha256:d0ce39e3d7263780a7edfadec48e99324fb3b3262133378aca984891d0793ebe

Observation aea0d5fa-1a47-408f-b966-5346630614bc · inbound

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution cites this paper.

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:27:56.335474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-05-15T10:27:56.185943Z digest=sha256:b0eb10dcd1fff63f7a60c88ed4c534d3a813bf9ff78ac05c6ce63029fc1108e4

Observation 36fc85fc-539e-4463-9562-b0c8006de306 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 242

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:02:45.420842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:10ca845463ee51dbb4ea2adf168ef4470802386eee4893534e61a81f0a4437dc

Observation 8a9b378f-7981-4cff-ba94-67a27c73d481 · inbound

Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models cites this paper.

Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:47:26.380429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T02:46:20.578680Z digest=sha256:7bbc9680895ef024c27c733582e785a5a6175efb501ccc919f8f810cff4e8584

Observation 603a0f70-7dc0-45ef-acb9-4082d8366ef4 · inbound

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs cites this paper.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:22:28.002945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:a42b3ef3ed46d395b118f4026bd9c2c659414837981d83e814e9d21ee7e0c908

Observation 116978b6-fa7d-444b-9bb9-b75b9f716a96 · inbound

Visual-RFT: Visual Reinforcement Fine-Tuning cites this paper.

Visual-RFT: Visual Reinforcement Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:16:16.557271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-13T22:16:16.528682Z digest=sha256:9a85c2290f35fae08fa071e8b3b3524bb05543a2eddaecbd78188cc410d95fbe

Observation 90502821-c995-4e39-828a-1d268bdfcf66 · inbound

LLM-TabLogic: Preserving Inter-Column Logical Relationships in Synthetic Tabular Data via Prompt-Guided Latent Diffusion cites this paper.

LLM-TabLogic: Preserving Inter-Column Logical Relationships in Synthetic Tabular Data via Prompt-Guided Latent Diffusion DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T01:07:19.886029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T01:05:38.969311Z digest=sha256:412537e21825f8aaadf4875531a622857755bec01482f66e5f3301652d56898e

Observation 1dd3c76e-bd63-468f-9143-dfa452f4a92c · inbound

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning cites this paper.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:19:22.201688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:d53e192ce7a25ef718dd071f1a22825d17e19dca5b9f12056954b2474612e2d3

Observation d149f89e-0910-4480-bdbb-6328b9b2f829 · inbound

Unified Reward Model for Multimodal Understanding and Generation cites this paper.

Unified Reward Model for Multimodal Understanding and Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:44:30.828729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-14T00:44:30.558048Z digest=sha256:65f15195cb9fca875b5c62da43f12420408c2abb4fc4e10d52fa9db4142a8fc8

Observation 54605935-9cbe-4788-9153-4689e2ae971b · inbound

R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning cites this paper.

R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:37:17.756376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-13T18:37:17.725984Z digest=sha256:4a3f1ab75fea56cc39ac4e960b140a800d4654fcaa177b793631ff564bd0196d

Observation ad47c6a7-71a2-478d-b321-b7802d3d672c · inbound

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement cites this paper.

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:31:43.568223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-16T12:31:43.494099Z digest=sha256:23b5edd4cfd80b85d76a42cfc2991ea562c6b9675613e5c836b2c9ef745bca5d

Observation 67bd6635-cb2d-48ec-9332-2d3ca94b3429 · inbound

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL cites this paper.

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:15:46.399734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-16T15:15:46.255296Z digest=sha256:35ed026b6cd8e388ac007be037a486b03aecebcc5f1aa596975f84678c310303

Observation 7d7f2f93-19a2-46c9-af1a-2229716c3e92 · inbound

AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning cites this paper.

AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:06:27.200708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-16T20:06:27.136345Z digest=sha256:df24bb17217cd24808dc27f4fdf960c35bd967dbcdfabcd5c1f365704123b243

Observation 35165ca8-d2ec-4554-95ea-87583becc981 · inbound

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO cites this paper.

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:22:18.727457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T00:21:51.621582Z digest=sha256:6901550773ec2c80fbb36ff4f7def5a8c16eb4cae62f75bfdf93f0e2d49d55f4

Observation 74d9fc96-deb9-4561-b7fd-adf08d6711fb · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 231

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:40:41.437202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:a195b659e738b37dc56cf0af202554c91d6c85be2260851e46a59878de1c7a1e

Observation 0732cb83-ecc1-484e-a2c3-88b974073d01 · inbound

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks cites this paper.

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:32:18.650809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-17T21:32:18.491541Z digest=sha256:c8d7de5c5f7662fd52423abf6d74992c9c40106532a828c0ebb1f9ee6aa1a901

Observation 971aa50b-a162-4170-82f0-7306414ca55a · inbound

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization cites this paper.

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T00:19:20.514762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-16T00:19:20.462455Z digest=sha256:b0c6560889118de306f7ecaf4b2f7a06bea50164a96d8d6457792fb4d0a56e39

Observation 85e0a5c0-80ae-4228-a69f-fe56867a6191 · inbound

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation cites this paper.

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:24:12.926621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-21T07:24:12.845841Z digest=sha256:c147d2cce342d214e2e3209cf444eb3941caf9e39a49fb545e247249aa02759d

Observation f962c689-0269-4631-9644-d409a64b5968 · inbound

Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub Scenarios cites this paper.

Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub Scenarios DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:52:19.444673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-23T00:50:05.344148Z digest=sha256:5c80c095ad91fc31a55780e8da08bd866399411071a4b59cfb148fd2287a8e13

Observation d96d1e8e-4f65-4208-ae60-924eb7d998fb · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:18:53.213897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:1d6f760f4c2704d75fc99156a6f21e814bd7b242bf5ca981291e905dedd6da3f

Observation bd13eaa7-9b6a-4cac-849b-0abf7dd9a7b5 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:04:22.832542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:c0d2a2df282abd48ced9254796c10557c684c69e483cb04e70ddc6c9718ffa08

Observation 096c1a33-4843-4c20-92da-e0ce2affb49c · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:40:06.573742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:843f3ec74d4d1a6bb86fef295829904d26bb6130caefa9c620801df033684b5d

Observation ea8125e8-dd16-432a-a319-8b2672e690c2 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:02:17.859435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:cb315cb76d56f509561a4dc394b53ec0e8a19495672f987474ae629c072f6b3f

Observation 82c25e79-80e3-4b0f-b478-56ec6b11807a · inbound

DAPO: An Open-Source LLM Reinforcement Learning System at Scale cites this paper.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:35:13.475103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:99995cb26c7f9fa8598e299cc45e57b221b92135b5f009021be4fa5f10fb446f

Observation 6802f323-9631-45de-ac66-59b5ea8423a7 · inbound

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning cites this paper.

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:47:10.211498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-16T12:47:10.146795Z digest=sha256:c584d52eac46b4d0838f99009c435908a0d43767c2f14747c94c334ff599581c

Observation 92cc2cb5-b885-4f17-893d-2f4b9dcb799a · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:29:57.336032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:180275d1e5b8fbcae7a93000678efe799a551a1830d7a03db46adc9b0d4fea2e

Observation 11f8cdf8-21c4-44ef-a526-4bca0cf0b6f6 · inbound

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems cites this paper.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:57:13.221779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T22:55:34.238427Z digest=sha256:1fd613d46907b6187944574be1190d4ea0b8c8adcf33fff5de4b3753ad8a2c00

Observation 1e93d0ff-364c-4c1c-8112-1ce67243dcc3 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:59:03.392088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:6d40a5a69ff536dce0661200042c3aeb4868ab8968729c9e95b40da574b58da3

Observation 68e98038-2889-4a71-9910-6bc93cbc006c · inbound

Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark cites this paper.

Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:42:16.458629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T23:40:25.111680Z digest=sha256:58cfa0210bf0e0147ee6e5d24768f2b2f13b0ff9bc782cd0f39f78cae64a2354

Observation 99ae216b-12e2-47a8-b664-f60cedaed612 · inbound

ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning cites this paper.

ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:48:34.759536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-16T15:48:34.705881Z digest=sha256:60f59c7f211f5f5c39bfed4ef9f7455c73b8c19689c3d1b5ff6ebbf3343a7603

Observation 35bee7cb-f4ae-47c3-9f8c-c53efb43866f · inbound

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models cites this paper.

Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:15:11.202017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T22:14:53.522206Z digest=sha256:a283868adc1d8a9ac79c6766ed77ff4eec24b0ef8f54b6c5cc720ee8e1a8dbbd

Observation 738d9a85-dec2-4ef6-b95b-21871b661bb4 · inbound

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning cites this paper.

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:02:41.372022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-16T11:02:41.335059Z digest=sha256:62e161b5b6f494c2d6773f3ea21c5d9ab4679d552919275fa1a61683e11d2583

Observation be9eb4e3-65b2-4cf1-b05b-22bdbf1e52b1 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:43:00.371525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:7e22fcfe82858e8b151289bf188c031939996cad73fdcb510abe4f5df3b13f77

Observation 3b310225-95b6-406f-b598-7f7f51011791 · inbound

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? cites this paper.

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:42:13.491102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T22:38:35.969273Z digest=sha256:88a43997b4342104a5044aef516ded61d1e16faebb67b9553cbaa2ed5549e09b

Observation bfa048df-635d-44b4-8afe-d09eac9af652 · inbound

Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model cites this paper.

Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T21:59:02.206491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-12T21:59:02.173820Z digest=sha256:113d69551a6335ee52453f8c9c67f4244ce94b9c715d5f571acbaa9f3a88f398

Observation cfaa2480-7902-44c2-ab06-62f7dda100ca · inbound

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning cites this paper.

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:18:43.778898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-15T15:18:43.724432Z digest=sha256:819283f6030f7b3f684f74162c136ba7cf78271af881ca90d6e7faa5095b2e8a

Observation a2e71d8e-7606-43d6-80e6-042a3c4bd402 · inbound

OpenCodeReasoning: Advancing Data Distillation for Competitive Coding cites this paper.

OpenCodeReasoning: Advancing Data Distillation for Competitive Coding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T19:21:42.194865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-17T19:21:42.081762Z digest=sha256:c11f55826b5f16d224ee89fc54d68a1f1d15478a71887db1c7f169616763e3fa

Observation 2ce22385-4116-418f-bb7e-be2b3ffc90df · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 119

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:42:10.935652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:9b2a89ae194aefb7430968fcde5031b74545b2835b2c33da898729e8dc78fde2

Observation bd97f8b2-89f7-4e8c-93fe-d37092050f0b · inbound

A Survey of Scaling in Large Language Model Reasoning cites this paper.

A Survey of Scaling in Large Language Model Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:22:09.055779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T21:20:07.238992Z digest=sha256:c0443a58053fdfb2b305e06bbc8c0cb0178b1661887a99c049e2079551e60c47

Observation 4f456f52-de83-49c3-8b8e-8e5b282cc5e4 · inbound

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks cites this paper.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T09:36:04.755391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:525e542222cf44fa98247d415ed6cbf7b7ef59bdd862661c3e3520af8be2f628

Observation 8dfef9c8-9d04-43a7-9dbc-8400afbf2c66 · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:23:51.684873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:acb8844635ce1316d5e9a0f10cd6126d48a350aaab1c9a60e2a6e58ef8dc019e

Observation 4280e525-0f88-4f93-b04a-b1fa86727d1d · inbound

ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs cites this paper.

ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:12:08.328703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T21:11:46.405944Z digest=sha256:7e1b158c18ffa2738c19dc85a16146b7d7c3076c31456445119e36e516827df9

Observation 0949eb1c-3dcc-4606-8923-aed3099d416c · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:56:07.786208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:981286bf45ac761d13acdc3fd376b3b6cfd02bf3d45250832ee5deeb6b7bd5da

Observation d65c834a-2372-447d-9e05-dbe0880a3c18 · inbound

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model cites this paper.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:13:57.420952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:e04a7e99acb1b1278a1a1dd11b86476f7489c813659d27e56289b0adcf433461

Observation ee895e6a-0b4b-4093-a63d-638f1b070568 · inbound

Exploring the System 1 Thinking Capability of Large Reasoning Models cites this paper.

Exploring the System 1 Thinking Capability of Large Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T20:07:02.597495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T20:05:13.849797Z digest=sha256:c39e8557861c97ac67fe74ed3bb6dd8f9cd1a0a0d6eed35774d83507fcfd157e

Observation 20418ee8-95e6-451f-97c9-cb463c22b704 · inbound

GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents cites this paper.

GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:10:58.123955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-15T02:10:57.976448Z digest=sha256:ddd7750e3fb075475e45df451f6ceb13dcfde60a44fbd779bb5c48308e6fe2de

Observation 995d756a-08af-4938-8ce5-992406dc104c · inbound

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning cites this paper.

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:31:04.799231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-16T10:31:04.728005Z digest=sha256:6756a2e5bac98e308f9984e02c77cf60874bd503f52bb9a85dfaf03bb31cd99d

Observation 63eae499-6855-45f0-9437-3404fb437802 · inbound

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs cites this paper.

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:42:39.053511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-13T18:42:39.023650Z digest=sha256:c3e0f7dbc88993e3ed855face58878742b4875d44eb541dba3029bf0866a79a9

Observation d8cef190-a21f-45ea-8c0f-5ddaead9f219 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:32:01.362542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:4b7e547aeb02eadeadafd62d0dd249a62a03d9002961bc0cde5ff2fe03a7e973

Observation 3f634a63-c95a-4e6b-a10d-c6a773c21517 · inbound

Design Topological Materials by Reinforcement Fine-Tuned Generative Model cites this paper.

Design Topological Materials by Reinforcement Fine-Tuned Generative Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:16:58.685680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T19:15:26.493687Z digest=sha256:56e81ccbcfc5fac0a80cbf16944fb272b8edeb6fae56996d6749e3ff58813e81

Observation 59837342-2064-481c-b8cb-4758d66da240 · inbound

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning cites this paper.

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:46:57.178986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T18:46:11.571831Z digest=sha256:7a2061b1369fab5648b9b985b81dcb6a386607d3f6b8df44221b38b357b2b78b

Observation 0c311622-675b-42de-9781-55f79ba65698 · inbound

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning cites this paper.

Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:15:09.400280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T21:14:13.351140Z digest=sha256:ae47ddaa0fb6f62870332d217dd65805303936dfc8ae7b9c79608923c18e1dfc

Observation ba59a3e8-a1ce-413b-a6f7-379e4d7c8f52 · inbound

ToolRL: Reward is All Tool Learning Needs cites this paper.

ToolRL: Reward is All Tool Learning Needs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T00:26:48.396196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-14T00:26:48.291431Z digest=sha256:7a97371fdd1607b7c75f4d68cba86d5cc832854d15b958cb2e37d487fb448692

Observation d4e41a9c-3bc4-432c-85fd-f29cb7206a30 · inbound

Learning to Reason under Off-Policy Guidance cites this paper.

Learning to Reason under Off-Policy Guidance DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:17:02.754382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-15T23:17:02.701393Z digest=sha256:7ffe17dc9535c7100657aa46ff77723290f2f2a2450408615993dfa96997ab56

Observation ddef81f6-addb-460a-8034-11e9c34cc2ee · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:36:58.742184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:18a4eb08171b62f1a968d1a33e12d99b267847b3598a5955e17a61348fe6bb8b

Observation c061af56-75aa-4fb8-a8f1-a3c5ab10911a · inbound

BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese cites this paper.

BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T22:04:49.942408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-17T22:04:49.915916Z digest=sha256:2ff8873be3eef610d776b947d3db8d0bbb71b1b44a302618882e0a495c32a58a

Observation 009483ec-02f4-4879-afa0-17775769c5d3 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:57:37.963947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:43f95798d39ba481033de1214dfb57c3b087d852918699e59a68693532291adf

Observation 843fc5fc-219d-41b6-8d90-e2ff72eedd0a · inbound

Reinforcement Learning for Reasoning in Large Language Models with One Training Example cites this paper.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:51:04.844671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:01b68b0bf0903205d3818cc437d856f6d6edaaa1416c07ca98b16640d080032f

Observation caaa20f3-36fb-4418-b8d3-d12c69e68d78 · inbound

Phi-4-reasoning Technical Report cites this paper.

Phi-4-reasoning Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.767220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:a57d618a417440d710b0f055d569df09e9ef7ed5113598aeb61f7b0869a91f4d

Observation 6479e86a-1f1e-4d7a-b287-167f95917dc3 · inbound

DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition cites this paper.

DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:32:05.672242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-15T09:32:05.642611Z digest=sha256:562fcff56cb9e9adbf1fad2953bd3d91086d7ec128ae8f9ce85088d8227bd246

Observation dd04b8e8-a459-453c-b41b-44d68b16a649 · inbound

Always Tell Me The Odds: Fine-grained Conditional Probability Estimation cites this paper.

Always Tell Me The Odds: Fine-grained Conditional Probability Estimation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T16:31:46.933633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:eb0d7baec37061f771825fe95c6b6e03110b1bcb161b78cc4ea38ad2e192b363

Observation 0322336e-a359-49ae-b64a-48297f7bbb6f · inbound

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference cites this paper.

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-09T06:35:25.153201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T15:59:04.724780Z digest=sha256:c9c84376239a02fa40b714bdc940d332e6a6cf761f9b92b9d7011017cc598e7e

Observation dacaf138-2ff8-4cdf-96be-7db464020e06 · inbound

SocialLM: Social Signal Processing of Patient-Provider Communication using LLMs and Contextual Aggregation cites this paper.

SocialLM: Social Signal Processing of Patient-Provider Communication using LLMs and Contextual Aggregation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T17:05:00.355095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T17:04:00.508789Z digest=sha256:1ba39663ad25a8dd8f952a7b2dbbefc837793c0ebab798629c7095e347e2c416

Observation ce75d320-2c9b-4c7a-8e03-01b105ceaa31 · inbound

Flow-GRPO: Training Flow Matching Models via Online RL cites this paper.

Flow-GRPO: Training Flow Matching Models via Online RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:45:17.030802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-11T18:45:16.641012Z digest=sha256:ab8472b1f302165185e82ad8422013045b8035d43c670f408e88d8710090c59a

Observation 2488832d-2931-470f-9507-c33112e66903 · inbound

LLMs Get Lost In Multi-Turn Conversation cites this paper.

LLMs Get Lost In Multi-Turn Conversation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:11:09.174697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-14T00:57:10.262350Z digest=sha256:765ae119c1f687437227d462c37b708526e8cab893696b6bb6d1d619ab887186

Observation a7dd1a63-4a2e-42ab-bd40-8e5d44fd080b · inbound

A Survey on Foundation Models for Personalized Federated Intelligence cites this paper.

A Survey on Foundation Models for Personalized Federated Intelligence DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:34:57.704097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T15:32:15.293888Z digest=sha256:ca8689fe1674e52537a6c9d7dce3e696d152e5e931948cd401e24c857ba81024

Observation 9ffc2f68-d28e-4814-b09c-179010b8042f · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.042133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:cae1aac95932110a7840e802ea511e4ac76f46dfe900dddc8a0037a0f8cae4ad

Observation 5421bf75-b042-4d23-904e-9e889de9bd6d · inbound

Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning cites this paper.

Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:51:45.690186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T15:49:44.263123Z digest=sha256:a7f72af361b5dbdf03e34d468c76ccb1243fd541dfb99d106ed14812ed4070e5

Observation 11809140-14c3-4f97-ae15-67bf556b4edb · inbound

DanceGRPO: Unleashing GRPO on Visual Generation cites this paper.

DanceGRPO: Unleashing GRPO on Visual Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:28:26.684350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-11T22:28:24.929046Z digest=sha256:eb7958a9d4a11291f8aec834c7a747aeb017a9fd4c14b42390c58ccdabe001d2

Observation 76b95b6f-5f57-413a-b0db-a92597925aed · inbound

Not that Groove: Zero-Shot Symbolic Music Editing cites this paper.

Not that Groove: Zero-Shot Symbolic Music Editing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T16:31:47.462937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T16:27:44.903349Z digest=sha256:f6612bd42174394dca0fd763426eb4b1388ce963931d8e5348d09df515ef8392

Observation de132211-36e5-4928-943e-0d0bfd0276ff · inbound

Qwen3 Technical Report cites this paper.

Qwen3 Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-09T06:35:28.514089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-09T06:35:27.813995Z digest=sha256:9bbd89180b6fc0bc9c92d6e190e7e6d996dc367eded6067fb5fd6e280d23bdd4

Observation 757208c1-62b3-4bbd-b38f-eb9bd58d792f · inbound

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production cites this paper.

ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:44:58.108944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T15:42:05.266854Z digest=sha256:5180e0e58d1d266bf74ef7ae0af2a3d6c7452396c6bd8be4b85cf2ef382e4a38

Observation 73db4015-62f4-4f9e-8870-f6bc33d7eda8 · inbound

Superposition Yields Robust Neural Scaling cites this paper.

Superposition Yields Robust Neural Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 63

Resolution
malformed identifier
local_arxiv, observed 2026-05-09T06:36:25.553937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T14:38:44.789822Z digest=sha256:823fba5fa7d82ebae13db4b99617216b5fdde0c956c0a78ac98323c2f27e145d

Observation d3df8d7b-dcd4-451e-8682-aedd51add6ac · inbound

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction cites this paper.

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:21:45.070163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.

source=pdf_text observed=2026-05-22T15:18:15.475294Z digest=sha256:fe71bc7049d9967b0c6cd2b94303f522e42c46aaebcd487d52c6c78ef8c1d4d7