Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 29 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2501.12948.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-28T06:31:03.373048+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T03:17:51.467877Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 9d11db15-e491-4f63-8eca-9369f2313f81 · inbound
Scaling and renormalization in high-dimensional regression DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 92948f84-05c2-4d2e-9fcd-845885754947 · inbound
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 762eac57-ccd7-48ff-bcfe-646e64a5986e · inbound
Retrieval-Augmented Generation for Natural Language Processing: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 27a7bd9b-26cd-48c0-bf15-ddcedfb63719 · inbound
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 74586e5a-1166-476f-ab26-1e2418675231 · inbound
Enhancing Clinical Trial Patient Matching through Knowledge Augmentation and Reasoning with Multi-Agent DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 6a332db2-c5c0-4cf7-943e-6ebdfb87ea26 · inbound
Training Large Language Models to Reason in a Continuous Latent Space DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 206a5d9f-1c6e-4102-880a-1b846521b70b · inbound
SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 1776ca05-fe61-4453-90f9-f98e4894e39e · inbound
Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation cfe1a6af-393e-4e2e-b8f2-dee6f9727076 · inbound
Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation b707e1b1-845e-4c85-8edb-415edac1f881 · inbound
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 8767a48a-0425-4028-a7f3-803aa0e0d932 · inbound
Process Reinforcement through Implicit Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation efd852f4-5519-4f3b-86e5-9a19c3a810d1 · inbound
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 32bd76dd-b6f7-4929-bc4c-c2f67ca14f06 · inbound
Distributional Statistics Restore Training Data Auditability in One-step Distilled Diffusion Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 3cf708f7-52c2-48eb-b279-124bee20a0d3 · inbound
LIMO: Less is More for Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 9a1ab93c-6bbf-442c-80fa-e03deec7a1a1 · inbound
Large Language Models for Multi-Robot Systems: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 38cfc498-da59-42ea-9fa1-cd41c6c5a4cc · inbound
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation cbcc6279-61f2-4c9a-bc04-55a63117a991 · inbound
Large Language Diffusion Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation f6b0c620-dc27-4df5-b667-8f25f95a4d3e · inbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 4c454568-6b64-4975-afca-cc83155628b1 · inbound
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 7ca610c1-2904-4f2a-b3e5-93500c54d47b · inbound
A-MEM: Agentic Memory for LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 11497c5e-639e-4a54-b725-c48e6ac33320 · inbound
Hallucinations are inevitable but can be made statistically negligible DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 4f7f7a4c-d5f1-4184-8a17-8f69d175f469 · inbound
Learning to Reason at the Frontier of Learnability DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 44924d88-cb1c-4f33-938c-4cb66f2ec26c · inbound
MoBA: Mixture of Block Attention for Long-Context LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 7f780ee7-29de-460e-9c07-66b590fa0526 · inbound
Supervising the search process produces reliable and generalizable information-seeking agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 84b52af0-bc9d-4b88-8ac3-3f9fd90962fa · inbound
Tokenizing Single-Channel EEG with Time-Frequency Motif Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 8555a31e-9eef-4501-81e5-0bff8f3eb4fe · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 8d8df19b-750f-4734-9816-e2ac6d4dd1b3 · inbound
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation aea0d5fa-1a47-408f-b966-5346630614bc · inbound
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 36fc85fc-539e-4463-9562-b0c8006de306 · inbound
Towards an AI co-scientist DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 242
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 8a9b378f-7981-4cff-ba94-67a27c73d481 · inbound
Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 603a0f70-7dc0-45ef-acb9-4082d8366ef4 · inbound
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 116978b6-fa7d-444b-9bb9-b75b9f716a96 · inbound
Visual-RFT: Visual Reinforcement Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 90502821-c995-4e39-828a-1d268bdfcf66 · inbound
LLM-TabLogic: Preserving Inter-Column Logical Relationships in Synthetic Tabular Data via Prompt-Guided Latent Diffusion DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 1dd3c76e-bd63-468f-9143-dfa452f4a92c · inbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation d149f89e-0910-4480-bdbb-6328b9b2f829 · inbound
Unified Reward Model for Multimodal Understanding and Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 54605935-9cbe-4788-9153-4689e2ae971b · inbound
R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation ad47c6a7-71a2-478d-b321-b7802d3d672c · inbound
Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 67bd6635-cb2d-48ec-9332-2d3ca94b3429 · inbound
LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 7d7f2f93-19a2-46c9-af1a-2229716c3e92 · inbound
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 35165ca8-d2ec-4554-95ea-87583becc981 · inbound
FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 74d9fc96-deb9-4561-b7fd-adf08d6711fb · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 231
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 0732cb83-ecc1-484e-a2c3-88b974073d01 · inbound
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 971aa50b-a162-4170-82f0-7306414ca55a · inbound
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 85e0a5c0-80ae-4228-a69f-fe56867a6191 · inbound
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation f962c689-0269-4631-9644-d409a64b5968 · inbound
Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub Scenarios DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation d96d1e8e-4f65-4208-ae60-924eb7d998fb · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 137
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation bd13eaa7-9b6a-4cac-849b-0abf7dd9a7b5 · inbound
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 096c1a33-4843-4c20-92da-e0ce2affb49c · inbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation ea8125e8-dd16-432a-a319-8b2672e690c2 · inbound
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 82c25e79-80e3-4b0f-b478-56ec6b11807a · inbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 6802f323-9631-45de-ac66-59b5ea8423a7 · inbound
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 92cc2cb5-b885-4f17-893d-2f4b9dcb799a · inbound
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 11f8cdf8-21c4-44ef-a526-4bca0cf0b6f6 · inbound
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 1e93d0ff-364c-4c1c-8112-1ce67243dcc3 · inbound
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 68e98038-2889-4a71-9910-6bc93cbc006c · inbound
Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 99ae216b-12e2-47a8-b664-f60cedaed612 · inbound
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 35bee7cb-f4ae-47c3-9f8c-c53efb43866f · inbound
Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 738d9a85-dec2-4ef6-b95b-21871b661bb4 · inbound
UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation be9eb4e3-65b2-4cf1-b05b-22bdbf1e52b1 · inbound
Video-R1: Reinforcing Video Reasoning in MLLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 3b310225-95b6-406f-b598-7f7f51011791 · inbound
When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation bfa048df-635d-44b4-8afe-d09eac9af652 · inbound
Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation cfaa2480-7902-44c2-ab06-62f7dda100ca · inbound
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation a2e71d8e-7606-43d6-80e6-042a3c4bd402 · inbound
OpenCodeReasoning: Advancing Data Distillation for Competitive Coding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 2ce22385-4116-418f-bb7e-be2b3ffc90df · inbound
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation bd97f8b2-89f7-4e8c-93fe-d37092050f0b · inbound
A Survey of Scaling in Large Language Model Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 4f456f52-de83-49c3-8b8e-8e5b282cc5e4 · inbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 8dfef9c8-9d04-43a7-9dbc-8400afbf2c66 · inbound
SmolVLM: Redefining small and efficient multimodal models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 4280e525-0f88-4f93-b04a-b1fa86727d1d · inbound
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 0949eb1c-3dcc-4606-8923-aed3099d416c · inbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation d65c834a-2372-447d-9e05-dbe0880a3c18 · inbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation ee895e6a-0b4b-4093-a63d-638f1b070568 · inbound
Exploring the System 1 Thinking Capability of Large Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 20418ee8-95e6-451f-97c9-cb463c22b704 · inbound
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 995d756a-08af-4938-8ce5-992406dc104c · inbound
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 63eae499-6855-45f0-9437-3404fb437802 · inbound
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation d8cef190-a21f-45ea-8c0f-5ddaead9f219 · inbound
Reinforcement Learning from Human Feedback DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 3f634a63-c95a-4e6b-a10d-c6a773c21517 · inbound
Design Topological Materials by Reinforcement Fine-Tuned Generative Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 59837342-2064-481c-b8cb-4758d66da240 · inbound
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 0c311622-675b-42de-9781-55f79ba65698 · inbound
Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation ba59a3e8-a1ce-413b-a6f7-379e4d7c8f52 · inbound
ToolRL: Reward is All Tool Learning Needs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation d4e41a9c-3bc4-432c-85fd-f29cb7206a30 · inbound
Learning to Reason under Off-Policy Guidance DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation ddef81f6-addb-460a-8034-11e9c34cc2ee · inbound
PRIMETIME : Limits of LLMs in Temporal Primitives DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation c061af56-75aa-4fb8-a8f1-a3c5ab10911a · inbound
BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 009483ec-02f4-4879-afa0-17775769c5d3 · inbound
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 843fc5fc-219d-41b6-8d90-e2ff72eedd0a · inbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation caaa20f3-36fb-4418-b8d3-d12c69e68d78 · inbound
Phi-4-reasoning Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 6479e86a-1f1e-4d7a-b287-167f95917dc3 · inbound
DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation dd04b8e8-a459-453c-b41b-44d68b16a649 · inbound
Always Tell Me The Odds: Fine-grained Conditional Probability Estimation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 0322336e-a359-49ae-b64a-48297f7bbb6f · inbound
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation dacaf138-2ff8-4cdf-96be-7db464020e06 · inbound
SocialLM: Social Signal Processing of Patient-Provider Communication using LLMs and Contextual Aggregation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation ce75d320-2c9b-4c7a-8e03-01b105ceaa31 · inbound
Flow-GRPO: Training Flow Matching Models via Online RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 2488832d-2931-470f-9507-c33112e66903 · inbound
LLMs Get Lost In Multi-Turn Conversation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation a7dd1a63-4a2e-42ab-bd40-8e5d44fd080b · inbound
A Survey on Foundation Models for Personalized Federated Intelligence DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 9ffc2f68-d28e-4814-b09c-179010b8042f · inbound
Seed1.5-VL Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 5421bf75-b042-4d23-904e-9e889de9bd6d · inbound
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 11809140-14c3-4f97-ae15-67bf556b4edb · inbound
DanceGRPO: Unleashing GRPO on Visual Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 76b95b6f-5f57-413a-b0db-a92597925aed · inbound
Not that Groove: Zero-Shot Symbolic Music Editing DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation de132211-36e5-4928-943e-0d0bfd0276ff · inbound
Qwen3 Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 757208c1-62b3-4bbd-b38f-eb9bd58d792f · inbound
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation 73db4015-62f4-4f9e-8870-f6bc33d7eda8 · inbound
Superposition Yields Robust Neural Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.
Observation d3df8d7b-dcd4-451e-8682-aedd51add6ac · inbound
InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-28T06:31:03.373048+00:00.