Pith. sign in

Paper Citation Record · LEDGER

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

As of 4 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 100 inbound Pith citation observations for arXiv:2506.13585.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13585 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T09:28:16.189617Z

measured 150 of 150 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 100 of 140 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T21:20:34.201020Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T22:26:37.204571Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact36
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e68edf5f-de20-434f-bd6c-dfeda10c4553 · outbound

This paper cites Simple linear attention language models balance the recall-throughput tradeoff.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Simple linear attention language models balance the recall-throughput tradeoff

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.364285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:281f68119036ae93f2c8ccbbe67d9a2b66564d7776d1ef251bd98b6deb976d6c

Observation 5b3e7e35-31d4-4657-b486-35d5ac92efa7 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:38:00.425257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:981dc0bd1c23818a2400961a3bf063570225b85caa56d7226e32cf1260a2cbe2

Observation 9c05ef52-b361-4e23-b9fb-94875aad42da · outbound

This paper cites Titans: Learning to Memorize at Test Time.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Titans: Learning to Memorize at Test Time

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:15.778565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:2d0937984315ed5808572832249120c563902e04f4c074b451b3fc0a1636318c

Observation ec78ae99-6974-4acb-bfb4-de531ad346aa · outbound

This paper cites Longformer: The Long-Document Transformer.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Longformer: The Long-Document Transformer

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.279029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:15cf4a4cd9d9f35d93faa0cb84423ab76d0ed1ac3f4c1b0f56adb6ad7f774593

Observation 7b29f544-0aea-41a5-8357-dcd458e463e4 · outbound

This paper cites Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.286365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:c3a13dd28de974fc0f39d0c7bffb9bbbf7895607af99ab3fb03f77cc97526d8a

Observation 56f03810-65e0-453b-8ed5-1da4fd01f4a0 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:31:34.026635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:0547cdb59bac54cfb8564f19ace1edb74ab68b498279a72a436023622b2ee870

Observation 593fad5d-b88f-46bc-837a-d9b2a412b8a8 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.300988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:a507d6afc6e4a67c36970dbc9f7f71f93a015e7b0b0418838e179a49924ccd77

Observation 21d32be9-0bca-4a8b-a714-84827c5d8f7e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.307524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:71adfa97db9f559b71b21b4f7adc4fbda4d3f1fa81f4af202a49e544c1568174

Observation 60964a14-e308-481f-ae78-743bba18e333 · outbound

This paper cites Jacob Dunefsky, Philippe Chlenski, and Neel Nanda.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Jacob Dunefsky, Philippe Chlenski, and Neel Nanda

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:28:16.313654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:5641818380ea79603ac2a5a02a9bee80c7b8af20431a79f72201b8e92033a615

Observation 9654d861-6642-4b76-88b4-5557070304b0 · outbound

This paper cites Zamba: A Compact 7B SSM Hybrid Model.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Zamba: A Compact 7B SSM Hybrid Model

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.320560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:bb1262086fff74dd1341cd4fdd40174f24faeda8cd29b7c7183274cd9ec60d3a

Observation e3659577-ad0c-4230-a534-68a211f70d99 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Efficiently modeling long sequences with structured state spaces

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T09:28:16.562028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:59ae28d2f9385ad1f7372845f5baeb382acea6a4f894bb54ff17251c71c3f6df

Observation eafeb127-343a-4d17-87f4-02c7f3e9b5d8 · outbound

This paper cites Rodimus*: Breaking the Accuracy-Efficiency Trade-Off with Efficient Attentions.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Rodimus*: Breaking the Accuracy-Efficiency Trade-Off with Efficient Attentions

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.371534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:350046bd66fa4a40bf18f0a6e1034553c4dc863ab40322c1fa1a561efd8c87d1

Observation 37b134cb-284e-42a6-9895-08da0793bb76 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.378103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:cfb57d95a78e0fa812e61b7ea98baea45ac51a8b35ec2f6702529da649c7a585

Observation 0f551b5e-2242-45cc-98a4-232f5899fd46 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:59:02.400889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:375ef3d14e7f46b88f43090112e2b5ca68d6fb3bae00c5cd7f789f652fcf300e

Observation f80e8e31-7220-4a05-bfc8-859117216744 · outbound

This paper cites Jamba-1.5: Hybrid Transformer-Mamba Models at Scale.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Jamba-1.5: Hybrid Transformer-Mamba Models at Scale

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.390827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:fa0fe4ec6319c041ffff2dfbf2e130b71a4dc55ee9dc99a526cfd2b8e5aab2c8

Observation c56f9962-2490-41a2-8bd5-18ab8b5d60d3 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.396386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:7ff9109100763dbfbd58eff6a6fb734e7691684e5d2dba0e2cd0886a20412c50

Observation 16742d04-cdeb-40b6-a297-36ef8ac30758 · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.403060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:b849212e0ba2ac7490c55f1621a5ae01ad61669e79ec36795d3170d400328131

Observation 67260a79-9d62-4e4f-986a-a8294be5bfab · outbound

This paper cites SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.409868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:609a9674aa939bf6e9ac9e69f318002aa92f62812ff93b8c439faff58666deb2

Observation 331728be-3b51-4110-b04a-41cbf0c5ed92 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Understanding R1-Zero-Like Training: A Critical Perspective

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.416359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:6662985101f246ac3b0fea97e880b6c6154ea61ffdc1d87a0f87737f404eea67

Observation 91649ab1-c636-4a27-89a8-ad28dba2fafb · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:15:46.352230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:86e88e77fe9b2dfaccb979db3c65308ab6b5e35f8aecdf851f46225a1645b49b

Observation 0dc7b9d1-f66e-4e67-a50d-3c2d18760fbe · outbound

This paper cites Parallelizing linear recurrent neural nets over sequence length.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Parallelizing linear recurrent neural nets over sequence length

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T09:28:16.585483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:227fa94dade2d27b21b6a3339c8f528a7c5f997f8076fb95c36bc171f0bb0731

Observation d4ef5545-0d5b-4647-9c1f-e04512a18721 · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:26:38.921226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:167f59bc0c40fa808475824273ebc43e2764135afdb178c27c35646ce1820f71

Observation 65ff9138-aa07-45f5-819a-2f8d2b8f82d2 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention A Theory on Adam Instability in Large-Scale Machine Learning

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:28:16.436746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:44e89c4467488799018cc1cca1284e3a2636dfd1ad0fed78d1e1679a5489b586

Observation 16f2f937-4798-4a27-8407-540cfc123fe9 · outbound

This paper cites Openai mrcr dataset.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Openai mrcr dataset

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T09:28:16.575258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:1579b01680d357296a5252d70069e7c70fa416bf0b08c38c66217ef0ce38eefd

Observation da63c11c-e263-4b29-bf34-bb5d6546aba4 · outbound

This paper cites Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:19:58.366890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:055b168ab7a527d3c247ae5f731d856ac212195aa7628ea878df601547147649

Observation 1c9fa2bd-f427-4dbb-a00c-f6156be8024b · outbound

This paper cites Humanity's Last Exam.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Humanity's Last Exam

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.460030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T22:08:32.844211+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:7a4feef6f811267e373db6ea638f5e163aa62124df5bb6ff0ba9d333b4f2c697

Observation 0ca1a511-b229-49fe-8dcc-d6e0fd5a6318 · outbound

This paper cites The devil in linear transformer.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention The devil in linear transformer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T09:28:16.570862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:10f878e36a48ae7a8cd0c899b7ff592201253aa1916a46fd97dd0d9a44dbdac3

Observation 858af31e-a12c-47ee-8cee-705f48de85b2 · outbound

This paper cites You Only Scan Once: Efficient Multi-dimension Sequential Modeling with LightNet.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention You Only Scan Once: Efficient Multi-dimension Sequential Modeling with LightNet

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.467787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:a90baf76e693888790b079f48c2433b6e20104e95de2e0bf0486c3c49be3f693

Observation 2abd9990-72fc-4936-8a72-533e2a802a8a · outbound

This paper cites Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:28:16.474483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:612c2bc434e32fed82289b3aa67084200d5a2aa503ee7b2c3a1246f967b9354d

Observation 86f5c81e-a8b6-4b0a-97e8-c95199e0703b · outbound

This paper cites Proximal Policy Optimization Algorithms.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Proximal Policy Optimization Algorithms

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.479837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:f341b9f05e01c56eb7336724e54af33c7ebcd648691eaad4c0a934b153c273e4

Observation d3aa5dac-3974-4304-85ac-82d203f22792 · outbound

This paper cites Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Seed1.5-thinking: Advancing superb reasoning models with reinforce- ment learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.485752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:e3dd6a5c0a49ccd81a743d490b12311302352006266d41471013c54c5e268f61

Observation 57125760-c6af-406e-96dd-e4250e091bd0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T09:28:16.492056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:be150b08cba7f61c66b74a1140e44170b4a37a53c647ba4b91b2f65f5492e174

Observation 0ccd4f7e-c438-4461-ab6b-800a4b5c448e · outbound

This paper cites Scaling laws for linear complexity language models.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Scaling laws for linear complexity language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T09:28:16.567025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:50b1a437ff787d9a159079cce81d51a04324e4371c51ace260839220cdfc5358

Observation 6b8eed45-b87b-48cf-bab6-8411df851386 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention HybridFlow: A Flexible and Efficient RLHF Framework

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.497594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:44f26f9c19b406329c00c4f8e325a6270a788a890b55b4146f5bbba6fdb81b0a

Observation 5a94e21b-abd7-44fa-b043-1f857267facf · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:28:16.503960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:b0001a3feb20ee24f3cde7e68c91bbe133d9cc90a5efd8b50c4a9aeb887a20e4

Observation dc341404-69b9-44f8-917a-5b76f4040c2a · outbound

This paper cites Deltaproduct: Im- proving state-tracking in linear rnns via householder products.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Deltaproduct: Im- proving state-tracking in linear rnns via householder products

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.510648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:8dff06b38dfe38b6a67c4482374edf3b7c81beaa3cb02fac50fd163e0e22087c

Observation 0e6f16b8-a732-4ca0-90e5-adf7914b8582 · outbound

This paper cites MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.518006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:83089d2540bd48c95293e08778b8424a0764878ce8a47767bceb3b1137517cf2

Observation 3a54b0c2-3903-4112-8156-c39f53c51fa1 · outbound

This paper cites Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.524229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:4c7738520b3d048d4bedadfa592a6760a614b220b682b8c8e5be8d235c3cfa73

Observation 798487c2-6b20-4806-aba4-8317a974ad20 · outbound

This paper cites Learning to (Learn at Test Time): RNNs with Expressive Hidden States.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Learning to (Learn at Test Time): RNNs with Expressive Hidden States

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:20:12.625336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:47285d62221a7db8bff828487ab44b5873e6b8b857598105ea21d81815e75c67

Observation 624bdbb1-782b-4c71-b378-26dc17b3111c · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Retentive Network: A Successor to Transformer for Large Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.537212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:0638d95f6d8339d2c72e84e358e2d591eba499e3c6c53d96da520f54ad0b878a

Observation ba43a89e-9bb5-4d29-85e7-7f2927d72a1f · outbound

This paper cites Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T09:28:16.580407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:0c7f06b6cd39c553637852f7000f0417976e7659b95db88d5a428c225cc898ba

Observation cabe8746-15d0-422b-ab9f-301c54448768 · outbound

This paper cites MesaNet: Sequence Modeling by Locally Optimal Test-Time Training.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MesaNet: Sequence Modeling by Locally Optimal Test-Time Training

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-04T02:07:39.171975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:38a8bebdc0bc942a46d55270f684a3fd5507f4f307c29e5838610a553560bfc7

Observation e33cebfb-ea64-479b-89db-b92606590f77 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:12:09.004283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:a66d615b6fb84f8cf5bdb2846f2a5f7075995ca4913aac3f7d50d4ae50234117

Observation aa04c82a-eca2-449a-a320-c53ea51779a6 · outbound

This paper cites Measuring short-form factuality in large language models.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Measuring short-form factuality in large language models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.345693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:18497465cfcfe36d32bd53a6aa0e1914023170f653f7135171e09df8241530ca

Observation 99e26778-80cd-4cc1-a9fa-d38686e4cf07 · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Agentless: Demystifying LLM-based Software Engineering Agents

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.327284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:2de20d559bb191e1facd9f50cd31ed715a8e20ebc705da72ec5312c30a456010

Observation f0db653e-4769-4fec-9817-2f1f7d4535b5 · outbound

This paper cites TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:39:31.128039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:3134f8cda5d664017fdd21aa3e981ba2863e04162ed1d0754cfaeb93b64d8073

Observation 634b95bd-8e8e-4906-bf3a-5e771fea1811 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:15:14.310613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:4b3b4f9ea7f1a35d32d5ae26066414cceebc9bd02e7bcbf1b53d9195ec5146e1

Observation fac225b2-26b8-4575-bb79-0f177dd04143 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:28:16.349490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:9bdf825ee86d21500c4ea0e607ae1911fdc8e24cad55afb82cf41b6172cdad58

Observation 5fd4e1fa-3aa3-4669-be5d-55a1ad0c4583 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:46:30.294205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:23ab11a9031c17fd8212add9f442d37ebbb188189268acf2ab9baaf85f366800

Observation 71b77b6d-887f-484d-b7d6-defffc02baff · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:21:05.971096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:8f20c058983c4112f093badb4372a8739c0c52e4a565b3656abc910f80a9f18b

Pith citing papers

Observation 3e09aa0f-8239-440b-992d-2301121851ed · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:0585644ad51dd62fe7ea32517f9368755aa26fbd9fd165ca94bd3272f309766a

Observation 6b011508-66c4-4c06-809c-a7c39b4c800c · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 124

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T19:32:00.612132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:3daba90961e0ebc15f97296bf5d6a0fc5c164f29a2f63b5971534cccdee6b8ca

Observation c73d9f34-1448-4c09-92c8-74a51a692335 · inbound

Group Sequence Policy Optimization cites this paper.

Group Sequence Policy Optimization MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T19:22:54.194429Z digest=sha256:18143a12c35fed7eb84050dbc7b09ff678cc11e868a334319467993bed22ff3a

Observation 970fa807-a60c-49d6-a484-bfb9ec99dfb1 · inbound

The Ratchet Effect in Silico: How Interaction Drives Cumulative Intelligence in Large Language Models cites this paper.

The Ratchet Effect in Silico: How Interaction Drives Cumulative Intelligence in Large Language Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:41:59.944495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-19T02:38:27.510872Z digest=sha256:09ae5a1c8d04bc1c52f8b599e0720b75e69bb45f0bbf1f49e4c5b0824a5c258a

Observation 32620d74-df48-4bcd-a248-defddb64ab0e · inbound

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models cites this paper.

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T17:50:08.399160Z digest=sha256:bf179863860fb0e07d74d761f4ed5a294e15d2bab106630075179701e235fab7

Observation b4ae12b1-fc5f-4937-9669-dc814ac88bb9 · inbound

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling cites this paper.

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:34:23.963723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T22:33:09.674822Z digest=sha256:f43d329e208f6d6fa279a98eed4dba36a9bedf2508a360a37af3ab0c54d4200a

Observation 7d1b4d7a-7c71-4105-8913-838fc92e4dd8 · inbound

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning cites this paper.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.574144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:ff1dfd572b898c180114e89392130f33a5e3339017e16d6b0308f62defa90b24

Observation 2183f214-97a3-4e5a-adfd-749caa639e52 · inbound

OneRec-V2 Technical Report cites this paper.

OneRec-V2 Technical Report MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:56:49.260399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T22:56:49.216473Z digest=sha256:a77eb3cd16d71d8300d6c2dc13f7507cec8c6b28f1851a0e6f715458740a73c7

Observation 6efab08b-9e24-4260-9bcf-bd7ac153841d · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:25.007260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:6ae81a6240dee3c140e2cd598a2eac790e8f672ff78778bb0f11d14a4df73737

Observation c6916d2d-64c6-4933-839e-720cf355372d · inbound

When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL cites this paper.

When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T20:30:35.536429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T20:29:25.874620Z digest=sha256:4a99e7b486cd23fe6d1dc80c30276bc219089a23eee2749128c9e93df2368c2c

Observation 16e52cfc-216f-4d74-b354-52098de5ed16 · inbound

The Art of Scaling Reinforcement Learning Compute for LLMs cites this paper.

The Art of Scaling Reinforcement Learning Compute for LLMs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T16:29:14.006669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:6f1e2780ba6462ee58e7d9ac1bd3f988854aaa5a0d05f8838dc5540920ad67e2

Observation 25b384bf-a634-4caf-a3b9-6bdb3acea3b1 · inbound

SSPO: Subsentence-level Policy Optimization cites this paper.

SSPO: Subsentence-level Policy Optimization MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:20:34.286180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T01:19:45.247967Z digest=sha256:9c8075a5c8051928d9f0bd376642b9b26ee66bea589cba85a3cc607456594dfd

Observation ffcfdc02-7916-4590-8457-9ce5e23b79b5 · inbound

From Ranking to Reasoning: Explainable Web API Recommendation via Semantic Reasoning cites this paper.

From Ranking to Reasoning: Explainable Web API Recommendation via Semantic Reasoning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:55:30.778116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:7335cd82c40ef2bfc63f35a1483a0ad0f7f6709ec663d4db8af6277bde30494b

Observation 49ab84e1-f9a9-48b3-bb11-c090793019c2 · inbound

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone cites this paper.

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-03T21:20:34.201020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:20:34.201020Z digest=sha256:382e42a089bc12787560a7e050c227d56a2b926d41dab43d37d0f1bddcd817a5

Observation ea722621-c861-4a8c-a7a6-6e6c108e1baf · inbound

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression cites this paper.

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:00:27.111637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T17:59:23.826110Z digest=sha256:16f9fe31134f357706f599426c9ccbb166c79e19cd7c6ddd11fd1b26b4fa9717

Observation fc2a974b-b169-4d1a-9190-b1f39d27718a · inbound

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training cites this paper.

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T01:48:50.986514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T01:46:21.744857Z digest=sha256:f72a7f81844d22e541c0e85f358fedbda3c222cdf6e43943c70765e8c0b563ca

Observation 20298989-d242-4b22-9495-ce66dcc40aef · inbound

SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios cites this paper.

SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:28:24.460251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-16T20:24:40.939455Z digest=sha256:a496ded3d5f6a00bbc34ef5afd72124c8033e3d6a6e15c55ae96f0fb81eb97d9

Observation bdb29954-907c-48f1-a257-01f58658a29b · inbound

Toward Training Superintelligent Software Agents through Self-Play SWE-RL cites this paper.

Toward Training Superintelligent Software Agents through Self-Play SWE-RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-21T16:10:20.223753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T16:07:48.570995Z digest=sha256:c2a1c233af1b799cb8501d67e90c86a42afbf31cf7036c969ee401b8259245c0

Observation 58dbb434-80ab-4321-8210-2d25144aa89d · inbound

Toward Training Superintelligent Software Agents through Self-Play SWE-RL cites this paper.

Toward Training Superintelligent Software Agents through Self-Play SWE-RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T15:02:10.359980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:02:10.359980Z digest=sha256:0577bda5acc1f412ecc3965bba6cabb231e773bc1f25ce253dd33c1252e932af

Observation e4ab3d3b-027d-4553-895e-a50d9a704d3c · inbound

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning cites this paper.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.349112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.349112Z digest=sha256:bc07d792d5550be59844475f5eb0d52c2f6eb649dd793a8bac35ed24eb11cccd

Observation c4443504-4134-4d24-bede-585fccf716e6 · inbound

Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design cites this paper.

Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:44:11.492588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T13:42:03.896311Z digest=sha256:d799796c4b33bfaeb6789f3bd1fcc4b5c6cf36a1d3083c2b9c5b27c06997b1e1

Observation 2a83887d-1d26-4f19-b7d5-de748c4df8af · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.357667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.357667Z digest=sha256:7f444ac966e7cf5c2e930cc3df9a812b6b202f2910a9c809189c3afcff6b9634

Observation de312d3a-41bd-4ffc-95b6-86e7050505b2 · inbound

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System cites this paper.

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:37:28.908077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T06:33:41.860803Z digest=sha256:dc86482562eb6d64312da501f5867d29c0818e96545da9000a09a4dc839f79d6

Observation 2e920905-a857-441f-801a-235ad2dd578a · inbound

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm cites this paper.

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T11:14:47.993011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T11:11:51.440058Z digest=sha256:1024e605dcb59ae893fb31afe77c98a88ad3bf5132d24fcbf3a788e9dd3b4ad1

Observation 2b21eaa7-957c-4b05-a7db-64ea4f9633dc · inbound

Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective cites this paper.

Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:37:10.393546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T02:34:43.248634Z digest=sha256:2dc3d9b28f41d0c1b932978bee835b04b129ed16032c127560bc7a10050b997e

Observation d4238400-8aa8-40c3-aa10-5ab80151159a · inbound

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens cites this paper.

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:46:43.270813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T21:41:48.690125Z digest=sha256:9c35328efce00d54def553d5c7a526b49e8341d87808dd000746c2792705fbdf

Observation 02ec5989-ebb3-4bc6-927f-9c4a4c8737eb · inbound

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens cites this paper.

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T22:52:04.889489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:52:04.889489Z digest=sha256:0dc06aed05ab79ce4c3e9769355889b6c99136f7804317d61945b68c66bf7a49

Observation d229a8a5-acd8-4255-a874-362312e2bd32 · inbound

Soft Sequence Policy Optimization cites this paper.

Soft Sequence Policy Optimization MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T21:41:30.634991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:41:30.634991Z digest=sha256:f3872c94fac59ab34671763fe5a1708271dbce0b6b5dc0e398a9675d14c8cb77

Observation afd8f254-9ddd-4193-9ac6-e7ee3efeba62 · inbound

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking cites this paper.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.153620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.153620Z digest=sha256:0e611ef9b3d1030e0c1fa79503a5dd9de84dd68d4eb970ce15051bcdc3ab3936

Observation 404904fe-7036-4f2d-a49d-842460e67567 · inbound

Stabilizing Policy Optimization via Logits Convexity cites this paper.

Stabilizing Policy Optimization via Logits Convexity MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:07.595577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:07.595577Z digest=sha256:e1d171e75b86c0a85a81232012fad09ac6a504030be1279f9cfd43ef8c9eb347

Observation 22a1b6f1-47b9-48a6-912b-1279eefc91f6 · inbound

STRIDE: Post-Training LLMs to Reason and Refine Bio-Sequences via Edit Trajectories cites this paper.

STRIDE: Post-Training LLMs to Reason and Refine Bio-Sequences via Edit Trajectories MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T19:08:34.612684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:08:34.612684Z digest=sha256:06a4fb1d4ba430dcdfd4c6ba851353d665829a11429b43b41e0a8bbb09ebf5ea

Observation 826de674-94c9-4999-ba5f-478f8030ec94 · inbound

MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue cites this paper.

MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:26:10.750491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T15:25:51.138776Z digest=sha256:59b46d5d1520172db13480b1c658c80c40b4babcc169e1a9f2fc6d53d7009d70

Observation 5ee11886-eb84-478a-9375-506b66d0ab64 · inbound

MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue cites this paper.

MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T18:44:32.165143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:44:32.165143Z digest=sha256:262242ed7580b2a0781ff4ad426b56c2a50ed00533ae917b18551a7f81a2275e

Observation ab4f3d66-e961-4bf2-80bb-bcca4c1a1447 · inbound

Policy Improvement Reinforcement Learning cites this paper.

Policy Improvement Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:48:22.821800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T22:47:05.132020Z digest=sha256:b26762f262fe66face4a1154cca2340cc3048dabd1c0c348392ff028a0620f98

Observation a8e293ed-5ea8-4fa0-a606-56021edcbb8e · inbound

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments cites this paper.

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T19:07:46.077831Z digest=sha256:d39303e3fad34978b820fc335dd9d5b90508b670889968d6100802f95e5a7413

Observation 47233751-3fc5-4d2f-86e9-1e806c38c06d · inbound

SAGE: A Service Agent Graph-guided Evaluation Benchmark cites this paper.

SAGE: A Service Agent Graph-guided Evaluation Benchmark MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:41:23.956104Z digest=sha256:891b704e633e9012e21655a7f30063ab05681b48e495227ae99690d60a023e32

Observation 4673c71a-edd0-4790-863a-1de397d6c26b · inbound

MEMENTO: Teaching LLMs to Manage Their Own Context cites this paper.

MEMENTO: Teaching LLMs to Manage Their Own Context MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T17:16:36.784237Z digest=sha256:54ec7880dd63b31b7acc86e84ff1477c9f7e86fcb1882bd9aa2dcc4751c7b4c6

Observation 1325ab61-421d-41af-b60d-d1e13849a183 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:12:06.985609Z digest=sha256:a09e64aefd3eb4d03496f74df6be091623136a26adba55305b9d4dffa5d54465

Observation 4761a986-7f68-48d8-9da8-d9427ce7edb6 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T01:04:12.454268Z digest=sha256:ff7d30d75dc7db64ed329b2b2179bd2b214586223dceb5e729c4797e27e680b4

Observation e4225109-fd0a-49b0-932a-51456e662cbc · inbound

Beyond Distribution Sharpening: The Importance of Task Rewards cites this paper.

Beyond Distribution Sharpening: The Importance of Task Rewards MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T08:07:14.691463Z digest=sha256:e400f1251db1bca1d63cad8d4473ac9e8abffa9ef09f227eca9bf9497e0a4bda

Observation 592c9e31-cc86-43bb-879a-b8e8f9e117a8 · inbound

Scaling Self-Play with Self-Guidance cites this paper.

Scaling Self-Play with Self-Guidance MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T01:31:06.090698Z digest=sha256:90547213f3f148c63f14bfd285421eb08caf0a9e63e02dafed04d7ce00ed3692

Observation d81cc469-82bb-4da0-ba32-304f0a53c70f · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:6759b3ec81da4890e5e593193ff2fe145f4098663e8bfbe8fa0033b702fc37cf

Observation c3306625-b3bb-4f70-aa84-832274d4fed5 · inbound

Cost-Aware Learning cites this paper.

Cost-Aware Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:36:30.931073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:11:01.131590Z digest=sha256:11ad3962642fc019c5183930d4c779971a6f6695f38ba66c8dad9aceaa022cef

Observation fd34b05a-6c76-4482-87b8-fc79236e14c5 · inbound

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length cites this paper.

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T18:13:25.735085Z digest=sha256:5a6ecb81993a654146968922bd1e7e4b588e42242f00d3a57162cdbfc64ad76b

Observation 13159349-17ac-41ca-ac49-0bd6baf553bb · inbound

When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs cites this paper.

When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T01:27:54.164991Z digest=sha256:0c8f330a9804c1fb5ce048d0fcfa89cb042b48f66561512ae4cffb5a51d8cbda

Observation 8f0e9aa9-b5d1-471d-aceb-7d0cedfb200d · inbound

When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs cites this paper.

When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T16:57:26.334739Z digest=sha256:c677de0a18c091567ef2a2a65b2c43848ae2da9599cef417109ae2f1045310f6

Observation b3e036a8-8ebb-43b7-bced-178e70046601 · inbound

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO cites this paper.

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T14:53:35.133157Z digest=sha256:2b7a551dab78c6dde3f5df869316cf714d385e0238dbe088e748aa32c12b37c4

Observation 9d543917-dfea-41ce-bc7e-b3bc3c1adb73 · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 164

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:c4bbf4eb525804ac2cf78d1a4cd99515a036cb5417ea426a98ea6b69539630d6

Observation 59b9e726-dfb3-404f-bc6a-cad1101ea784 · inbound

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR cites this paper.

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T15:39:54.001327Z digest=sha256:951d880c287c0054c78a7cd2b0c63b6dfbb91e0fe1d97ed2342b8b8dfa398c83

Observation e18527fe-41b6-422f-9a24-faf07d29ef46 · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T13:55:22.422923Z digest=sha256:77333801421c6bf56725deef4f2d6cb5b36f3c9d1588273625926ccc5eb905c5

Observation 43fff8e8-7ccc-47bc-9c6c-7260f7cd1fa7 · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T09:04:04.494942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-21T09:02:28.407064Z digest=sha256:c06cae29c9fc257095c6031a25b3a824f4d5e8759f02f59cc67671fe132fbae5

Observation 016ba4b9-6d44-450f-bac6-581e01c96a98 · inbound

On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR cites this paper.

On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:40:53.063991Z digest=sha256:f938b543efa09eaf9456727bf64dadca871a5d3c462abbc5dc267de3a1acba61

Observation df9dd6c4-b80e-44f1-94e4-e6a3681bcc5f · inbound

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction cites this paper.

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T09:59:48.604813Z digest=sha256:95281d9054bb91d07e612c76719f10bd16cde20c8b717f1f66b8ab8bf62b0686

Observation d87f61ed-4dd0-4d57-a869-e9d4312d91d4 · inbound

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients cites this paper.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:311d32bb5fb6294ae58aafaf6d4d5a3ad3f5d7d1dd2b071581f76f2f6f9eedd1

Observation ad615347-e005-4078-8bd6-bdcabfffad1e · inbound

KL for a KL: On-Policy Distillation with Control Variate Baseline cites this paper.

KL for a KL: On-Policy Distillation with Control Variate Baseline MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T03:31:07.462474Z digest=sha256:16df4d79a6f73c2ce0bd15a22a9031e17d196750baabf9aa8a65356ac0a00345

Observation 73390f12-bb59-4489-a813-743d78cc460c · inbound

Priming: Hybrid State Space Models From Pre-trained Transformers cites this paper.

Priming: Hybrid State Space Models From Pre-trained Transformers MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:70f0c1b626e336b9a443afd1e9a83058433db57465b06796da70d60851a555dd

Observation 528b2895-8b71-464c-92a7-c2eac2d97257 · inbound

CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging cites this paper.

CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T01:43:06.644640Z digest=sha256:5805c12b5df889860e1a35ddc1f6ca81b8e8fcbbe8ef0372aac22c41ba6cb17e

Observation e7494e4d-7876-4197-ba33-222885f30225 · inbound

CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging cites this paper.

CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:35:46.943118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T22:54:20.299335Z digest=sha256:1c7a53563a4d439be5b6531514c7d984cc568cfa88d9f59e4e7c7d56316ce009

Observation b9beb139-041a-48c5-8ca1-e0ea276e2ab6 · inbound

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning cites this paper.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:0e41e118e6863a81a8e61fa72fdb7b974722debd72e22657fc0f7288b4569a5c

Observation f7d85ecf-0f21-4be1-bb20-4448c4dd40e9 · inbound

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization cites this paper.

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-12T00:59:44.364491Z digest=sha256:e8aa023da82693d7b8822c10c7f5ec9304c8bc5f39b2adc0f189de89691efc77

Observation 81dff8b9-4bfc-4a25-aa92-56ef2365ab0b · inbound

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping cites this paper.

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:33:40.994346Z digest=sha256:b7f00ee5d8fbffda285f745fe65d598d2b3e848878ce427c970cc01e64aaecb2

Observation a3e68dec-3125-4f92-92e4-0e351725efc6 · inbound

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning cites this paper.

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:22:06.767513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:19:27.345348Z digest=sha256:2ddc6773c1f5c8ad07a2dd781edd62081b2564c46d2c8b4bc05b1ef12295f5d6

Observation 243d28b7-4cab-4dcb-aa7b-30b7bef80327 · inbound

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction cites this paper.

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.573116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:57:49.286939Z digest=sha256:7d886069d2674b8fa3aef9f84707488d963912cd244dbb4228093dd2e5d75bdf

Observation f4e1da17-d336-4264-ab27-989813359986 · inbound

Learning, Fast and Slow: Towards LLMs That Adapt Continually cites this paper.

Learning, Fast and Slow: Towards LLMs That Adapt Continually MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:18.451790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:00:31.452781Z digest=sha256:fb81c9cd8d9acc76441792a48e9921624dbab032374cb4e80bec7e28c54c1adf

Observation 535f1e6b-0f6e-4ca1-be81-d4523d591ec8 · inbound

Learning, Fast and Slow: Towards LLMs That Adapt Continually cites this paper.

Learning, Fast and Slow: Towards LLMs That Adapt Continually MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:19:45.756862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T05:19:05.368681Z digest=sha256:8e2616000a1aa5817f97a480ce1e5aa916436e05478c1363a1b27529547ce4f3

Observation 973ec6af-f331-4fc5-8496-d04104a314fa · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:22:54.727347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T20:21:26.640493Z digest=sha256:a57502194a02b1fc542ac54681e3dac0b4cf48e037e56e41f57b1741e8c9c5f2

Observation 43feeb7a-5639-4b33-af6b-461beb8e7121 · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:43:45.447227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T21:42:49.452347Z digest=sha256:499b8e30cea125cc67ce1c052a276026a2aec031209c1486c9ce189f747af664

Observation 4ffc632f-bad1-4fe8-b374-29c830ddcb33 · inbound

F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking cites this paper.

F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:52:52.301459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-14T19:51:45.657658Z digest=sha256:89f99c2dfc872444af1b893cf0a8fd30b022a3920f2cea09f158fe1aea97a618

Observation 219c7ddb-d5e2-42b4-976c-0e857e6f353b · inbound

MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning cites this paper.

MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:49:25.833263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T19:48:53.213389Z digest=sha256:ab5dea28ae646dfe1fe815634d73f21cd1f67c47062b7946968e7c9ed1e134ca

Observation 1073cbd1-ab0b-49d0-ba04-7735432bf065 · inbound

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models cites this paper.

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:47:53.443277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T19:45:26.458859Z digest=sha256:b7e89689e78feb345893dbf33d126736126da317cdf780c01f5eb94cdc0f4b4a

Observation 6262dd51-39d7-44ee-85ed-93f7170c6f54 · inbound

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models cites this paper.

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:05:06.539198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T06:04:17.574622Z digest=sha256:7354cb384be4a918d79533a094d89912eb2fd83cddf560a739ea2234cd8f483c

Observation e540ea9f-088f-4a8f-85db-d5f7c9dc722c · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:19:43.083892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:9ffd67e1ae766ada9be022736a935985bd8d624997133d9792a7dfc9d139f3dd

Observation d46f416a-957b-4119-96f0-016ef6d78a3b · inbound

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning cites this paper.

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T18:02:42.410827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-19T17:58:05.817581Z digest=sha256:849bbbcc171b36611f599186f53f7b1eaf805e5d30d4851de68ecf0b6de6aebe

Observation dab4ccb6-f27a-449f-87aa-0db2c25d18af · inbound

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation cites this paper.

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.036772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T19:37:09.244578Z digest=sha256:d14a14e37a1522efd493cbae916ad41325e8cba3b4483a2b00248711ebf17d6b

Observation b8f3cae1-2e6f-429a-bcd8-e626ad505f2e · inbound

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making cites this paper.

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 91

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:53:23.362870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-20T14:48:55.203993Z digest=sha256:f0489e36b5065260d6810d48e7f2502581ba9c4d1fa996a7b4ac497ecc66732d

Observation 28a354d0-1b5e-4951-b3ac-d8ea05325266 · inbound

How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning cites this paper.

How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:43:19.695402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T13:38:45.835754Z digest=sha256:01007ff6707ee3622452c315cdeb56780acd3ed33c0dee9e299a8beac2bce363

Observation 972dc449-351d-45e5-b653-90d1a3d809b5 · inbound

BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting cites this paper.

BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T11:48:14.582596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T11:47:37.503669Z digest=sha256:8ed8a2f34bab078dd7b08f7c5636689732340e0722d1df9b36182aa9d1d5c764

Observation a461d5d7-ff41-4cf7-9c3a-72fc634dacd6 · inbound

BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting cites this paper.

BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T19:04:59.947587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:59:23.215826Z digest=sha256:40055bc23a571c771940b869fd191d30641db4239ec1f8f7a462c61515242d26

Observation 56b928d6-8271-49f6-a332-83d15bb20553 · inbound

KVBuffer: IO-aware Serving for Linear Attention cites this paper.

KVBuffer: IO-aware Serving for Linear Attention MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:03:15.274678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T12:00:06.171816Z digest=sha256:ec3b8307e64f8c5415827714de5b4c02ae06b203f1d29142cea36a778a4b2160

Observation 219827f1-287f-4231-8e00-68d4379b366b · inbound

When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR cites this paper.

When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.420080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T07:31:23.194325Z digest=sha256:e4a1a066d41a9b7bb654dcfaa05bd2703ed2c3eb02699c8885564091475dc9b9

Observation b9511ce5-506c-4d82-ae68-815637aac990 · inbound

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards cites this paper.

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:59:41.109714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T05:55:45.654673Z digest=sha256:a0ce64b761340106defd375f075b3786476a8f0da804452fa0b3ffbdf812a707

Observation 44906f4b-b4bc-4132-b2a4-7acccd1d3ac0 · inbound

One-Way Policy Optimization for Self-Evolving LLMs cites this paper.

One-Way Policy Optimization for Self-Evolving LLMs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:01:15.588609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T08:01:05.911650Z digest=sha256:160d45db49a43c6eab0c4b3346d9811ba5ff045e469a475caac1c4dc656d23d9

Observation 5e0859d6-e3d1-4d7d-a264-7c97c2c886c0 · inbound

Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals cites this paper.

Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:51:15.796005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T07:50:44.907952Z digest=sha256:70d8c8c6a198074ded851d621d4c46e9c8ad2d2e86fcd6e8e36cad615a0739a5

Observation ec7be72c-99cc-44f9-9651-00d30ef08ec7 · inbound

Extreme Region Policy Distillation cites this paper.

Extreme Region Policy Distillation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:24:02.212528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T23:14:56.223606Z digest=sha256:7e1b82adb37b94e16faa2960390e24c1c2c272fa2ad3292eb5b853a8466f8ee3

Observation bde85b5b-d6b2-4f13-9184-fd7844d8d025 · inbound

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL cites this paper.

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:43:30.588305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-29T14:41:13.191919Z digest=sha256:12cbfcc5226946a3beb6162203dcd47cbbffaf247e9676e69cdc5a81d173df6c

Observation c0ee9ffe-4d9e-4825-88e0-d45f00133734 · inbound

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning cites this paper.

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:23:24.140080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T12:22:39.655615Z digest=sha256:09c3d0f34ebf2b7c1c6978591d8be4ec359b83082844d00d86830a0b71f18f35

Observation 29c8779f-fdda-409e-80f1-508e656c1a1f · inbound

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning cites this paper.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:43:14.669708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:16f7bc719fe41090b7d4e7be30a9824692ed9d61c85228d5ae91664fb4ba1ca1

Observation ef2d275e-88f3-46ad-b6ba-6e2b3a692a95 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 220

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.686591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:cb57ef0f3431dbe042689baf64a93a22d5a5f1c1f65b7e655ed5e8338af569d3

Observation 2dbb7668-e0ee-4cf6-b1c9-42c231482f3b · inbound

Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses cites this paper.

Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:26:22.117948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T14:25:08.052988Z digest=sha256:7d375784cc5d8bb4236a65f7a3694c8a412ecd0aa58b604fd521e23e321734b3

Observation 3102dbf0-819d-480e-bf9a-ed70f4e3c6e5 · inbound

Building Better Activation Oracles cites this paper.

Building Better Activation Oracles MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T14:24:45.029320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-30T14:19:37.262201Z digest=sha256:c9fdc171abe8c9ef512d0008b5f0929566558b1af0aade8c9625fbe82e91931b

Observation 1a887651-8d7f-4103-9808-8dffc4b75f73 · inbound

CodegenBench: Can LLMs Write Efficient Code Across Architectures? cites this paper.

CodegenBench: Can LLMs Write Efficient Code Across Architectures? MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:06:23.733023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T13:38:48.308683Z digest=sha256:e28838737d987b7bee22bacc9ef09bd430aac439a756add74d9bf6fb20527972

Observation fe087c66-1ca6-4dc4-825d-69e0e5a4391e · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:06:41.085294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:b581f228b988450448058e45f2eada8ee475f96903d2020ddc987fc8db9c8623

Observation 1d23178c-503c-493e-9650-a536e97ea135 · inbound

AsyncWebRL: Efficient Multi-Step RL for Visual Web Agents cites this paper.

AsyncWebRL: Efficient Multi-Step RL for Visual Web Agents MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T11:56:55.541482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T02:45:25.694378Z digest=sha256:e0eb7af16fd9df75c84180dcc1f944214424189615ff8c09b563e52008569323

Observation 8ae9c60e-72cb-40cd-b58a-313cf0ae28c2 · inbound

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short cites this paper.

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:17:29.552493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T17:18:03.459289Z digest=sha256:10883941b79c48cd862786bc68589c60d63b000a2440464636574b0001bb4248

Observation 79b6cc93-3e73-4351-b30f-205024855454 · inbound

Rethinking the Divergence Regularization in LLM RL cites this paper.

Rethinking the Divergence Regularization in LLM RL MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:27:29.403496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T17:14:22.229521Z digest=sha256:d31b28da6ec0165aff7c9f3810bfcb75bde8ed60f0a84e0970b3616c144accd8

Observation 78afcae2-efc6-422d-97f7-29e17030d033 · inbound

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff cites this paper.

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 100

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:27:26.422018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-27T18:49:47.876179Z digest=sha256:2390cdc23c5a52a1256c4543bf1a0cdae27d93c972438995f5e55e6a122f94cd

Observation f3be0ae9-5ee3-49b2-a6ba-95ec85a59737 · inbound

Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs cites this paper.

Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 93

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T22:52:45.490001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T22:49:03.124658Z digest=sha256:d8679f572f9192141fb4953711b369de1b28b6f0f7ff32c0c5dee03097d08c18

Observation 01ac48bd-6ae6-43fa-a5de-407bc10e5e52 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T09:47:59.866198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:4f59f89236882aa2a56619676458b82b1145f4f6674e1391156298da91961b28

Observation b6b38a68-9110-4cd3-bba9-22a29141e4f2 · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 288

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T09:07:47.893630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:6ceb5d62e5bbcaa2ecd962f2888255941669508b9d6140cc49da5563bd00a76c

Observation fffc33e7-5c0c-41f3-9f00-c91e0f7b3f80 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T09:37:49.398551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:e0cc23a88cf3cd215ed7b2391d60ed78748420d16713d049e8a1cc1b2ce7bd4c