Pith. sign in

Paper Citation Record · LEDGER

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs

As of 6 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2506.12876.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12876 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T09:01:16.991413Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:33:45.192139Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T07:16:28.884137Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact23
  • verified fuzzy5
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b7a8745-443f-4d20-84d3-979a884d812c · outbound

This paper cites GPT-4 Technical Report.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:02:14.533711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:5fc3e8ba66d1e874e270e9207463bc2e7339360d7edcbe2955f5adb95f39d31b

Observation 12a7463b-ae81-4132-86b9-7725ce243f1c · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.510181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:bd7b493b3ac15445a3ba19aca27c70d5b3d94683bef4aac6ec9bd12082e796a3

Observation 2b846821-4029-47fd-8e49-53ca091ada16 · outbound

This paper cites Conditional Gradient Methods.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Conditional Gradient Methods

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:02:14.502332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:96149153b3acdec42f69d2aa4b818e28254819ed98b0ab1bcc1503ac2786d0a4

Observation b8c0716b-0594-427a-b5e8-3dd8ef99673c · outbound

This paper cites Language models are few-shot learners.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Language models are few-shot learners

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T09:02:15.329803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:f3d5f157c192c8a11e87e22e14cf66dd8ed879616e49de76c669a1c3561f2c2d

Observation e9d74994-b7e5-449e-a920-06b4435db414 · outbound

This paper cites LoRAShear: Efficient Large Language Model Structured Pruning and Knowledge Recovery.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs LoRAShear: Efficient Large Language Model Structured Pruning and Knowledge Recovery

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.482082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:903fab5ac2c9646d280ccb64c4c26816176d95f364b1e7b30e8f85bec4d32fd9

Observation bdcbcc22-64a4-44ed-b14d-8d77749cc551 · outbound

This paper cites Task-Specific Expert Pruning for Sparse Mixture-of-Experts.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Task-Specific Expert Pruning for Sparse Mixture-of-Experts

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:02:14.475061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:e3a17ee58c5e1c29d6a18a18955b94b9165cd56b15927683ecbc031353a14e98

Observation 6635c67d-0394-4f16-8376-a3484e1ddbd7 · outbound

This paper cites Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:02:14.488887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:9dbee5652e6335ff74078c155d61906cdcd1a6ad2c04e21a0877dd54b0c4999a

Observation 810e8927-0e8b-4329-8bb2-14774f4a8dd6 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:02:14.516002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:f07032ccd8c358dc23a0760ae0f3dc8cfc3208ec16e0e3fc8e2620f3e10779df

Observation b760cdbf-06ce-4c04-bc65-bf92194acf81 · outbound

This paper cites Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:02:14.539340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:009c77bb710880c7a25d6aaba78062b1e7cd02f3cd50c1937da3529b35972527

Observation 25088480-6b77-4779-a666-6d74713839d7 · outbound

This paper cites Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.450120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:c18fc8e954196236680ff3116726594943b35912b60580ff720342d880e16873

Observation a4a77b2c-73d3-4780-8bc1-da80d7f3da57 · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:02:14.468374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:31fa502addd63397e0cbd9ac7205e7a4b7709f6666b953a18745e1cd71324419

Observation 169e84d6-5917-43a8-afc1-1daf4e1d77a4 · outbound

This paper cites SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.432385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:5b05301d8b2b85f26e2e705ff287f4a27f9d8fb8c90f1964dafa16389b8aeccd

Observation ae558d44-2087-42b7-a42f-9186c26ae340 · outbound

This paper cites He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:02:14.438600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:03a2d2cd37c99430c41b6c502f1267e291c1453370a6cdcc663397b791af91fd

Observation 0c45894e-7e30-48af-ae59-49fa6f643542 · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:02:14.418983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:6268b10cd0d6fa5847aa4ee76e4eb0b7c833b68b266f74492cb5d968f8fa20b5

Observation bcdd5286-1cf7-481f-8d0a-493237f2079f · outbound

This paper cites Measuring Massive Multitask Language Understanding.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Measuring Massive Multitask Language Understanding

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:02:14.425268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:a574bd7d6096d0de774645f337f35010bad85c5bd4c74c39a15ef6032a232cb1

Observation ea9901ba-def2-41da-b32b-3e175a3b8147 · outbound

This paper cites NASH: A Simple Unified Framework of Structured Pruning for Accelerating Encoder-Decoder Language Models.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs NASH: A Simple Unified Framework of Structured Pruning for Accelerating Encoder-Decoder Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.444063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:5113f30f09f09350976669601d49280d08356904b14d7343a835b5f51b189518

Observation 107f5e7c-ba76-4f8c-951a-31ef40885737 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.462154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:f969a8e45c0cfb6d2d5df75d39cd145cd8d24e9ad833d676c70e216327259385

Observation 7c1ba13a-6e5b-43b2-be41-5d9fde396e16 · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.495620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:6f04711b3c00d29b86e55f2b8867dc9a9861fcd17ccb14a99751773b1abeb54a

Observation 7e2c1e45-5b1a-4a8b-84bd-163db5c66d47 · outbound

This paper cites Accelerating Sparse Deep Neural Networks.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Accelerating Sparse Deep Neural Networks

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:02:14.394137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:09483e56a88d5b71d6d5f83cfdf5c83de99d4887907fa9102abacce9f751b584

Observation 1597a848-3014-4030-b0a8-c9a05d339192 · outbound

This paper cites On Efficient Training of Large-Scale Deep Learning Models: A Literature Review.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs On Efficient Training of Large-Scale Deep Learning Models: A Literature Review

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:02:14.400163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:d5c44f6823bbc363a1e8610b7bde402e23ce4db51b89574f21c0924f1dc13a9d

Observation cc848090-bc8a-497f-a8a7-fcd089bc488c · outbound

This paper cites LLM Pruning and Distillation in Practice: The Minitron Approach.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs LLM Pruning and Distillation in Practice: The Minitron Approach

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.406550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:404e2ad8bb79d89e8eccd2f0af412244bd04e9c842dd1af5d77c6a624f38c86c

Observation 50fb2b9b-0ebf-4bc4-b112-afebe82b8c26 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs A Simple and Effective Pruning Approach for Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:02:14.413055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:526da2c55f34db791c12ee98ebaa9b9e8e102b43fa13192e861f4035d4fc5218

Observation b0bc6a4f-15bd-453c-b8be-7c725b3b8c12 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Gemma: Open Models Based on Gemini Research and Technology

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:02:14.455939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:eb15d93de8f4d2fc521a040299dccfdae1ff754945e5c49c5a68e116ba3f81cf

Observation 6606bfda-ff8d-4616-b902-ea66d1954656 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:02:14.527739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:0457118d5af21d1dac13e8bb4a669ca2d4814a1ca7f33b3c11635e8df7076214

Observation a4826869-d6d3-49e4-adbe-945c7e9bd8b8 · outbound

This paper cites Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.552122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:df2f09be0007929daad383332900c35fced9449bfd4fa24bd139ebd10ed99811

Observation 1bec7733-1361-4ff9-ad4c-546b2a85b73f · outbound

This paper cites Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.387859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:84e8e8d1d33aa420592392ed8fb8b6ebf5deeec1693d50ccccb634351b3cf52d

Observation e5d46116-a700-4821-9c58-317345c79fdf · outbound

This paper cites LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.375780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:bbe2626b938ba8063823a3df993758a30b6e27f587651d6f22308140e6e4295a

Observation 1a547d06-e3b0-44bc-bbb5-8e2bdb0af8b5 · outbound

This paper cites Pruning as a Domain-specific LLM Extractor.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Pruning as a Domain-specific LLM Extractor

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.381972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:6408a82c9e0043191ea23585318707af2ccfb84b46b55128d09699e1b06ab986

Observation aee10fc1-99f2-4a51-8972-51ec7728b6b9 · outbound

This paper cites APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.522226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:d69968261c7a74a43717d0ae54d1c7d0f1e09b835336beebf645cae6701f9352

Observation 8b521e1b-059c-4494-a251-aecb8bbc2a7a · outbound

This paper cites Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Learning N:M Fine-grained Structured Sparse Neural Networks From Scratch

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:02:14.545469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:47c852dd6f81369e0cef1b9550114c09d8e024206a2f15751eecbe2ea1a5ba56

Observation 4471b47c-5dea-412a-a350-02f619e59392 · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs A Survey on Efficient Inference for Large Language Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:02:14.368980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:4dcc7bf6dc005c38bc356ab6bcd033fb055f5fa8fd1fedc6822776c2f9cd1802

Observation c24e6127-b66a-4e3f-9b09-c487d76fd08c · outbound

This paper cites 0 2000 4000 6000 8000 10000 Iteration 0.5 0.4 0.3 0.2 0.1 0.0 f(mt w, ) f(m0 w, ) Train with 1 Sample (a) Training set =.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs 0 2000 4000 6000 8000 10000 Iteration 0.5 0.4 0.3 0.2 0.1 0.0 f(mt w, ) f(m0 w, ) Train with 1 Sample (a) Training set =

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T09:02:15.318129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:6f3e4670b98fd0a1948ef419f96d4ed6951204cb238e75af8c295593f10f6ee9

Observation 093589cb-b717-4d86-8b24-76a8b00483c2 · outbound

This paper cites an unresolved cited work.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-19T09:02:15.314085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:b5925d8347915bcd142484354e1e99482aa125ff7d290a1a4916a50501c6c33b

Observation 0ddf2fd8-5a0b-4ebd-818a-dd91be01bfaa · outbound

This paper cites an unresolved cited work.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-19T09:02:15.321817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:ec72224d9344665c53d96614f5a5b670aa7f51256d35215c7b7b835f9aa259d0

Observation 3a484256-1522-40cd-9b2c-e38eb5df94f0 · outbound

This paper cites Figure 5: Loss residual curves of training on LLaMA-2-7B model with 1, 32, 128, and 320k samples.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Figure 5: Loss residual curves of training on LLaMA-2-7B model with 1, 32, 128, and 320k samples

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T09:02:15.325633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:267056632b503a49baa1d4e7a2f897b10f73717b6096093678569a2927254983

Observation d00f1756-82b6-49a0-8d0d-6e89cf87ac14 · outbound

This paper cites The MaskLLM method suffers from severe memory explosion and exceeds the memory limitation of 8× A100 GPUs (> 640 G).

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs The MaskLLM method suffers from severe memory explosion and exceeds the memory limitation of 8× A100 GPUs (> 640 G)

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T09:02:15.305044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:4b002904851cefd50fc5e10fb51fe89f2e919e0d9eb52e5831deba19e980e67a

Observation 297d6eb4-d07e-4d33-84f3-1242c6d30cfe · outbound

This paper cites Similarly, letting δ = 0, it degrades to the update with only loss residual, which is also a unbiased estimator of the standard policy gradient.

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs Similarly, letting δ = 0, it degrades to the update with only loss residual, which is also a unbiased estimator of the standard policy gradient

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T09:02:15.309160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:01:16.991413Z digest=sha256:0270b05a3435343fa92d904c98584d9d5de35ed90b9dee9de3aef3fb1af54f54

Pith citing papers

Observation 1018947d-9799-41ad-96ba-0ab476a704d6 · inbound

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization cites this paper.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:54:43.618312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:66114e0389350f210434b0de9e846c7f4d075686ccccd9c93baab5e41b6e1a08