Pith. sign in

Paper Citation Record · LEDGER

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs

As of 4 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2607.04371.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04371 v2

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T19:41:48.225019Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 481f871b-300d-4b8d-b7dd-58aa949534f8 · outbound

This paper cites SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding.

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T19:41:48.225019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:41:48.225019Z digest=sha256:da4d757acdba07196309c2b1cf9502c362f98706bf536930c90040d9dbcfda61

Observation e9445dc3-3ed0-4734-af37-6118be652910 · outbound

This paper cites Venmugil Elango, Nidhi Bhatia, Roger Waleffe, Rasoul Shafipour, Tomer Asida, Abhinav Khattar, Nave Assaf, Maximilian Golub, Joey Guman, Tiyasa Mitra, et al.

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs Venmugil Elango, Nidhi Bhatia, Roger Waleffe, Rasoul Shafipour, Tomer Asida, Abhinav Khattar, Nave Assaf, Maximilian Golub, Joey Guman, Tiyasa Mitra, et al

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T19:41:48.225019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:41:48.225019Z digest=sha256:c4413739f4c519241adcf24084898e195f7d1b8d64ec8426a577439f3753684c

Observation 5c1feb8a-e9cf-41d4-88fa-e0d554fcd4a7 · outbound

This paper cites Yaniv Leviathan, Matan Kalman, and Yossi Matias.

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs Yaniv Leviathan, Matan Kalman, and Yossi Matias

Reference 3

Resolution
verified exact
doi, observed 2026-07-11T19:48:13.157049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T19:41:48.225019Z digest=sha256:816fa89fc39db606a2c716511f1d6f6094e0a5209626a23f07b765195693f82b

Observation b4c6027c-a690-466f-8a28-fdc5b9ba75e1 · outbound

This paper cites Compact Language Models via Pruning and Knowledge Distillation.

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs Compact Language Models via Pruning and Knowledge Distillation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T19:41:48.225019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:41:48.225019Z digest=sha256:61276097c2e90e2bd927204f3ab174c729f95aa36c552ff6d09b33a3f0444a75

Observation 34c2917f-6c0a-4939-a47e-53946387a420 · outbound

This paper cites org/CorpusID:271328221.

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs org/CorpusID:271328221

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T19:41:48.225019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:41:48.225019Z digest=sha256:56a116b6da21d5bd791881b6b4fedecef6f4e02b7cf622240132485b0a5ad331

Observation 05212816-e4fc-40cd-819c-7773a08d16cd · outbound

This paper cites URL https://doi.org/10.

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs URL https://doi.org/10

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T19:41:48.225019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:41:48.225019Z digest=sha256:a9c6782e06f7ffb96073d2fa473247886d798decda6f8f1a83873c5c5da25b1b

Observation fe49c38d-1044-44ea-94c9-9c1697f51c5f · outbound

This paper cites Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning.

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T19:41:48.225019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:41:48.225019Z digest=sha256:c96b9dc5ecf61c4f6d470036de70f89b262676ab1024a5a14f2a2e674328454b

Observation 052a5d1a-d075-4b19-927d-2ba205bc81ed · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T19:41:48.225019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:41:48.225019Z digest=sha256:611f029e52dd964ef890ece573b6fd0a98bff727561cfeda5cc55fdd8e484364

Observation 75c94ad7-f8fb-4c72-932e-a3b685627ad8 · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T19:41:48.225019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:41:48.225019Z digest=sha256:6bce0cc55ab59230e6b153b5bbe40ea6d402e9835e90d5cce09ca242dae40425

Observation 69f43dcf-873f-48ff-bbd2-d6eed2f11802 · outbound

This paper cites 17 Appendix A Mamba SSM Pruning A.1 Method We chose which SSM channels to prune by estimating their contribution to the Mamba layer output in the following manner.

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs 17 Appendix A Mamba SSM Pruning A.1 Method We chose which SSM channels to prune by estimating their contribution to the Mamba layer output in the following manner

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T19:41:48.225019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:41:48.225019Z digest=sha256:713b8116b2c4117eab3d200fa30b11aa11e24156ea4f2c0f6f542b30866c4035

Observation 1b264ea6-46b9-4a6d-964f-1577307d0d34 · outbound

This paper cites Per-request prefill of the 990K-token prompt is roughly1.2×faster on Super Turbo than on Super.

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs Per-request prefill of the 990K-token prompt is roughly1.2×faster on Super Turbo than on Super

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-07-11T19:41:48.225019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:41:48.225019Z digest=sha256:957dcfd94ed25c97b8995105928d12d6d09336ce141492d34139ea1aedf84683

Observation 319c2fb7-d39b-4070-beca-7544883b3811 · outbound

This paper cites The projection-based method outperforms the coordinate-selection baselines at each tested latent size.

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs The projection-based method outperforms the coordinate-selection baselines at each tested latent size

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T19:41:48.225019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:41:48.225019Z digest=sha256:c033797993ee0320429b0fa3ac47a85e6cfc6a70d560077d2653ccd93a5e3024

Pith citing papers

No inbound Pith citation observations are available.