Pith. sign in

Paper Citation Record · LEDGER

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity

As of 10 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2501.16295.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16295 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T13:40:31.736571Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:27:49.926091Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:27:51.662445Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1e54f7f-f4be-46ba-bd09-0f5e4050a66d · outbound

This paper cites BlackMamba: Mixture of Experts for State-Space Models.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity BlackMamba: Mixture of Experts for State-Space Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.658057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.658057Z digest=sha256:09c5b1ae5a03e61babbed242f570ac0b4226354acef8a859e58d768cf2f78900

Observation 966c4c3c-53b9-4ea2-8352-04f941ff620c · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.665609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.665609Z digest=sha256:cc9c563b6e83e2acd307418d68ba175e2a9eb99fa86a5ca0fb3753399eb63431

Observation fa1eca84-31aa-4b5a-aa4d-8addbb37b868 · outbound

This paper cites Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.672333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.672333Z digest=sha256:805f3999f48862a25962270fc7a00a5e587cf83fa7fddd6a83b5a67d0765481d

Observation 8099324c-7760-4942-859e-eda2129029a9 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.675443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.675443Z digest=sha256:798bd5c8fd8c9a3f800e4e5f45ab9b4ead9996877014097b2c8de1f69ed98fe2

Observation e0113b73-1045-4717-84db-9e43b3b38861 · outbound

This paper cites Mixtral of Experts.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Mixtral of Experts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.684781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.684781Z digest=sha256:e4850754dd2e3f3fd3d562437f2ddd8bd641207497cfd2914048199bccf3b0dc

Observation f2a9dc46-5651-4177-8449-d4332f734b13 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.691269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.691269Z digest=sha256:32386830eab8d1d80dea77901b366646652b4cd1b288f2f4b44266222945b4cd

Observation a1c2de6c-835e-4a82-aeb2-60512abf1740 · outbound

This paper cites MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.694571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.694571Z digest=sha256:c76cd3157d9d6708d5bf37f5b2a429ce3bbd4a1e9e8da257f2096d76b093cbf6

Observation 0a70059c-3f42-41f5-86b0-cbb1009252f1 · outbound

This paper cites DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.697617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.697617Z digest=sha256:825545760fb763fb4883c294f8f0d7922dff79ab877157f81b8cada003ed6262

Observation 2482b423-b6b8-466f-82ed-b59f64d72a31 · outbound

This paper cites Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.700705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.700705Z digest=sha256:24f6c482f22f3c75ed865bb498499471414785d15d21f119690f06bd72f23028

Observation 738b8640-4580-4cb1-b3ef-3a9f2a5a2ea9 · outbound

This paper cites VL-Mamba: Exploring State Space Models for Multimodal Learning.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity VL-Mamba: Exploring State Space Models for Multimodal Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.704039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.704039Z digest=sha256:fe4a77f23a13d41d5b3f3e3911a8bed182f918ce247b4477444ea54ff1a5a947

Observation a5400b72-263e-45e9-900a-ea901e4fecd1 · outbound

This paper cites GLU Variants Improve Transformer.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity GLU Variants Improve Transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.707420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.707420Z digest=sha256:b31d77efa5522e0406b2654c29f2a89a3076f447b95aba693e408b339289e54a

Observation 90168bb9-d3d3-41a1-9f32-d9750a85bdeb · outbound

This paper cites ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.717857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.717857Z digest=sha256:c98c6fab63456769607a4a14269815ed28acb0366e7caced283340e1af69d71d

Observation 1e1cd6be-58ee-4e87-a7d3-7c128ff908d7 · outbound

This paper cites Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.724880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.724880Z digest=sha256:9aba124ca0b672110470d34eb68b6386e49339d8bcbdfc1abc413f4b564f6743

Observation 398e87de-fad2-47ad-bc3b-faba6250862c · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.727960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.727960Z digest=sha256:18dd670d97d0f8e94399687a3cb4764e2e017aeb0cbfbd24e0e80d96720b2ed3

Observation 6d64e3e0-bdfa-4af8-b635-676b2c3e3308 · outbound

This paper cites Specialized Foundation Models Struggle to Beat Supervised Baselines.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Specialized Foundation Models Struggle to Beat Supervised Baselines

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.730810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.730810Z digest=sha256:0c7212a61bb69fb58d3d361f2f720949015a8a84c2929511751ca78f8e624455

Observation bf36deb8-d066-4e15-9593-2faf5df2f54f · outbound

This paper cites Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.733836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.733836Z digest=sha256:6308172442e0b0c864a51a483b256baa32063746d4fe066a579cdeee09aabd8f

Observation a7c3202c-204c-4847-8af2-cb80fbce470b · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.736571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.736571Z digest=sha256:8fb57cd015f8dfc31602d9f592807e415beed84423f8da962cfd390c8645b111

Observation 75d6bb2c-7e3c-4715-9fcd-5e3d7d6ed844 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.714813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.714813Z digest=sha256:f2295e043543425db53ba4551ff23249a14ccad3a994d0cf4f7ed385a1440369

Observation 2e2c2643-cacd-4412-9d24-3376fe7ae497 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.688207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.688207Z digest=sha256:6d1e80079640927416aaa07cd4fe7a59f678e653725e6e3b6ba4d7b902e1b424

Observation 16726869-fae9-4895-8147-a60da36e5519 · outbound

This paper cites MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.681892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.681892Z digest=sha256:3686a9b02feb4e17e5b1d11233c4f1c9d0b522d54b62db4b88e99bd45e1a45a4

Observation 06815254-d7f4-444c-be73-7f0e53d1ebc6 · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.668994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.668994Z digest=sha256:1b81788337c1593fad5f4afc8f461ef6bf202ef46be57ecf24526d1cf5be5ee0

Observation 4f1c4b22-55f7-4c60-978a-c1dfb61d4965 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Efficiently Modeling Long Sequences with Structured State Spaces

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.678996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.678996Z digest=sha256:4117c4e2780e2e1ba92f4a730f6bfbc9a49887e0ef7d746a40f5b37547c72410

Observation b1286cb2-0b10-49ae-8446-e159a33033df · outbound

This paper cites VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.662404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.662404Z digest=sha256:7257377d20d91825c215e72ac65248f4fb8e1d0ab8547ed427f98c81e3f943ba

Observation 2576ec84-5f26-41ae-942b-d852836e6f68 · outbound

This paper cites CAT: Content-Adaptive Image Tokenization.

Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity CAT: Content-Adaptive Image Tokenization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T13:40:31.721604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:40:31.721604Z digest=sha256:3c08f91b54a17976e2d293f276c27ac22e631b6f41fe5c3fc5281d528cf23c40

Pith citing papers

Observation f13cb217-ff8e-47c5-a4bf-d67e8501f29f · inbound

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction cites this paper.

Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:27:51.667553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:27:49.926091Z digest=sha256:589386e343046e93d3554347bb3d4e4a3595a5ad5feb507cbd4de5b287ce1cbc