Pith. sign in

Paper Citation Record · LEDGER

Hymba: A Hybrid-head Architecture for Small Language Models

As of 19 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 43 inbound Pith citation observations for arXiv:2411.13676.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13676 v1

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:20:32.846517Z

measured 143 of 143 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:04:17.554378Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 103 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 680b0417-8773-43a0-8e7f-9b09b3a7a5ef · outbound

This paper cites Attention is all you need.Advances in Neural Information Processing Systems, 2017.

Hymba: A Hybrid-head Architecture for Small Language Models Attention is all you need.Advances in Neural Information Processing Systems, 2017

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.284974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.284974Z digest=sha256:55d5dccc1df47f3a89d242c2ea7b830d8b7fbcc6fea49d3b6ee5a27414aa7749

Observation b7babe13-0db6-457e-82c4-2e552c45e11a · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Hymba: A Hybrid-head Architecture for Small Language Models Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.291392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.291392Z digest=sha256:b5566aada3b4b06d4fffcdeb8cbd4d946a8ea1d84226d4b8bf92a3abba3a48a8

Observation 89b6ad65-3af9-477f-a3ac-e7c287dcefb3 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Hymba: A Hybrid-head Architecture for Small Language Models Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.296885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.296885Z digest=sha256:edbdf7afe41852fa65984d0c1757d61d83c8984d1cdea62ed91fe11a4e2acbfe

Observation b6e571fe-5d20-44ee-849c-cfb101d740c9 · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

Hymba: A Hybrid-head Architecture for Small Language Models An Empirical Study of Mamba-based Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.302339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.302339Z digest=sha256:82797dbf4d40c18672f5f202935448b671722b9b5299f769e96bd3b2490d578e

Observation ccac880c-419a-4823-9a74-99cfa328a5ef · outbound

This paper cites Simple linear attention language models balance the recall-throughput tradeoff.

Hymba: A Hybrid-head Architecture for Small Language Models Simple linear attention language models balance the recall-throughput tradeoff

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.309050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.309050Z digest=sha256:273b6f206c24c9ade2a7e33da69ebc876cefe9dc3b193e6fdb31e3c69bbe948f

Observation 365e0b9a-e682-4f90-99ec-adb327a270a4 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Hymba: A Hybrid-head Architecture for Small Language Models Jamba: A Hybrid Transformer-Mamba Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.314399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.314399Z digest=sha256:e15f79e510437a50e0b5d3d015bf7df5ddf836a6d7e34e3369db3bb1e6108f8e

Observation 36a6edb2-9ea8-40f7-8f2a-0ad9df9c4855 · outbound

This paper cites Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling.

Hymba: A Hybrid-head Architecture for Small Language Models Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.321340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.321340Z digest=sha256:fef342930354c3f307b1a92a11f2516df5e5c5d41984652e86dc3043b690a292

Observation 1eab19ac-ab9f-458d-ae84-d9f07161d8e0 · outbound

This paper cites Quantizable transformers: Removing outliers by helping attention heads do nothing.Ad- vances in Neural Information Processing Systems, 36:75067–75096, 2023.

Hymba: A Hybrid-head Architecture for Small Language Models Quantizable transformers: Removing outliers by helping attention heads do nothing.Ad- vances in Neural Information Processing Systems, 36:75067–75096, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.327340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.327340Z digest=sha256:667548ed28ff1f55647a766d094123a8216158a407c175e3ac0f5a07109be2a2

Observation d6f7b289-1374-4bd3-a17f-466c8f3a51eb · outbound

This paper cites Attention is off by one.

Hymba: A Hybrid-head Architecture for Small Language Models Attention is off by one

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.332796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.332796Z digest=sha256:15e74c15e4d187c791adec0fb4d680a7762484920d141c272a6a00b9551da5c5

Observation f8d69819-4e30-42c2-9f5e-cb8f74ce3df0 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Hymba: A Hybrid-head Architecture for Small Language Models Efficient Streaming Language Models with Attention Sinks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.337908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.337908Z digest=sha256:4cf75c609ad19b28ea0c1c339378778da47854c1575fd541449e6ec0c528d935

Observation baa5f6ec-1d50-4f0b-9471-2fe7120524d9 · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

Hymba: A Hybrid-head Architecture for Small Language Models Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.344744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.344744Z digest=sha256:2fc899310964984732b74ac533db6c1acf199c38ef7f60f602aa2d4877b4121b

Observation 1ad8a78a-1823-43bd-9217-9601d3c26c4f · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024.

Hymba: A Hybrid-head Architecture for Small Language Models Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.351411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.351411Z digest=sha256:c002c1f4034d8fa5b537cb13dbfd938d7e65b1eed4d28654b5ed2f46ff54d4f4

Observation cf706014-e8d5-4e99-b15d-8b48e360f942 · outbound

This paper cites Dora: Weight- decomposed low-rank adaptation.

Hymba: A Hybrid-head Architecture for Small Language Models Dora: Weight- decomposed low-rank adaptation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.357034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.357034Z digest=sha256:7a112437c1d8b963dbfde41ab5d9e7efb80742e15193a89ae8aefc439b60d599

Observation cd50ac20-f563-4bda-b000-3d2fb9916443 · outbound

This paper cites RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models.

Hymba: A Hybrid-head Architecture for Small Language Models RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.362269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.362269Z digest=sha256:345d8fc3c4fcf7b7c63a6c92e04f9cc93f301b07c57a4d70323a00dc735ee0c4

Observation 0e9bd003-c636-4f45-b859-121c1be6fdf6 · outbound

This paper cites Repeat After Me: Transformers are Better than State Space Models at Copying.

Hymba: A Hybrid-head Architecture for Small Language Models Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.368162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.368162Z digest=sha256:d9618ced67504f3ce0ad134549f82587877dd0ecd4abae8fb2acb4f60b6cc3ba

Observation 7131e50a-bb25-4aa6-83ab-f318ddd29d06 · outbound

This paper cites Decimamba: Exploring the length extrapolation potential of mamba, 2024.

Hymba: A Hybrid-head Architecture for Small Language Models Decimamba: Exploring the length extrapolation potential of mamba, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.373999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.373999Z digest=sha256:16baf62e91b89a7c959bd2717d052fb3f02567696778ab4a5ed36ab48b4328b4

Observation 0bf35d70-0d9a-4231-82d4-edfbee70b0c8 · outbound

This paper cites Zamba: A Compact 7B SSM Hybrid Model.

Hymba: A Hybrid-head Architecture for Small Language Models Zamba: A Compact 7B SSM Hybrid Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.379199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.379199Z digest=sha256:861b7b729941c24593286952b07b87bbb481bf20f42f46570a15d7a7a685eee8

Observation abd1283e-8c31-433e-8508-0dd2500a3c4d · outbound

This paper cites Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models.

Hymba: A Hybrid-head Architecture for Small Language Models Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.384743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.384743Z digest=sha256:bb25643d44ccca6691588e7415083585042d01ccfe2e4485214659031e4a033b

Observation d6726c42-5169-4d68-8ad3-6e22d63dc772 · outbound

This paper cites Talking Heads: Understanding Inter-layer Communication in Transformer Language Models.

Hymba: A Hybrid-head Architecture for Small Language Models Talking Heads: Understanding Inter-layer Communication in Transformer Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.390532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.390532Z digest=sha256:af50ff93a1b01a2d3ff81cf796e5359be9fd9dfa7a8f5118cc24117b7690601c

Observation d8706e9e-2cf8-4435-8259-39f7b32885db · outbound

This paper cites The Hidden Attention of Mamba Models.

Hymba: A Hybrid-head Architecture for Small Language Models The Hidden Attention of Mamba Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.396513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.396513Z digest=sha256:5d9b8686b39a79149e1b59da701513872a4a67749ba676f0329b2f57935d6447

Observation ff8cedd4-32ed-4362-bb4e-110aeabb4438 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.

Hymba: A Hybrid-head Architecture for Small Language Models Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.402349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.402349Z digest=sha256:87bc53a156496ab94c95463e8b9f7cfc7f005e402728f3141b3523245511f09b

Observation ec508037-f6d4-41a1-b849-9ac4e6e33d87 · outbound

This paper cites Longformer: The Long-Document Transformer.

Hymba: A Hybrid-head Architecture for Small Language Models Longformer: The Long-Document Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.407666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.407666Z digest=sha256:d83be9f1e36c918dd2ecdafac18e1773cab82c8ae4b97e1c79d2b3afaea760ed

Observation a54f002f-c183-4e53-88e0-19641b7e1091 · outbound

This paper cites MiniCache: KV Cache Compression in Depth Dimension for Large Language Models.

Hymba: A Hybrid-head Architecture for Small Language Models MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.412901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.412901Z digest=sha256:4bbe0958bce0f730553a48a5f86f40e2e99ebaae75de744990cb008eebf75167

Observation 616dfada-358d-4b3c-9072-8de28a846d50 · outbound

This paper cites SQuAD: 100,000+ questions for machine comprehension of text.

Hymba: A Hybrid-head Architecture for Small Language Models SQuAD: 100,000+ questions for machine comprehension of text

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.418068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.418068Z digest=sha256:b636268fde659be89b4d0f7d75bc1586018df291b15904aa172d76d24267b7fd

Observation 88ab6c05-250c-437a-a1af-02dbdfa49a68 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Hymba: A Hybrid-head Architecture for Small Language Models Training Verifiers to Solve Math Word Problems

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.423182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.423182Z digest=sha256:2107a7656cb794beb50928f0b3ca6e05ab15c745c91b18f05973ea83972510d2

Observation 1330a48e-f9b6-4671-b370-e0fb58ba821d · outbound

This paper cites Codeparrot/github-code · datasets at hugging face.

Hymba: A Hybrid-head Architecture for Small Language Models Codeparrot/github-code · datasets at hugging face

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.429431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.429431Z digest=sha256:00fa0f5402f67a1187d20d18279d6b7b8d3124927f36980dfdb698bb3133d248

Observation 19db1aa1-d488-4222-b898-838bd9e87b48 · outbound

This paper cites Lm-infinite: Zero-shot extreme length generalization for large lan- guage models.

Hymba: A Hybrid-head Architecture for Small Language Models Lm-infinite: Zero-shot extreme length generalization for large lan- guage models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.434742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.434742Z digest=sha256:65339e863875baadf3be98dce4278e7242d04304671a6842087ec099dc5cb6f6

Observation c68a54bb-3d63-4a46-81c4-53fda8165f8b · outbound

This paper cites A framework for few-shot language model evaluation, 12 2023.

Hymba: A Hybrid-head Architecture for Small Language Models A framework for few-shot language model evaluation, 12 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.439564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.439564Z digest=sha256:0e8b12927456c44d3cd9b95c833294add78ea06c4b3e56f53374fa0583080989

Observation 2bbaa4bd-2308-4c11-ad31-97bc5337a56e · outbound

This paper cites Simple lin- ear attention language models balance the recall- throughput tradeoff, 2024.

Hymba: A Hybrid-head Architecture for Small Language Models Simple lin- ear attention language models balance the recall- throughput tradeoff, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.444967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.444967Z digest=sha256:cfe998dd10d1a0497bdaaa975909cacdf811b78cf0328d4c4842db98d69714ba

Observation 90ff8672-6050-444f-8575-294829ebe974 · outbound

This paper cites Vision Transformers Need Registers.

Hymba: A Hybrid-head Architecture for Small Language Models Vision Transformers Need Registers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.451809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.451809Z digest=sha256:1ff8141a25ca12cb5ae06616c91707c1b15703752198ec414bee077720e553ab

Observation 56be0203-c52e-44a3-81e0-d720192ebaa3 · outbound

This paper cites Attention if off by one, 2023.

Hymba: A Hybrid-head Architecture for Small Language Models Attention if off by one, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.457365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.457365Z digest=sha256:a3b55436ceeb1904024dbf45e7a6d423506981ddb0c9740573e4c280a9685951

Observation fc87f738-bec5-4b84-9cce-ced81552c1d1 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Hymba: A Hybrid-head Architecture for Small Language Models The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.463210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.463210Z digest=sha256:752ef746c31922d61069dd5033bab5b588261cacc37356afc3aa03580f9b49ba

Observation 4a05abdf-034d-4bee-9039-8037ad9f7cac · outbound

This paper cites PPT: Pre-trained Prompt Tuning for Few-shot Learning.

Hymba: A Hybrid-head Architecture for Small Language Models PPT: Pre-trained Prompt Tuning for Few-shot Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.469793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.469793Z digest=sha256:b93a8517192e448769701d799f56b0a344b1025f0a2c74da214193d4dedd4935

Observation 7c9000d7-3e32-42df-96a3-42a970f93439 · outbound

This paper cites The Adventures of Oliver Twist.

Hymba: A Hybrid-head Architecture for Small Language Models The Adventures of Oliver Twist

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.475413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.475413Z digest=sha256:3f05ac20d2fcac82fbebac2252091adb12dbd53c5acf4647650254dcea0138db

Observation 3061283c-4d28-4fd3-9414-6f5b291602c8 · outbound

This paper cites Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation.

Hymba: A Hybrid-head Architecture for Small Language Models Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.480631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.480631Z digest=sha256:aaa939c9572aad82ffc956d1f9bac522075804b17d94ef1e77f848163ff02a88

Observation ab30e457-35f0-4c87-ac38-b0664f91c465 · outbound

This paper cites Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar.

Hymba: A Hybrid-head Architecture for Small Language Models Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.487084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.487084Z digest=sha256:0a6404e0753ceb6ee599265ad7efa2b31aa6c789d9358a5b3b8748cced19063b

Observation cae4d716-681c-43bd-8289-73cc594b3cfa · outbound

This paper cites Smollm-corpus, 2024.

Hymba: A Hybrid-head Architecture for Small Language Models Smollm-corpus, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.491936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.491936Z digest=sha256:46367843e0fab597616408ba15ebc390a0ab186dc2a6ca98186c7f6b7f68ae4d

Observation caf31356-346a-43a1-8f23-506738118c19 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Hymba: A Hybrid-head Architecture for Small Language Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.496829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.496829Z digest=sha256:604cc4544df54cfa8e85c1ebb651f4df47a87660a174bdd2037039e9d2a7804d

Observation b5d9d355-6f37-489e-be80-50920c2dc077 · outbound

This paper cites The Llama 3 Herd of Models.

Hymba: A Hybrid-head Architecture for Small Language Models The Llama 3 Herd of Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.501507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.501507Z digest=sha256:70c5a27512334deff0a7e50fede9da1bd79ae1a80b0a80b7407f46b1530b159d

Observation 296da1dc-63fe-4361-af32-02859fd87607 · outbound

This paper cites JetMoE: Reaching Llama2 Performance with 0.1M Dollars.

Hymba: A Hybrid-head Architecture for Small Language Models JetMoE: Reaching Llama2 Performance with 0.1M Dollars

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.506393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.506393Z digest=sha256:b80630186c1a642ad6d4ac5510a5975af38d0fa3a4af4eb9688efc3ba5861dd8

Observation 3df9b504-4256-49ef-879d-38694bac4cb1 · outbound

This paper cites Dynamically scaled rope further increases per- formance of long context llama with zero finetuning, July 2023.

Hymba: A Hybrid-head Architecture for Small Language Models Dynamically scaled rope further increases per- formance of long context llama with zero finetuning, July 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.513029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.513029Z digest=sha256:452782ceed41ba68f94826845e94a12698780f4f804ae4ffde9ed53152de34e0

Observation a0d2a8e5-58a7-4721-86c3-cd51d1d08bfa · outbound

This paper cites Llama 3.2: Revolutionizing edge AI and vision with open, customizable models.

Hymba: A Hybrid-head Architecture for Small Language Models Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.518667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.518667Z digest=sha256:04a1d4ccd7bdd0e09bb1496808a3bb5e375d724e2a26108b672fbfab89cbab21

Observation 05674b16-12c6-4989-90b2-3bc14d7d68d4 · outbound

This paper cites Smollm - blazingly fast and remarkably powerful, 2024.

Hymba: A Hybrid-head Architecture for Small Language Models Smollm - blazingly fast and remarkably powerful, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.523538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.523538Z digest=sha256:47d4b63b8c6e326b6fb67335ba85e52cd4cd8ae8ec017f30070148efa74ccc5a

Observation 1ceef134-b3aa-44cc-bede-2b90a580dda4 · outbound

This paper cites Smollm2 - with great data, comes great performance, 2024.

Hymba: A Hybrid-head Architecture for Small Language Models Smollm2 - with great data, comes great performance, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.528148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.528148Z digest=sha256:cecfa7cd41a5a83bef4519b57abb60e5610aeb608595a55b236d646c2e871678

Observation fbcb8e86-336d-4f74-8260-ca0adc84773a · outbound

This paper cites Amd-olmo: A series of 1b language models trained from scratch by amd on amd instinct™ mi250 gpus., October 2024.

Hymba: A Hybrid-head Architecture for Small Language Models Amd-olmo: A series of 1b language models trained from scratch by amd on amd instinct™ mi250 gpus., October 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.532645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.532645Z digest=sha256:97819dba817148cfd4bf2d4d4dd2a146de5a8c3ae130878832af1b0bb02bc6fb

Observation e3896155-8905-435a-9ee9-1aea20503140 · outbound

This paper cites Stable LM 2 1.6B Technical Report.

Hymba: A Hybrid-head Architecture for Small Language Models Stable LM 2 1.6B Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.537094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.537094Z digest=sha256:f496df2fefa9a3666cc97feae12025ddab1a1adffa921ef01d5e60534be98069

Observation 4e6c3d94-b939-42ac-98bb-d7d51d4ec30d · outbound

This paper cites an unresolved cited work.

Hymba: A Hybrid-head Architecture for Small Language Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.542119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.542119Z digest=sha256:1e95139c417c83b39041e7511e5f8ac4bf07a70e2ecffb36ba96fb4782018d7c

Observation be1b41a3-91d5-4def-997e-10f0fa9990c9 · outbound

This paper cites HuggingFaceTB/cosmo-1b.

Hymba: A Hybrid-head Architecture for Small Language Models HuggingFaceTB/cosmo-1b

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.547512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.547512Z digest=sha256:1902c5c168b7a03c569a8a1a0050a4aa52860331fc5c16ee16ae6f0c32796a48

Observation b37f3336-fccb-4338-b0c2-fb463294b7c4 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

Hymba: A Hybrid-head Architecture for Small Language Models Textbooks Are All You Need II: phi-1.5 technical report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.553121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.553121Z digest=sha256:86b765b7d301c40641b87125e266d579cce0699a7805d04bc5ea03169d5f9f63

Observation 033b67d2-727e-4f18-a939-b2ea0cd480fd · outbound

This paper cites H2O-Danube-1.8B Technical Report.

Hymba: A Hybrid-head Architecture for Small Language Models H2O-Danube-1.8B Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.559404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.559404Z digest=sha256:5ca809eb826bc09d5f19dc9ecb9065d4ebb1baa825322feabf69a7fb9b7fd371

Observation 0413a501-64b7-4ce5-8b63-038f12d9bb01 · outbound

This paper cites Openelm: An efficient language model family with open training and infer- ence framework.

Hymba: A Hybrid-head Architecture for Small Language Models Openelm: An efficient language model family with open training and infer- ence framework

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.605592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.564734Z digest=sha256:bf892f133574e24a646d3d00ac780b1a544e0383af103315af59306d542b2867

Observation 46161be2-7e94-47eb-81a3-a9a7b49db430 · outbound

This paper cites The On-Device Intelligence Update.

Hymba: A Hybrid-head Architecture for Small Language Models The On-Device Intelligence Update

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.589797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.569423Z digest=sha256:0495c1e0ce48b03e602466a5eb4b8e8fa01f9b4ade2b82b91d874de74213e651

Observation 5f05db3c-05b3-4b1e-8d73-52bf3addbfe6 · outbound

This paper cites Introducing Meta Llama 3: The most capable openly available LLM to date.

Hymba: A Hybrid-head Architecture for Small Language Models Introducing Meta Llama 3: The most capable openly available LLM to date

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.573954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.574279Z digest=sha256:9e7b38c6e969afe9fb35de3334199923ca65d52e5b938df8a68f36331de9627e

Observation 15cd0a8d-46ce-4adb-8f68-8c6f1be10571 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale, 2024.

Hymba: A Hybrid-head Architecture for Small Language Models The fineweb datasets: Decanting the web for the finest text data at scale, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.556144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.578717Z digest=sha256:d38fbf1a208c36039af212c68831d97ead1f2b572c7c3b0f4bed73ddfc0340a4

Observation f0af0be5-a3eb-465b-b186-2cb0ece0b238 · outbound

This paper cites OpenCeres: When open information extrac- tion meets the semi-structured web.

Hymba: A Hybrid-head Architecture for Small Language Models OpenCeres: When open information extrac- tion meets the semi-structured web

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.537885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.583890Z digest=sha256:44bb712de36cb435c87410d6cb48700e752c6128cb9778692bdaf3e01e0e2c95

Observation 5ecc7b57-4cab-466a-b602-4555974fedff · outbound

This paper cites Know what you don’t know: Unanswerable questions for squad, 2018.

Hymba: A Hybrid-head Architecture for Small Language Models Know what you don’t know: Unanswerable questions for squad, 2018

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.519467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.589064Z digest=sha256:d900fc5a3660c424c5c768f80b5007a9b4ca99060be644137d9371ed705361fd

Observation 0645f9af-d932-463c-9387-0862a26e8a30 · outbound

This paper cites Scaling Laws of RoPE-based Extrapolation.

Hymba: A Hybrid-head Architecture for Small Language Models Scaling Laws of RoPE-based Extrapolation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.594000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.594000Z digest=sha256:16ecc04bf7929fd098234a3238de7fa74f54783c346e2a1ffad211ba28d82c35

Observation b32b6aa2-5522-4e9c-9d22-54a4c927d5c1 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

Hymba: A Hybrid-head Architecture for Small Language Models YaRN: Efficient Context Window Extension of Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.599060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.599060Z digest=sha256:284c9e58fa5e14a147858b06941ea3394799710e8712b97b8834d710f6595a2c

Observation e016935e-bac5-493b-a86e-0ac1c4da9bef · outbound

This paper cites Inf bench: Extend- ing long context evaluation beyond 100k tokens.

Hymba: A Hybrid-head Architecture for Small Language Models Inf bench: Extend- ing long context evaluation beyond 100k tokens

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.503186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.604969Z digest=sha256:9d4dfdf9e41410aa5dde51848d89588238d8b01c7deb6e6c7bbf76459010116c

Observation 31c78362-9544-490e-8e41-50254e2057c0 · outbound

This paper cites Lost in the middle: How language models use long contexts.

Hymba: A Hybrid-head Architecture for Small Language Models Lost in the middle: How language models use long contexts

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.483400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.609923Z digest=sha256:858055aba85fff55db17fa2a6e675a871b76b75baa1d357c60e0e54b83d50a0d

Observation 53dd1e57-cae1-4b27-8e87-721fafa76a15 · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Hymba: A Hybrid-head Architecture for Small Language Models Zephyr: Direct Distillation of LM Alignment

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.614443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.614443Z digest=sha256:518f4a316eae9ee26d5672a0e5ab15f0b15c7e9c3f9215f82b34fe9c7ffe37c4

Observation 2409a0ef-cda0-4394-b326-c54fcc136470 · outbound

This paper cites Lmflow: An extensible toolkit for finetuning and inference of large foundation models.

Hymba: A Hybrid-head Architecture for Small Language Models Lmflow: An extensible toolkit for finetuning and inference of large foundation models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.465379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.619984Z digest=sha256:ba1af42612f4b20c5eade7240f56c8fb49a03d88e3f6ed8430effdb86dc9bf67

Observation 0dcee516-9775-4707-b5fb-406db6415ad2 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Hymba: A Hybrid-head Architecture for Small Language Models RLHF Workflow: From Reward Modeling to Online RLHF

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.624660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.624660Z digest=sha256:f97635ec1c5c24a6522cc4f84774e1a3dfaade0f820bf88da47ebe3ab101c6a4

Observation c642de7c-5d9c-4e4f-a650-22168a54b64c · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Hymba: A Hybrid-head Architecture for Small Language Models Qwen2.5: A party of foundation models, September 2024

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.630702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.630702Z digest=sha256:87380aee7603b923eae28622fe640b58c41c76457094551c7b4616aaee4ce114

Observation 64695960-143f-4a5e-9ed5-cd415cdc9296 · outbound

This paper cites Patil, Ion Stoica, and Joseph E.

Hymba: A Hybrid-head Architecture for Small Language Models Patil, Ion Stoica, and Joseph E

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.437227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.636151Z digest=sha256:b340b88ccf806c8104dcb98ed9d87517d6fd24ce931ed431c7c769a15c2e0982

Observation a3746c78-2e6c-4896-a81b-4d88e5fffe61 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Hymba: A Hybrid-head Architecture for Small Language Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.641414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.641414Z digest=sha256:540fe39e46d0b857bce44680c074f0ce42f990d29bc1fd7d46f9deb7e6e26f7e

Observation b756f457-6c40-490f-9ee0-57b8af45360a · outbound

This paper cites an unresolved cited work.

Hymba: A Hybrid-head Architecture for Small Language Models Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:20:34.419609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.646242Z digest=sha256:44485a30431090eaae1fcf8d69e1084ecaf1b003907925f422edf2bbf7a2dc73

Observation 772127fc-d81c-482e-a74c-8c86a214e4d9 · outbound

This paper cites Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$.

Hymba: A Hybrid-head Architecture for Small Language Models Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.651174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.651174Z digest=sha256:23e3426e6d57691063d802bfac53be0c812da03a856ed3ae25a78d34d09e682b

Observation 047ad6c2-2cf8-43d1-80bd-7ca9ba659b86 · outbound

This paper cites TransNormerLLM: A Faster and Better Large Language Model with Improved TransNormer.

Hymba: A Hybrid-head Architecture for Small Language Models TransNormerLLM: A Faster and Better Large Language Model with Improved TransNormer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.657559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.657559Z digest=sha256:8df17ee6c17f4131550b6f4e3a32b07aeacb07101b59c4b1b986110d13d0a7e5

Observation 9724b39c-ef0b-4e9d-8b6c-615ab811ea0d · outbound

This paper cites Scaling Laws for Neural Language Models.

Hymba: A Hybrid-head Architecture for Small Language Models Scaling Laws for Neural Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.662525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.662525Z digest=sha256:5f486474b9dcb93bb191c0f69fc2213b16c0248df3899f0c785fb660d45b3a40

Observation 1368313a-1517-4bd9-8dbc-342565407ef3 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Hymba: A Hybrid-head Architecture for Small Language Models Pythia: A suite for analyzing large language models across training and scaling

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.399440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.667530Z digest=sha256:2ce122b194f4578aad6b1ba359b9f3c651880d1344ebf4dea8fb769170dfc5b8

Observation 24de5703-f677-42d0-a3f8-0c4e18948c10 · outbound

This paper cites GLM: General Language Model Pretraining with Autoregressive Blank Infilling.

Hymba: A Hybrid-head Architecture for Small Language Models GLM: General Language Model Pretraining with Autoregressive Blank Infilling

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.674221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.674221Z digest=sha256:57f66e1349a8b86c985e3cf0d604af96852ceacdf0ed56ca8a89e4a0c38a5d1e

Observation d38f3495-d6c1-404e-894e-47dc525869ea · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Hymba: A Hybrid-head Architecture for Small Language Models OPT: Open Pre-trained Transformer Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.681439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.681439Z digest=sha256:0de9d8642ac4a33f2df89807415ec5a606c5aae4b14a08811467972ee4991696

Observation 178da746-bc93-45b6-b9c6-b366210b9a6e · outbound

This paper cites Mistral 7B.

Hymba: A Hybrid-head Architecture for Small Language Models Mistral 7B

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.687367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.687367Z digest=sha256:16adf35fd5186d1364ff08f1659f198c5684cdf62ea2f4665e1fb2a4183a6070

Observation 554820f4-fc6b-4b8b-ba72-9e6d399c26ec · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Hymba: A Hybrid-head Architecture for Small Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.697912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.697912Z digest=sha256:45e212d508f9f0ba6f2dd8d6dfc6c825de0563789d6425014354c86e7349c52c

Observation 5c4f27a7-868a-489d-a502-dda45593d24a · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Hymba: A Hybrid-head Architecture for Small Language Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.703994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.703994Z digest=sha256:82280df06b64a379728fd6bd4771a806904f19ca92522926bff5b58fbd46be1f

Observation ed7d84d0-6053-4dcd-8a16-0f28dc7cabb4 · outbound

This paper cites GPT-4 Technical Report.

Hymba: A Hybrid-head Architecture for Small Language Models GPT-4 Technical Report

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.711332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.711332Z digest=sha256:b6ff74d9005db83c0c11d6dcf55c593dbcf1356877c1d8fa7240e526add9bb34

Observation 2738873a-0eeb-4cfb-bde8-7bdcc9a3f562 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Hymba: A Hybrid-head Architecture for Small Language Models RWKV: Reinventing RNNs for the Transformer Era

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.717475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.717475Z digest=sha256:79dab5d952fcd503fe59a797a22148a5f8c1a6e9d31e6f9ca178e1c2e44e44db

Observation 5700e08d-826e-4db2-b63f-f7907257b64c · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Hymba: A Hybrid-head Architecture for Small Language Models Retentive Network: A Successor to Transformer for Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.723486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.723486Z digest=sha256:cad3cff904bdca422d5cbbf29642b8606d218677d770e00ab9e793f14e9fe174

Observation 6bf116a2-ac4f-44ad-b0c3-78be8e9fcb54 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Hymba: A Hybrid-head Architecture for Small Language Models Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.731248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.731248Z digest=sha256:b15f2578981aedfdd4550a99d4000346b9b9630843872fe571f44c3355ae28e5

Observation bfa4a184-e756-450c-8edd-6282fb87788d · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear atten- tion.

Hymba: A Hybrid-head Architecture for Small Language Models Transformers are rnns: Fast autoregressive transformers with linear atten- tion

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.381569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.738147Z digest=sha256:c0c6f571f19d5f687778117ae12e5b7b94086d0680ab636f220e5785ae58f3aa

Observation 7f23ff8e-6b62-4c9d-8e1a-f79088d7ba6a · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Hymba: A Hybrid-head Architecture for Small Language Models Efficiently Modeling Long Sequences with Structured State Spaces

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.744134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.744134Z digest=sha256:e5e711d060bd5902b0b0b961c7da89cccb675a61cffa63fcb3322d14cb4b4caa

Observation ff1bf771-e877-43d0-99c0-003eb8e3e7a8 · outbound

This paper cites Combining recurrent, convolutional, and continuous-time models with linear state space layers.Advances in neural information processing systems, 34:572–585, 2021.

Hymba: A Hybrid-head Architecture for Small Language Models Combining recurrent, convolutional, and continuous-time models with linear state space layers.Advances in neural information processing systems, 34:572–585, 2021

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.749783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.749783Z digest=sha256:5c272ece0a094cf1c622753d985ffabdcdfc3b1ea55a2865643a3c785cda124b

Observation 427a0b1a-4761-496c-bc1b-3c27e6828253 · outbound

This paper cites Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks.

Hymba: A Hybrid-head Architecture for Small Language Models Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.755868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.755868Z digest=sha256:5d3bdaed006f6d0a7a9f03b23abfcdad308c5515a028e7151134cff199dc9ca4

Observation 221fdb52-126f-4b54-bf02-399f2e468f51 · outbound

This paper cites Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models.

Hymba: A Hybrid-head Architecture for Small Language Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.762403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.762403Z digest=sha256:bb7665a5f2f6e68791d71c2510954715563dce0733a7b9bc9f0cdb63e6bbe9aa

Observation 9c15ad46-0b6c-42ef-a3f2-853f11c72e3d · outbound

This paper cites Block- state transformers.

Hymba: A Hybrid-head Architecture for Small Language Models Block- state transformers

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.353639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.768129Z digest=sha256:99538537ba59afee94e5ac35d5d8dfdad1ff7613641429395b91ebb3f3f7253e

Observation 1abb3645-4ada-4ffe-9b84-44fbd7aabdc4 · outbound

This paper cites Diag- onal state space augmented transformers for speech recognition.

Hymba: A Hybrid-head Architecture for Small Language Models Diag- onal state space augmented transformers for speech recognition

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:34.337491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.773195Z digest=sha256:d187912a76186eddfc1196a70811fb04b3198db231d18d4645cb36eb2262f9df

Observation 13209097-82c5-4c4e-a91b-74485a254335 · outbound

This paper cites Parallelizing Linear Transformers with the Delta Rule over Sequence Length.

Hymba: A Hybrid-head Architecture for Small Language Models Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.778566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.778566Z digest=sha256:9c6fefc43eac2f281eebe0ebbb4c186d1edb7f1fbb043b031102642f055d70db

Observation 01b40957-50f6-4b34-b5f6-444ab50118ab · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Hymba: A Hybrid-head Architecture for Small Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.785778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.785778Z digest=sha256:6fd3235ba163fa17c3f3896e82b8dc645530f85a5207b9311538580625852f2c

Observation ea642133-f126-4478-9e44-bf271734c227 · outbound

This paper cites Memory Transformer.

Hymba: A Hybrid-head Architecture for Small Language Models Memory Transformer

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.792479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.792479Z digest=sha256:11c7400a588df84526d222abd796e8042476bf606b1b3e04674e4d3dd8c59a46

Observation d948f96a-c773-4dc5-8c8b-3a17c9314eac · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Hymba: A Hybrid-head Architecture for Small Language Models A framework for few-shot language model evaluation, 07 2024

Reference 91

Resolution
verified exact
raw_fallback, observed 2026-08-12T16:20:33.077400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.798325Z digest=sha256:565b9a8ba863d0d3a128540b6e9bde33ea2af946b6f5d37e1ebe247eb09bd058

Observation 0f644df1-24e6-4a17-a1c0-472e4ff55e6f · outbound

This paper cites an unresolved cited work.

Hymba: A Hybrid-head Architecture for Small Language Models Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:20:34.320092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.803882Z digest=sha256:9a8bbb331f9a1afe2fdcec2702f511ad1488bab798dd103b812a7cb65610c1ab

Observation 9ea2b237-2054-4c78-bd6b-69498b9596f1 · outbound

This paper cites an unresolved cited work.

Hymba: A Hybrid-head Architecture for Small Language Models Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:20:34.304626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.809178Z digest=sha256:d2d9224648794f177cfb38ee63f9537cbbbeca9f763be0dabe21eee2c91f8ab0

Observation a9e769b4-4418-4fb2-9c30-200f13ac80a3 · outbound

This paper cites an unresolved cited work.

Hymba: A Hybrid-head Architecture for Small Language Models Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:20:34.286627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.814570Z digest=sha256:d9b150be0b161130d0e9efae17edb86fc64392d1f3bf8b96c08bb591edc5b97b

Observation 93b27070-6193-4306-97f8-32226790fcb6 · outbound

This paper cites an unresolved cited work.

Hymba: A Hybrid-head Architecture for Small Language Models Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:20:34.271386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.820863Z digest=sha256:e6b455b553a22285a855ce9e1876ea7e5c92710d89834edc2526fe2a05178282

Observation c240863e-5391-4c8f-992a-72e7375d71c2 · outbound

This paper cites an unresolved cited work.

Hymba: A Hybrid-head Architecture for Small Language Models Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:20:34.256058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.825954Z digest=sha256:e4028873e452e708ccf8011e5f723e93fc36caf199932b127fef55b0209243de

Observation 27524bd6-e83a-4b79-8182-9f5bbb45d8b9 · outbound

This paper cites an unresolved cited work.

Hymba: A Hybrid-head Architecture for Small Language Models Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:20:34.240168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.830982Z digest=sha256:c89843de07b5b00559117479896e539e83f7214c7b6d55eadf75811ba5e84848

Observation e177609c-60b2-4ec8-ba05-bfd6d75b6359 · outbound

This paper cites an unresolved cited work.

Hymba: A Hybrid-head Architecture for Small Language Models Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:20:34.221010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.836548Z digest=sha256:14472354ea3570c40fa09c5ae827b4e5c46f59c7fc21114085aab87ef6ef9e74

Observation 8146592d-b368-49c7-af4b-0a347ab669e0 · outbound

This paper cites an unresolved cited work.

Hymba: A Hybrid-head Architecture for Small Language Models Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:20:34.201109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.841292Z digest=sha256:bad4c404ad56c9eae82aae30154855deef3422c09188a732e090b60df5621f3b

Observation a9cd995e-1094-4218-948d-012d8ccbd4df · outbound

This paper cites an unresolved cited work.

Hymba: A Hybrid-head Architecture for Small Language Models Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:20:34.183542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T16:20:32.846517Z digest=sha256:f4cc551085ea10fe857e7807c5ac9309e547fa0fdbd07bd9ee594825490bb31f

Pith citing papers

Observation e8dd4dfb-103d-4fdd-ba9e-d44c2d708ff1 · inbound

Titans: Learning to Memorize at Test Time cites this paper.

Titans: Learning to Memorize at Test Time Hymba: A Hybrid-head Architecture for Small Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:15.324590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T22:08:14.982302Z digest=sha256:4524d3196c9c8994bfc1dfef11195204918ed25f0aee2ccf462c6d9132367957

Observation 3cf72c39-56e3-4580-92ce-4b5fb40f27dd · inbound

ARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer cites this paper.

ARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer Hymba: A Hybrid-head Architecture for Small Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:12:09.285754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:12:09.285754Z digest=sha256:27cc5cf3dba988691aedf404c2b20d49ebe176156366feb99927cca50623e78b

Observation fea57a3b-6a94-4969-a6af-02e7ca4a10e0 · inbound

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices cites this paper.

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices Hymba: A Hybrid-head Architecture for Small Language Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:05:16.456536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T01:03:26.037233Z digest=sha256:ea592906fafa13cfdeaffd97b91edc335ddb0502a6b13975331883f80a466f98

Observation f1b3ad46-c1d1-457c-8807-7c1adc1348d9 · inbound

Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration cites this paper.

Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration Hymba: A Hybrid-head Architecture for Small Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:11:57.910134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T19:11:08.841835Z digest=sha256:bd5fa3da643c3d609437e8697c631252f879a68a55036afdff7f145c89d9d84c

Observation fe8bfa86-92cd-4487-8296-e4d8b474fb52 · inbound

WuNeng: Hybrid State with Attention cites this paper.

WuNeng: Hybrid State with Attention Hymba: A Hybrid-head Architecture for Small Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:17.554378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:17.554378Z digest=sha256:8f57d3e311f6ceec17dff540ce6a6def594a93b7224055b7174ac183f5d6f331

Observation baeac036-1bd8-40a0-aafe-0c79739615b3 · inbound

Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation cites this paper.

Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation Hymba: A Hybrid-head Architecture for Small Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:22.035105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:56:22.035105Z digest=sha256:2070bcd3fd60e5b5765217052b63ad924530968a1917234af885342b7b9df54d

Observation 6e606860-5a8e-47d4-8058-ddd8ec7e988d · inbound

Overflow Prevention Enhances Long-Context Recurrent LLMs cites this paper.

Overflow Prevention Enhances Long-Context Recurrent LLMs Hymba: A Hybrid-head Architecture for Small Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:12:58.603905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:12:58.603905Z digest=sha256:7793dd1a3fb3e3db98c2a7790bfdecf1fba159f273ceb39dbb2b08805d3af514

Observation 82ae23a5-1ebd-4d31-bb29-1c836f3c8e44 · inbound

Balancing Computation Load and Representation Expressivity in Parallel Hybrid Neural Networks cites this paper.

Balancing Computation Load and Representation Expressivity in Parallel Hybrid Neural Networks Hymba: A Hybrid-head Architecture for Small Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:23.650165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:23.650165Z digest=sha256:42e2c4897ee9746b18249b768f3d418037894ccb8429a27ac11df4b162b5ca17

Observation d21d9efc-827a-407c-8150-307bcf69d8c2 · inbound

Small Language Models are the Future of Agentic AI cites this paper.

Small Language Models are the Future of Agentic AI Hymba: A Hybrid-head Architecture for Small Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:55:51.030625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T11:55:50.897500Z digest=sha256:4a538c7a9ef003bfda083af095b89025d85251acd01693c5747d8d8f03a7405e

Observation 3b22685c-af14-40c0-8907-133b90e7a683 · inbound

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques cites this paper.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques Hymba: A Hybrid-head Architecture for Small Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.896718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.896718Z digest=sha256:5b988bdcd80036ab713dd95cf5b629bca2151d036f1c4ef0584d8ce126918a13

Observation 5d536d28-56ee-48fd-98f7-a858af3525ea · inbound

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning cites this paper.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Hymba: A Hybrid-head Architecture for Small Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.591124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.591124Z digest=sha256:f85795a7210ccc340d3dad0790e8649d59d2270c858a99e2f6c91b28369fcd72

Observation 25719a15-1504-4c2e-84f1-b2e79e8d364b · inbound

SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot cites this paper.

SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot Hymba: A Hybrid-head Architecture for Small Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:48:38.651983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:48:38.651983Z digest=sha256:b71f65193fec6036fe127c4920fd26eb092bb63d913dde56bdbfc7657676d07f

Observation b8137f2f-5bd4-4a54-b85c-6ec7b1077466 · inbound

Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection cites this paper.

Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection Hymba: A Hybrid-head Architecture for Small Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:44.573565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:44.573565Z digest=sha256:260a5ea52b9133416022495a9a933816b32fc3eb72a63b75725fdd58ffdec384

Observation 0d34d1ff-41e9-479a-94c8-75fe139c49a6 · inbound

System-performance and cost modeling of Large Language Model training and inference cites this paper.

System-performance and cost modeling of Large Language Model training and inference Hymba: A Hybrid-head Architecture for Small Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:36:00.876438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:36:00.876438Z digest=sha256:396d2bd4dbbce80a2f063de5f83caf8f995c8212f24d0d08f3996597765c571f

Observation 4e4c4713-4027-417f-aa9d-b53bbadf77ac · inbound

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance cites this paper.

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Hymba: A Hybrid-head Architecture for Small Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:04.935445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:44:04.935445Z digest=sha256:559ecc7242e2c3c67d99a282be77d88002bdc2f253959c62ad72d4035a0e26f8

Observation 11be0f2c-da88-4262-8e93-b57e802bb9f9 · inbound

SpikingBrain: Spiking Brain-inspired Large Models cites this paper.

SpikingBrain: Spiking Brain-inspired Large Models Hymba: A Hybrid-head Architecture for Small Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:51:45.749016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T18:51:06.243305Z digest=sha256:5217fc6582fb56b840ab9f14ecb9b72d438bc495261ada0e21d2681d3c278d77

Observation 5a8b7ae4-45b6-48a6-93be-61fa93b10ba5 · inbound

Short window attention enables long-term memorization cites this paper.

Short window attention enables long-term memorization Hymba: A Hybrid-head Architecture for Small Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:11:21.791249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T12:10:42.646127Z digest=sha256:e38556937a5c924a1e53e369b9b305da63a372a7674d2ecf5133694fd365a25d

Observation 272efa11-27ed-4cf9-b197-38d0d2d85232 · inbound

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights cites this paper.

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights Hymba: A Hybrid-head Architecture for Small Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:21:15.114330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T10:18:04.431436Z digest=sha256:4f781d1e2f494e25acacea990772b622553e401e515c9b9d1b198cdc397d3df3

Observation 2877e21b-05ea-4b5d-99a9-8fb659bc4e14 · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture Hymba: A Hybrid-head Architecture for Small Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:49:11.040289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:5d5be8d8f0e2ae57295cb9d996897927521db0af408c997bbc16ce39b501498a

Observation 50527f6e-efcb-4e44-ac2f-79c9f4ddbb18 · inbound

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression cites this paper.

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression Hymba: A Hybrid-head Architecture for Small Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:27.018953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T17:59:23.826110Z digest=sha256:d5d9d4c34647195ce213095a0d25ace9964c0c24ad80349fc7c9a5f56aded007

Observation 72f4a784-24bb-4a8d-aaa1-e1595598e1d5 · inbound

Hidden State Poisoning Attacks against Mamba-based Language Models cites this paper.

Hidden State Poisoning Attacks against Mamba-based Language Models Hymba: A Hybrid-head Architecture for Small Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T18:18:14.502702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T18:15:05.108462Z digest=sha256:50a60a9862b47ad8c69662c9ac8584710a83b1c63f7594c1e037d945e0411a68

Observation 6be18e09-4e1c-4982-a403-69dfe3ce19d0 · inbound

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models cites this paper.

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models Hymba: A Hybrid-head Architecture for Small Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:41.380193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:41.380193Z digest=sha256:4a8c1dab136141633deb57e147eb93822d0ee23d5c3983f84e7d867bcf8671ad

Observation a624811f-df97-43d8-a275-32d8201fa572 · inbound

When to Think Fast and Slow? AMOR: Adaptive Entropy Gate for Hybrid Models cites this paper.

When to Think Fast and Slow? AMOR: Adaptive Entropy Gate for Hybrid Models Hymba: A Hybrid-head Architecture for Small Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:50:52.833832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T11:48:50.587732Z digest=sha256:8e64c31354d9158b0df27f16ff5715751d3abe19fbd6dc28caace22808a87dcf

Observation 00fb7d85-b6dd-487d-9948-49c2878c645e · inbound

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning cites this paper.

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning Hymba: A Hybrid-head Architecture for Small Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-15T12:09:16.468055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:09:16.468055Z digest=sha256:e337238d04a8f6a155842b71f61f8c6113d50de96984ec12e9ad3de5dcd32bbc

Observation 15f60ccd-f6c6-456f-89e0-71df9f3af120 · inbound

LPC-SM: Local Predictive Coding and Sparse Memory for Long-Context Language Modeling cites this paper.

LPC-SM: Local Predictive Coding and Sparse Memory for Long-Context Language Modeling Hymba: A Hybrid-head Architecture for Small Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:25:31.076370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T11:22:08.937342Z digest=sha256:cca3b8d13df206b19d643427bd29fb15ce8568ad41bbbfedda262d13f52b51b5

Observation 46ca10e9-ace1-44fb-b446-132a4ae48b7f · inbound

The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency cites this paper.

The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency Hymba: A Hybrid-head Architecture for Small Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:11:00.959927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:46:11.859634Z digest=sha256:24b42549e659ac3f7973a96614a2d80ef3981e7157601518b16834a487fd9d94

Observation 8bfaa0de-f961-4411-9688-08702050cf8b · inbound

A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models cites this paper.

A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models Hymba: A Hybrid-head Architecture for Small Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:29.249694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T14:24:53.745784Z digest=sha256:2d84ac283102d15b1c7b0e7c35e33eab18417482da30edcbb31e20814cb7b769

Observation c966df19-ca58-42f5-a39a-dbba1d7478aa · inbound

SpikingBrain2.0: Brain-Inspired Foundation Models for Efficient Long-Context and Cross-Platform Inference cites this paper.

SpikingBrain2.0: Brain-Inspired Foundation Models for Efficient Long-Context and Cross-Platform Inference Hymba: A Hybrid-head Architecture for Small Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:06.799830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T12:18:23.898779Z digest=sha256:41d76cf45efebab72ad040b0af529ae45dfdbf0e80295d9791b786321273b4c1

Observation 4b7653f2-4ada-4594-ba78-d7232a61fe89 · inbound

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators cites this paper.

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators Hymba: A Hybrid-head Architecture for Small Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:56.106800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T01:26:48.026538Z digest=sha256:e121ea786113d41c1e273a15c6cac3a65a2f8706942e6ee78533038d8a8baa62

Observation f0f0f8a2-f532-4519-9664-242e3c1996ef · inbound

A Single-Layer Model Can Do Language Modeling cites this paper.

A Single-Layer Model Can Do Language Modeling Hymba: A Hybrid-head Architecture for Small Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:01:25.421310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T04:40:12.906234Z digest=sha256:c8492b28bf19e0492bb7db5ce3c45722a7d988a20d9eac2c1bb2ad80fa2fc6b9

Observation 33411379-0fd7-478e-9fe3-057c3e355ed0 · inbound

Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference cites this paper.

Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference Hymba: A Hybrid-head Architecture for Small Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:31:14.110433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T07:28:35.020478Z digest=sha256:48177f818475d8095e27e87b277d51bb3bf415ef74605e0787bf796a75f3ceae

Observation cdca6b76-ee6b-47b0-a7b0-a94465b616f3 · inbound

Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference cites this paper.

Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference Hymba: A Hybrid-head Architecture for Small Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:43:59.519478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T21:37:51.638904Z digest=sha256:0dedbcd5aa87172c6206d98ede0f884196c78126bcec9618844a87fda46c23cc

Observation 3e582e7f-61ba-4159-8a71-56245017eaf3 · inbound

Forget Attention: Importance-Aware Attention Is All You Need cites this paper.

Forget Attention: Importance-Aware Attention Is All You Need Hymba: A Hybrid-head Architecture for Small Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.037769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T14:38:40.948032Z digest=sha256:0097c53e7be24bc0b4aebdbee402dea3f7822434932bf0e45fb77aeb626c8e66

Observation b3841260-6ed5-4578-b69f-1fae69a1aa46 · inbound

MOSAIC: Efficient Mixture-of-Agent Scheduling via Adaptive Aggregation and Inference Concurrency cites this paper.

MOSAIC: Efficient Mixture-of-Agent Scheduling via Adaptive Aggregation and Inference Concurrency Hymba: A Hybrid-head Architecture for Small Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-28T11:42:04.506675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T11:34:24.019560Z digest=sha256:cfc00fd0133ea3092946db44bc3c86a8004b5357bcae2544bbecccec4c9732ec

Observation 05a06ea1-e837-4055-bad9-33086f26cb31 · inbound

MOSAIC: A Workload-Driven Simulation and Design-Space Exploration Framework for Heterogeneous NPUs cites this paper.

MOSAIC: A Workload-Driven Simulation and Design-Space Exploration Framework for Heterogeneous NPUs Hymba: A Hybrid-head Architecture for Small Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:36:55.115884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T03:21:27.016397Z digest=sha256:df1b67a96728ffcd88ff93d07aa40fffb1b4869264b83da936774762f2c58bba

Observation df0f8b44-3048-4a66-82ff-512de8879512 · inbound

Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving cites this paper.

Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving Hymba: A Hybrid-head Architecture for Small Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.124721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T02:25:46.234576Z digest=sha256:fdfe307fb86d51d8107a6b63001551f66d09444daca982139ecefcd3bce979b6

Observation deb35f06-6355-4517-a654-b688b7971bd0 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Hymba: A Hybrid-head Architecture for Small Language Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:46.637432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:ac26454fa81d76d210ca7d22125674002348ce2c671b1d42869b6c74efcf7ba0

Observation d2d90e3c-333c-481d-9d6c-ac69050bca9d · inbound

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference cites this paper.

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference Hymba: A Hybrid-head Architecture for Small Language Models

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:17:46.087572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-11T00:12:13.830917Z digest=sha256:7d1c2d68aa7eb8feb2d5597544ea23885d0f7f5167c36a85805e50e03afab352

Observation 01e24ed9-e8bc-4d5b-b6bb-4b998673ba42 · inbound

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression cites this paper.

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression Hymba: A Hybrid-head Architecture for Small Language Models

Reference 85

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T03:14:31.609923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-08T03:07:23.648382Z digest=sha256:1c5b8d34978dcb8ed0872c3dc808316e2252c69a19fd7b6d24ea4138c60fb6ed

Observation d8c53253-2a75-4048-a577-8758951f92c0 · inbound

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale cites this paper.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Hymba: A Hybrid-head Architecture for Small Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T06:39:21.267463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:39:21.267463Z digest=sha256:5d4a475c611e5584732655a9c93ea027ac0d8a9da41bd72b6ac26231d7bcac06

Observation ac2f0a18-6a03-484b-8ddc-19bca428c2da · inbound

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale cites this paper.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Hymba: A Hybrid-head Architecture for Small Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T02:03:22.922600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:03:22.922600Z digest=sha256:8306a903e2f3b055cfffb4686b73e714a7f990dc98ab7f35f779cd187b55754c

Observation 79aaefee-9a18-4600-b590-70c339f69543 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Hymba: A Hybrid-head Architecture for Small Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:03.507389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:03.507389Z digest=sha256:61fc8cd0f6d4032ff9a26697bc73db1446a19b6730dabdd1ced97bb5ff8b3646

Observation 5b4b4e78-9383-48ee-aba3-01b541d41ce0 · inbound

Memory for Large Language Models cites this paper.

Memory for Large Language Models Hymba: A Hybrid-head Architecture for Small Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T02:37:54.543725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:37:54.543725Z digest=sha256:8facb042b353c6ee332bab16922a0b4dae05a92569811497921073fd1fbac24e