Pith. sign in

Paper Citation Record · LEDGER

MixFormer: Linear Transformer with Mixture of Memory Experts

As of 12 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2608.09468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09468 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:02:06.937137Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bd4a140-c4f3-4aca-8f13-fe8cf5344647 · outbound

This paper cites Attention is all you need.Proceedings of the 2017 Advances in Neural Information Processing Systems, NeuraIPS, pages 5998–6008, 2017.

MixFormer: Linear Transformer with Mixture of Memory Experts Attention is all you need.Proceedings of the 2017 Advances in Neural Information Processing Systems, NeuraIPS, pages 5998–6008, 2017

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.657671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.733397Z digest=sha256:5adb6c9c1804c1286391a5f895e499c3693e03246670aa91adbb2c371eaae71b

Observation 2a19aaee-0952-4df1-9c0a-856f4aa08113 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MixFormer: Linear Transformer with Mixture of Memory Experts LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.739145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.739145Z digest=sha256:b6894d5c569943de1b9ee4582db76770b40f09c723bda320cdd467cebce692a0

Observation f1eba3b0-f82b-450e-84bb-51e0e2505169 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MixFormer: Linear Transformer with Mixture of Memory Experts Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.744470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.744470Z digest=sha256:0577da59e2b38b35c59e1aeb31d9635316ed1e180784f34bd3842273159c530b

Observation 0fb668da-41b8-48ee-9f55-4115ea066474 · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407, 2024.

MixFormer: Linear Transformer with Mixture of Memory Experts The llama 3 herd of models.arXiv e-prints, pages arXiv–2407, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.749457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.749457Z digest=sha256:618717d881931b8ee0117387cca931765d767050b931b07e11c7ae0f12068581

Observation 77a00a8d-2fa9-42c1-81bf-23e2aa35a402 · outbound

This paper cites Molmo and pixmo: open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024.

MixFormer: Linear Transformer with Mixture of Memory Experts Molmo and pixmo: open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.630860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.755293Z digest=sha256:8c3d85e04b62ade0d916805952da54da5b5ffa554bf843a0a52c240355cad68c

Observation 6a6a3153-d04c-4a01-aef2-f61433ec6a6b · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

MixFormer: Linear Transformer with Mixture of Memory Experts NVLM: Open Frontier-Class Multimodal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.760415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.760415Z digest=sha256:ff2f91c5462efffa16e3b6ae5c6bd157aa9a7944761538f5907b0b54b833c8d1

Observation c4a816df-d03b-4ba8-8028-3b5009247283 · outbound

This paper cites Efficient transformers: a survey.ACM Computing Survey, 55(6), 2022.

MixFormer: Linear Transformer with Mixture of Memory Experts Efficient transformers: a survey.ACM Computing Survey, 55(6), 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.615287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.767004Z digest=sha256:44ce9a31844cca82890ff34875d38ca5862ce0aaf5b78708a06a70d80fae4d00

Observation 2f607eb3-f53f-4eee-b132-8c578e47b8de · outbound

This paper cites A survey on efficient training of transformers.

MixFormer: Linear Transformer with Mixture of Memory Experts A survey on efficient training of transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.598109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.772520Z digest=sha256:40bf7eec59bfd599c58674715430e898da042234a9c26f4059227580fb7950be

Observation e1c29fcc-6bc8-4563-a0c0-8ef0102c0970 · outbound

This paper cites an unresolved cited work.

MixFormer: Linear Transformer with Mixture of Memory Experts Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:02:07.581779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.777736Z digest=sha256:ab0c2483a71eb2bfa66c14fe5c62105e2ffdc68c160e5d3e23cd81c31df18247

Observation 1eb31702-f9b0-4ebf-b0f6-2acd135bf81d · outbound

This paper cites A survey on vision transformer.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1):87–110, 2022.

MixFormer: Linear Transformer with Mixture of Memory Experts A survey on vision transformer.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1):87–110, 2022

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.566504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.783024Z digest=sha256:fc5cdd9ff8794487d0ef9fb7db19b83dab0579d29b5d4f7165b0a3ce54b16fbc

Observation 1df62038-dde7-43b7-b1fb-94f733cdd8a7 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

MixFormer: Linear Transformer with Mixture of Memory Experts Efficiently modeling long sequences with structured state spaces

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.551246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.789114Z digest=sha256:383d38f96cc50b7be3c066258e4583032746e384dcc4832dda53f8d6059f4aa7

Observation dfa9049a-e736-42ae-b409-4e28161aa6c2 · outbound

This paper cites Diagonal state spaces are as effective as structured state spaces.

MixFormer: Linear Transformer with Mixture of Memory Experts Diagonal state spaces are as effective as structured state spaces

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.534221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.794479Z digest=sha256:31a221739e5910005f48329ac56d95b722e74ea701a9a22b2cafbe5e1888a49b

Observation 5a49165a-d784-4664-b347-f09b3875177e · outbound

This paper cites an unresolved cited work.

MixFormer: Linear Transformer with Mixture of Memory Experts Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:02:07.519106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.799632Z digest=sha256:99ff195c4ed120a009e58d66aa4848f75575e907c1ebcadf4e92c56462d33ebc

Observation fedf7564-1b39-4c14-a9b6-39fad529bd22 · outbound

This paper cites Liquid structural state-space models.

MixFormer: Linear Transformer with Mixture of Memory Experts Liquid structural state-space models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.503407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.804447Z digest=sha256:1d40ef227f16c4da89bb246af253116d75c32884112858d6e43183551e1979ea

Observation 7484d9bd-91c1-4cbc-97c9-821d3219d5db · outbound

This paper cites Simplified state space layers for sequence modeling.

MixFormer: Linear Transformer with Mixture of Memory Experts Simplified state space layers for sequence modeling

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.488124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.809142Z digest=sha256:038b3c9bbdf8bf0fd07662a76b5c274d4392b58ef10d4b1bee93fc22f7d2d28e

Observation affa02d5-02da-4bd1-9626-0ad06068fbf0 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

MixFormer: Linear Transformer with Mixture of Memory Experts Retentive Network: A Successor to Transformer for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.813888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.813888Z digest=sha256:9809ef0d3f614f18f89268197de06f6967a806e9228227e018f651e4a0b38b51

Observation a4d0d1ab-6067-46f6-b07d-9b14f63f7418 · outbound

This paper cites Gated linear attention transformers with hardware-efficient training.

MixFormer: Linear Transformer with Mixture of Memory Experts Gated linear attention transformers with hardware-efficient training

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.472604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.818921Z digest=sha256:db05cd816021aca01600bd7bdacbc13c0555ec1a7fae97cc3667244e5efcb909

Observation d97789cb-2202-4d0a-abf9-2067798803ac · outbound

This paper cites Transformers are rnns: fast autoregressive transformers with linear attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Transformers are rnns: fast autoregressive transformers with linear attention

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.456246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.823323Z digest=sha256:9ee6fbf571dbe8af36d3d6011a28d96bf559d2f3c7204a116d04f810d83a2022

Observation 3efb5ef3-2a30-4674-972b-c354604964b1 · outbound

This paper cites Efficient attention: attention with linear complexities.

MixFormer: Linear Transformer with Mixture of Memory Experts Efficient attention: attention with linear complexities

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.439782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.827677Z digest=sha256:edcdfa5be83fad5eb5fe67d4282cb777a5413eaecbe364ef889c7fc9dffa5168

Observation 7af4b9ed-8273-4831-8def-e00b15592c32 · outbound

This paper cites Linear transformers are secretly fast weight programmers.

MixFormer: Linear Transformer with Mixture of Memory Experts Linear transformers are secretly fast weight programmers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.423943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.832401Z digest=sha256:0d84648868835b75121ba5b8ad956791ae34585637c21d788c86d3d668d4808e

Observation 349fb060-3e53-4895-95f2-46bc8956a66f · outbound

This paper cites Flatten transformer: vision transformer using focused linear attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Flatten transformer: vision transformer using focused linear attention

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.408559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.837260Z digest=sha256:a32598e0af48b040079b30714a43dc07f8a5a8cfe67b475772281a52dabdd75f

Observation 902f38e1-f551-4875-817b-b88076078383 · outbound

This paper cites Polaformer: polarity-aware linear attention for vision transformers.

MixFormer: Linear Transformer with Mixture of Memory Experts Polaformer: polarity-aware linear attention for vision transformers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.389436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.841842Z digest=sha256:d5ef31611fba37c0c1dce436836a9a4b2d9257f748cb664ddefe26c884cd5714

Observation e04374ce-f3de-4284-ac12-4545da574cf5 · outbound

This paper cites Linear Attention Mechanism: An Efficient Attention for Semantic Segmentation.

MixFormer: Linear Transformer with Mixture of Memory Experts Linear Attention Mechanism: An Efficient Attention for Semantic Segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.846478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.846478Z digest=sha256:93f38820eaa40ad8162b2e2c2d99775d20c9a971bbe7e7368fc92164a8830041

Observation ec9b2a3f-6171-4bbc-836e-9a14c1f6e354 · outbound

This paper cites Random feature attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Random feature attention

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.370028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.851757Z digest=sha256:98272c0094b589543a75dd3e0159b641ffac3fed19db4fc186cac9da6922d20b

Observation df083ca3-9ea6-4d5e-8adb-151c05bc334b · outbound

This paper cites Rethinking attention with performers.

MixFormer: Linear Transformer with Mixture of Memory Experts Rethinking attention with performers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.354293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.856505Z digest=sha256:61223f9802cb44180c2422b95e42313349fbbd5942318fa20ed15f47b2da1a9e

Observation 207c6656-18d6-4fc9-ab07-ec9c10cf617b · outbound

This paper cites Mamba: linear-time sequence modeling with selective state spaces.

MixFormer: Linear Transformer with Mixture of Memory Experts Mamba: linear-time sequence modeling with selective state spaces

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.337411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.861100Z digest=sha256:7d6285823d0d427a7198e8cd3a9e7799c27499bb25955441e513254cd2ee86d9

Observation 6bcdf29d-0268-42a2-936e-30e01f8fa144 · outbound

This paper cites Transformers are ssms: generalized models and efficient algorithms through structured state space duality.

MixFormer: Linear Transformer with Mixture of Memory Experts Transformers are ssms: generalized models and efficient algorithms through structured state space duality

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.321854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.865999Z digest=sha256:a73222447c99265359e422917047610ddd0aef1752f2a0e3248b473cd4e5342b

Observation b3b9a948-da53-456b-918c-639b4a1da352 · outbound

This paper cites An Attention Free Transformer.

MixFormer: Linear Transformer with Mixture of Memory Experts An Attention Free Transformer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.870460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.870460Z digest=sha256:0951f9f82bd194b13247738e6540e3c9b5d8ec7ce9995905e9dee6ccebe66e0f

Observation 366b724a-9899-4afc-89e5-c6030e9a6aa5 · outbound

This paper cites Rwkv: Reinventing rnns for the transformer era.

MixFormer: Linear Transformer with Mixture of Memory Experts Rwkv: Reinventing rnns for the transformer era

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.305242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.875703Z digest=sha256:0db06d2c3aa44b941fa86465ed7aab97b04998c059bdb0be7e95d8ffe926dbcb

Observation e551ac0a-a7a5-42e3-954b-488100443664 · outbound

This paper cites Group normalization.

MixFormer: Linear Transformer with Mixture of Memory Experts Group normalization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.880665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.880665Z digest=sha256:1b2b5e739766b7340ef37d02a3ed4ee01f7fcb85b3ea5dbfac2a30e25a8b40cc

Observation fded8756-0e96-4f2e-b5e6-fe98c74c77b2 · outbound

This paper cites EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction.

MixFormer: Linear Transformer with Mixture of Memory Experts EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.885235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.885235Z digest=sha256:e1fbb271c10f71d42b128972e76c842b902f211c263fbf822e82124b74871481

Observation dd78c79f-08f4-42e2-ae0f-67a5679d4a35 · outbound

This paper cites Soft: softmax-free transformer with linear complexity.Proceedings of the 2021 Advances in Neural Information Processing Systems, NeuraIPS, 34:21297–21309, 2021.

MixFormer: Linear Transformer with Mixture of Memory Experts Soft: softmax-free transformer with linear complexity.Proceedings of the 2021 Advances in Neural Information Processing Systems, NeuraIPS, 34:21297–21309, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.274337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.890292Z digest=sha256:abeeb8edfd47859093d50ed8d828f6f51c1c56eda37a65f9f84885a3bf4054e7

Observation f91b6789-a3d1-4e26-9a39-a2b67165695a · outbound

This paper cites Nystr¨omformer: a nystr ¨om-based algorithm for approximating self-attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Nystr¨omformer: a nystr ¨om-based algorithm for approximating self-attention

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.254217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.895575Z digest=sha256:62494772111249537d24d63554218558ddd0ebe7d9e7998aa5e94ad24662aa78

Observation 594cde54-6c43-4f90-abf2-a79a4cdf328c · outbound

This paper cites Outrageously large neural networks: the sparsely-gated mixture-of-experts layer.

MixFormer: Linear Transformer with Mixture of Memory Experts Outrageously large neural networks: the sparsely-gated mixture-of-experts layer

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.236513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.900238Z digest=sha256:e518dd8b7c5e929cd85cc6c5ce355f23c7184bd3bfa7f8887ad9cb9aa0437142

Observation 11156257-0005-4f0a-aff0-29a3e7c52132 · outbound

This paper cites Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization.

MixFormer: Linear Transformer with Mixture of Memory Experts Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.217429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.904806Z digest=sha256:8808967a146cd80e65a92c6a706286b13ec898c3a63680fd753e6ce16d9475d0

Observation 12c2d22a-8cee-4abb-9529-43cf6d45b9cb · outbound

This paper cites Long range arena: a benchmark for efficient transformers.

MixFormer: Linear Transformer with Mixture of Memory Experts Long range arena: a benchmark for efficient transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.198005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.909593Z digest=sha256:987eb6e1287caef352da0e06d648c35b8ccdc3d90fd8e8590d50138dfbf85c79

Observation 985d0c57-9069-4181-bcf6-3803bf82d05e · outbound

This paper cites Listops: a diagnostic dataset for latent tree learning.

MixFormer: Linear Transformer with Mixture of Memory Experts Listops: a diagnostic dataset for latent tree learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.180622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.914154Z digest=sha256:d8da38c2386c8752e0956c218f0ff58c02aeecbe7d9e1c15cc1dd3f2fefa0fb2

Observation b2114500-42d9-4d0a-8f99-32bfeb3f4060 · outbound

This paper cites Learning word vectors for sentiment analysis.

MixFormer: Linear Transformer with Mixture of Memory Experts Learning word vectors for sentiment analysis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.162071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.918732Z digest=sha256:8ca7b78661718f477864e6dc9927f6b49c41d832642aa50d6782ee53d9b5a2ac

Observation e8449253-7c69-42bd-b3e2-7f360bc81a69 · outbound

This paper cites The acl anthology network corpus.Language Resources and Evaluation, 47(4):919–944, 2013.

MixFormer: Linear Transformer with Mixture of Memory Experts The acl anthology network corpus.Language Resources and Evaluation, 47(4):919–944, 2013

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.146128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.923391Z digest=sha256:cf5e47c5f3c832a7bb5cd65472545d5b7e7899472440a9ac2c9eab0dd599bbd4

Observation a66fb356-0427-48a0-ba50-37fff294600b · outbound

This paper cites Learning long-range spatial dependencies with horizontal gated recurrent units.Proceedings of the 2018 Advances in Neural Information Processing Systems, NeuraIPS, 31, 2018.

MixFormer: Linear Transformer with Mixture of Memory Experts Learning long-range spatial dependencies with horizontal gated recurrent units.Proceedings of the 2018 Advances in Neural Information Processing Systems, NeuraIPS, 31, 2018

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.128493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.927739Z digest=sha256:050eac2c233fc0d85ceedb91137d39174dec9e5f41528b689a70363def84a4bc

Observation 867a0f1f-591f-4288-9ec6-cfd59ae19a5d · outbound

This paper cites Learning multiple layers of features from tiny images.

MixFormer: Linear Transformer with Mixture of Memory Experts Learning multiple layers of features from tiny images

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.932524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.932524Z digest=sha256:dcfcf143e72d60352c281eafc9819a5d6205ee37e5e7b8bfba73b214cf64617e

Observation 12adb45a-901c-4675-b703-276e1affadcb · outbound

This paper cites Emnist: Extending mnist to handwritten letters.

MixFormer: Linear Transformer with Mixture of Memory Experts Emnist: Extending mnist to handwritten letters

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.101239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.937137Z digest=sha256:7203732149eb73960df85e892c049ad97518980dea8bc45d309ef312570e8630

Pith citing papers

No inbound Pith citation observations are available.