Pith. sign in

Paper Citation Record · LEDGER

MixFormer: Linear Transformer with Mixture of Memory Experts

As of 13 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2608.09468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09468 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:02:06.937137Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bd4a140-c4f3-4aca-8f13-fe8cf5344647 · outbound

This paper cites Attention is all you need.Proceedings of the 2017 Advances in Neural Information Processing Systems, NeuraIPS, pages 5998–6008, 2017.

MixFormer: Linear Transformer with Mixture of Memory Experts Attention is all you need.Proceedings of the 2017 Advances in Neural Information Processing Systems, NeuraIPS, pages 5998–6008, 2017

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.657671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.733397Z digest=sha256:9e16d3da6f7de7f81c3e03d2fe62be793d207cad0101380a69be8d82a0da22ad

Observation 2a19aaee-0952-4df1-9c0a-856f4aa08113 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MixFormer: Linear Transformer with Mixture of Memory Experts LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.739145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.739145Z digest=sha256:0a4391ef3ad860ff0eaa5268525ee80de78d968a8873e9cdb126322dec0856a6

Observation f1eba3b0-f82b-450e-84bb-51e0e2505169 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MixFormer: Linear Transformer with Mixture of Memory Experts Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.744470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.744470Z digest=sha256:f3cacee6af664f5000b02bac6ee219beb91a84bea907c457adde963774210e5b

Observation 0fb668da-41b8-48ee-9f55-4115ea066474 · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407, 2024.

MixFormer: Linear Transformer with Mixture of Memory Experts The llama 3 herd of models.arXiv e-prints, pages arXiv–2407, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.749457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.749457Z digest=sha256:5d1a5ac9b2533e18c7099b6c6e13aaad1c561de3e44f4770ee91947401607360

Observation 77a00a8d-2fa9-42c1-81bf-23e2aa35a402 · outbound

This paper cites Molmo and pixmo: open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024.

MixFormer: Linear Transformer with Mixture of Memory Experts Molmo and pixmo: open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.630860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.755293Z digest=sha256:0dc55e0a58fe3965fa1048474f0bbdbf11fbce07e0997d7f6f635179de7c99b9

Observation 6a6a3153-d04c-4a01-aef2-f61433ec6a6b · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

MixFormer: Linear Transformer with Mixture of Memory Experts NVLM: Open Frontier-Class Multimodal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.760415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.760415Z digest=sha256:9baeade4d3ca2a1e665de0f4fdd14f70e9a7465218dc355bc5ada3d8b95c5956

Observation c4a816df-d03b-4ba8-8028-3b5009247283 · outbound

This paper cites Efficient transformers: a survey.ACM Computing Survey, 55(6), 2022.

MixFormer: Linear Transformer with Mixture of Memory Experts Efficient transformers: a survey.ACM Computing Survey, 55(6), 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.615287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.767004Z digest=sha256:ab05e78af2813ead1a84506a641257face7490469617cb234ac87507e0e0a125

Observation 2f607eb3-f53f-4eee-b132-8c578e47b8de · outbound

This paper cites A survey on efficient training of transformers.

MixFormer: Linear Transformer with Mixture of Memory Experts A survey on efficient training of transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.598109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.772520Z digest=sha256:5dcb0c04add1df7b10e1c848dd3b61405e2a321d1e7d9c5695b85404d3d4a69a

Observation e1c29fcc-6bc8-4563-a0c0-8ef0102c0970 · outbound

This paper cites an unresolved cited work.

MixFormer: Linear Transformer with Mixture of Memory Experts Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:02:07.581779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.777736Z digest=sha256:ba5d108a0013d728483755ee7c6b96cd4ba88e9b693d84a49b1aeec0c52546cb

Observation 1eb31702-f9b0-4ebf-b0f6-2acd135bf81d · outbound

This paper cites A survey on vision transformer.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1):87–110, 2022.

MixFormer: Linear Transformer with Mixture of Memory Experts A survey on vision transformer.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1):87–110, 2022

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.566504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.783024Z digest=sha256:af87e1e6ab54d2ff0bb957529a188fd15787f9426ba9688bbeb08feca8a3bf16

Observation 1df62038-dde7-43b7-b1fb-94f733cdd8a7 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

MixFormer: Linear Transformer with Mixture of Memory Experts Efficiently modeling long sequences with structured state spaces

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.551246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.789114Z digest=sha256:f986af83647ba422fcb54cf74c19319806c7229ca46ebcde2abe28196edba21b

Observation dfa9049a-e736-42ae-b409-4e28161aa6c2 · outbound

This paper cites Diagonal state spaces are as effective as structured state spaces.

MixFormer: Linear Transformer with Mixture of Memory Experts Diagonal state spaces are as effective as structured state spaces

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.534221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.794479Z digest=sha256:4761821033e54cd9359304d380a3ecd2b6c03756223b981a0a78e2a9ba53a30e

Observation 5a49165a-d784-4664-b347-f09b3875177e · outbound

This paper cites an unresolved cited work.

MixFormer: Linear Transformer with Mixture of Memory Experts Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:02:07.519106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.799632Z digest=sha256:218006cde910da47624b86a9e6f614545e41cc9e98f4e882d590e97e4cbbbdda

Observation fedf7564-1b39-4c14-a9b6-39fad529bd22 · outbound

This paper cites Liquid structural state-space models.

MixFormer: Linear Transformer with Mixture of Memory Experts Liquid structural state-space models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.503407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.804447Z digest=sha256:6cddd8db175a117962a079c14bedc20e30aaaefbf0f7ec7a48b45b45e1c14298

Observation 7484d9bd-91c1-4cbc-97c9-821d3219d5db · outbound

This paper cites Simplified state space layers for sequence modeling.

MixFormer: Linear Transformer with Mixture of Memory Experts Simplified state space layers for sequence modeling

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.488124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.809142Z digest=sha256:e238f13c7485af3eae2d8d43e0074dcb923a3fcb875b080e53602805f3236821

Observation affa02d5-02da-4bd1-9626-0ad06068fbf0 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

MixFormer: Linear Transformer with Mixture of Memory Experts Retentive Network: A Successor to Transformer for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.813888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.813888Z digest=sha256:c8237b8f99b805af0e32852fc1a876ea2d5dd6d5d1d76790bc6385899571ed62

Observation a4d0d1ab-6067-46f6-b07d-9b14f63f7418 · outbound

This paper cites Gated linear attention transformers with hardware-efficient training.

MixFormer: Linear Transformer with Mixture of Memory Experts Gated linear attention transformers with hardware-efficient training

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.472604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.818921Z digest=sha256:053398abce44f28a58c51effaf205f66f8264ee6c14cfe08eff75ddb265ec11e

Observation d97789cb-2202-4d0a-abf9-2067798803ac · outbound

This paper cites Transformers are rnns: fast autoregressive transformers with linear attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Transformers are rnns: fast autoregressive transformers with linear attention

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.456246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.823323Z digest=sha256:59b77cf937479087c2fd9a195a69ab03a41fc29e6541cce3bb58276bfa68dc0a

Observation 3efb5ef3-2a30-4674-972b-c354604964b1 · outbound

This paper cites Efficient attention: attention with linear complexities.

MixFormer: Linear Transformer with Mixture of Memory Experts Efficient attention: attention with linear complexities

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.439782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.827677Z digest=sha256:894cd34aac15effbd1fb6798b4d35f6cb4cceaf5b906bc81c84b3a9faad9e0a2

Observation 7af4b9ed-8273-4831-8def-e00b15592c32 · outbound

This paper cites Linear transformers are secretly fast weight programmers.

MixFormer: Linear Transformer with Mixture of Memory Experts Linear transformers are secretly fast weight programmers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.423943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.832401Z digest=sha256:bc1956c797626c408268e9127864c1cc3a8f1df256503d851fce9c4eca48403f

Observation 349fb060-3e53-4895-95f2-46bc8956a66f · outbound

This paper cites Flatten transformer: vision transformer using focused linear attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Flatten transformer: vision transformer using focused linear attention

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.408559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.837260Z digest=sha256:196e80409b9ee7c998fb1accff51976a4db7d17e7178e478c41b79796857deba

Observation 902f38e1-f551-4875-817b-b88076078383 · outbound

This paper cites Polaformer: polarity-aware linear attention for vision transformers.

MixFormer: Linear Transformer with Mixture of Memory Experts Polaformer: polarity-aware linear attention for vision transformers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.389436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.841842Z digest=sha256:85884318e8f4e3321508ccc469ed8def4c938b9ab109b48266295a73aae02323

Observation e04374ce-f3de-4284-ac12-4545da574cf5 · outbound

This paper cites Linear Attention Mechanism: An Efficient Attention for Semantic Segmentation.

MixFormer: Linear Transformer with Mixture of Memory Experts Linear Attention Mechanism: An Efficient Attention for Semantic Segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.846478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.846478Z digest=sha256:6daeb0fcf57833b87a12b7c14971b1409cc57689744a54739e7e5eefea36f8e3

Observation ec9b2a3f-6171-4bbc-836e-9a14c1f6e354 · outbound

This paper cites Random feature attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Random feature attention

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.370028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.851757Z digest=sha256:56f34ecbe826bf46ce2a90eb66ee1f8f57ec0acd9b516525b6980108eaeb73ab

Observation df083ca3-9ea6-4d5e-8adb-151c05bc334b · outbound

This paper cites Rethinking attention with performers.

MixFormer: Linear Transformer with Mixture of Memory Experts Rethinking attention with performers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.354293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.856505Z digest=sha256:067d9b5277595e8c27c022529cf63c2833668cb9f9c56c2a76c66e97b2dcd527

Observation 207c6656-18d6-4fc9-ab07-ec9c10cf617b · outbound

This paper cites Mamba: linear-time sequence modeling with selective state spaces.

MixFormer: Linear Transformer with Mixture of Memory Experts Mamba: linear-time sequence modeling with selective state spaces

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.337411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.861100Z digest=sha256:2b0b8d5aca01d5f7363feb97dd78e6b824d753569c6efea5793b072d57e14188

Observation 6bcdf29d-0268-42a2-936e-30e01f8fa144 · outbound

This paper cites Transformers are ssms: generalized models and efficient algorithms through structured state space duality.

MixFormer: Linear Transformer with Mixture of Memory Experts Transformers are ssms: generalized models and efficient algorithms through structured state space duality

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.321854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.865999Z digest=sha256:8989f601404d79a375a16736253102f745dada6fb43a57e259170d5f0055bbe9

Observation b3b9a948-da53-456b-918c-639b4a1da352 · outbound

This paper cites An Attention Free Transformer.

MixFormer: Linear Transformer with Mixture of Memory Experts An Attention Free Transformer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.870460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.870460Z digest=sha256:33fc5f508334c5fb26f47899933bb4a255ebcade7260beefd57df16d67877fbb

Observation 366b724a-9899-4afc-89e5-c6030e9a6aa5 · outbound

This paper cites Rwkv: Reinventing rnns for the transformer era.

MixFormer: Linear Transformer with Mixture of Memory Experts Rwkv: Reinventing rnns for the transformer era

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.305242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.875703Z digest=sha256:f061b493d7ab98171c995a44a6ae38dcd6a121348938715b88f16b9f7c26c6ef

Observation e551ac0a-a7a5-42e3-954b-488100443664 · outbound

This paper cites Group normalization.

MixFormer: Linear Transformer with Mixture of Memory Experts Group normalization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.880665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.880665Z digest=sha256:34500b0b8bea90db36d69ecbe9ea87eebbe3ac8c76234bdb11ab3b39d7e38b6f

Observation fded8756-0e96-4f2e-b5e6-fe98c74c77b2 · outbound

This paper cites EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction.

MixFormer: Linear Transformer with Mixture of Memory Experts EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.885235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.885235Z digest=sha256:7317f86af860a7897f1bfb8db956c1fa80b62e1b89833ff48d8fb30670bae7c7

Observation dd78c79f-08f4-42e2-ae0f-67a5679d4a35 · outbound

This paper cites Soft: softmax-free transformer with linear complexity.Proceedings of the 2021 Advances in Neural Information Processing Systems, NeuraIPS, 34:21297–21309, 2021.

MixFormer: Linear Transformer with Mixture of Memory Experts Soft: softmax-free transformer with linear complexity.Proceedings of the 2021 Advances in Neural Information Processing Systems, NeuraIPS, 34:21297–21309, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.274337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.890292Z digest=sha256:b79bd8d43194f00da1ca26a82335c0d21220a81ec77c0a71031ddd7a51138eea

Observation f91b6789-a3d1-4e26-9a39-a2b67165695a · outbound

This paper cites Nystr¨omformer: a nystr ¨om-based algorithm for approximating self-attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Nystr¨omformer: a nystr ¨om-based algorithm for approximating self-attention

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.254217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.895575Z digest=sha256:a6b03d29ff6b1875bef2e7c607921432745098d8a8c399cc8ba9d8fbc8d680db

Observation 594cde54-6c43-4f90-abf2-a79a4cdf328c · outbound

This paper cites Outrageously large neural networks: the sparsely-gated mixture-of-experts layer.

MixFormer: Linear Transformer with Mixture of Memory Experts Outrageously large neural networks: the sparsely-gated mixture-of-experts layer

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.236513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.900238Z digest=sha256:01e8277dc5e5d4badc7c4d9fc5d9980594a8e7948f3acdf75cd32247521a060f

Observation 11156257-0005-4f0a-aff0-29a3e7c52132 · outbound

This paper cites Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization.

MixFormer: Linear Transformer with Mixture of Memory Experts Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.217429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.904806Z digest=sha256:525b88eda42271a2484d0c65d1805120cc0b3a4fa30cb05ed6414dd210476fba

Observation 12c2d22a-8cee-4abb-9529-43cf6d45b9cb · outbound

This paper cites Long range arena: a benchmark for efficient transformers.

MixFormer: Linear Transformer with Mixture of Memory Experts Long range arena: a benchmark for efficient transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.198005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.909593Z digest=sha256:632651aa371c660d455af980dbce165440d14e9110527496915440955123edfc

Observation 985d0c57-9069-4181-bcf6-3803bf82d05e · outbound

This paper cites Listops: a diagnostic dataset for latent tree learning.

MixFormer: Linear Transformer with Mixture of Memory Experts Listops: a diagnostic dataset for latent tree learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.180622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.914154Z digest=sha256:315816dbc2c35882379627648fcb891b86c1a82728b18f7317c5bbed167a9dad

Observation b2114500-42d9-4d0a-8f99-32bfeb3f4060 · outbound

This paper cites Learning word vectors for sentiment analysis.

MixFormer: Linear Transformer with Mixture of Memory Experts Learning word vectors for sentiment analysis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.162071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.918732Z digest=sha256:ae262e466fa1fe9d401976d695f1c587912e39c5f589355860101ec161015a30

Observation e8449253-7c69-42bd-b3e2-7f360bc81a69 · outbound

This paper cites The acl anthology network corpus.Language Resources and Evaluation, 47(4):919–944, 2013.

MixFormer: Linear Transformer with Mixture of Memory Experts The acl anthology network corpus.Language Resources and Evaluation, 47(4):919–944, 2013

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.146128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.923391Z digest=sha256:f6439a611425a7b4d71775f6deb367406c231dc775990c661932a9d1608d2fc7

Observation a66fb356-0427-48a0-ba50-37fff294600b · outbound

This paper cites Learning long-range spatial dependencies with horizontal gated recurrent units.Proceedings of the 2018 Advances in Neural Information Processing Systems, NeuraIPS, 31, 2018.

MixFormer: Linear Transformer with Mixture of Memory Experts Learning long-range spatial dependencies with horizontal gated recurrent units.Proceedings of the 2018 Advances in Neural Information Processing Systems, NeuraIPS, 31, 2018

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.128493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.927739Z digest=sha256:9e3f827de4fe7a65313f11ce35481924c68bb52943d5a2973a3f5edad855ebd9

Observation 867a0f1f-591f-4288-9ec6-cfd59ae19a5d · outbound

This paper cites Learning multiple layers of features from tiny images.

MixFormer: Linear Transformer with Mixture of Memory Experts Learning multiple layers of features from tiny images

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.932524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.932524Z digest=sha256:abcdbbb7d2fc46e90a99d182ffeb3e6b604a7cef83e3e07d11238fd66328d971

Observation 12adb45a-901c-4675-b703-276e1affadcb · outbound

This paper cites Emnist: Extending mnist to handwritten letters.

MixFormer: Linear Transformer with Mixture of Memory Experts Emnist: Extending mnist to handwritten letters

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.101239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:02:06.937137Z digest=sha256:76e37e2af291330c2db4a0a252936502a63a31948491adf5f2f978e3c4970ac4

Pith citing papers

No inbound Pith citation observations are available.