Pith. sign in

Paper Citation Record · LEDGER

MixFormer: Linear Transformer with Mixture of Memory Experts

As of 12 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2608.09468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09468 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:02:06.937137Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bd4a140-c4f3-4aca-8f13-fe8cf5344647 · outbound

This paper cites Attention is all you need.Proceedings of the 2017 Advances in Neural Information Processing Systems, NeuraIPS, pages 5998–6008, 2017.

MixFormer: Linear Transformer with Mixture of Memory Experts Attention is all you need.Proceedings of the 2017 Advances in Neural Information Processing Systems, NeuraIPS, pages 5998–6008, 2017

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.657671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.733397Z digest=sha256:1e5cbeca2fa849408f8f75b23fc9cdc5bbe527f93abb4c7c5abed504184a2999

Observation 2a19aaee-0952-4df1-9c0a-856f4aa08113 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MixFormer: Linear Transformer with Mixture of Memory Experts LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.739145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.739145Z digest=sha256:fd52a6a51db82f6e1b11d437eeeea6c10f8e1d43589fcbdbb3ff470a44345432

Observation f1eba3b0-f82b-450e-84bb-51e0e2505169 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MixFormer: Linear Transformer with Mixture of Memory Experts Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.744470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.744470Z digest=sha256:99b1ad9babadda8c0dea1300278b4bec4c4f27c61c70b3746e214bc8f407cf47

Observation 0fb668da-41b8-48ee-9f55-4115ea066474 · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407, 2024.

MixFormer: Linear Transformer with Mixture of Memory Experts The llama 3 herd of models.arXiv e-prints, pages arXiv–2407, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.749457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.749457Z digest=sha256:3add7a7d734e1f659d7b0c91e4e77dd5fa4150e0fd982e146bb828e03b14bfbc

Observation 77a00a8d-2fa9-42c1-81bf-23e2aa35a402 · outbound

This paper cites Molmo and pixmo: open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024.

MixFormer: Linear Transformer with Mixture of Memory Experts Molmo and pixmo: open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.630860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.755293Z digest=sha256:de4a18e9b4e4aeed979d6b1b9f854ebab4ddf3a848418d601ae26a263a9ffb99

Observation 6a6a3153-d04c-4a01-aef2-f61433ec6a6b · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

MixFormer: Linear Transformer with Mixture of Memory Experts NVLM: Open Frontier-Class Multimodal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.760415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.760415Z digest=sha256:4b87b33090c174a5d047a32b4d16119e9cd7e40c708c31dd57258809ea035f56

Observation c4a816df-d03b-4ba8-8028-3b5009247283 · outbound

This paper cites Efficient transformers: a survey.ACM Computing Survey, 55(6), 2022.

MixFormer: Linear Transformer with Mixture of Memory Experts Efficient transformers: a survey.ACM Computing Survey, 55(6), 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.615287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.767004Z digest=sha256:89c1ba5e33cad35f29d582a8c7d590426ebe6c093dd3413d4de9433108eaccc9

Observation 2f607eb3-f53f-4eee-b132-8c578e47b8de · outbound

This paper cites A survey on efficient training of transformers.

MixFormer: Linear Transformer with Mixture of Memory Experts A survey on efficient training of transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.598109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.772520Z digest=sha256:fcecc91d44594860c2c0ae79da038329a4b2b9bc728c371f7f7b8307cff9eb97

Observation e1c29fcc-6bc8-4563-a0c0-8ef0102c0970 · outbound

This paper cites an unresolved cited work.

MixFormer: Linear Transformer with Mixture of Memory Experts Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:02:07.581779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.777736Z digest=sha256:1879ef2bd7b5d70b6bd10a51f991c003a73c2cfe713565a04add998aea82060e

Observation 1eb31702-f9b0-4ebf-b0f6-2acd135bf81d · outbound

This paper cites A survey on vision transformer.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1):87–110, 2022.

MixFormer: Linear Transformer with Mixture of Memory Experts A survey on vision transformer.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1):87–110, 2022

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.566504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.783024Z digest=sha256:1ae957644e46e54ca87c51b289816a87c975772ca57a305e102847a87ad60d8e

Observation 1df62038-dde7-43b7-b1fb-94f733cdd8a7 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

MixFormer: Linear Transformer with Mixture of Memory Experts Efficiently modeling long sequences with structured state spaces

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.551246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.789114Z digest=sha256:64a9d71b189a54908b6d4552a2c76caf3eb7c46f7fba40273e84d368e05492db

Observation dfa9049a-e736-42ae-b409-4e28161aa6c2 · outbound

This paper cites Diagonal state spaces are as effective as structured state spaces.

MixFormer: Linear Transformer with Mixture of Memory Experts Diagonal state spaces are as effective as structured state spaces

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.534221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.794479Z digest=sha256:9edb5afbbe2ff8ac5d5dbde70860f602608075733c703630287e6608115f7873

Observation 5a49165a-d784-4664-b347-f09b3875177e · outbound

This paper cites an unresolved cited work.

MixFormer: Linear Transformer with Mixture of Memory Experts Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:02:07.519106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.799632Z digest=sha256:fcdfe54703e02c7d733b337d629e8531b7e2ee352146eebec30ce69bb8da35b9

Observation fedf7564-1b39-4c14-a9b6-39fad529bd22 · outbound

This paper cites Liquid structural state-space models.

MixFormer: Linear Transformer with Mixture of Memory Experts Liquid structural state-space models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.503407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.804447Z digest=sha256:937a0ac8dd3a9a5673d458ae48b7ea02c2ac7febbbafddfd6f9009193c249d26

Observation 7484d9bd-91c1-4cbc-97c9-821d3219d5db · outbound

This paper cites Simplified state space layers for sequence modeling.

MixFormer: Linear Transformer with Mixture of Memory Experts Simplified state space layers for sequence modeling

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.488124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.809142Z digest=sha256:ba94523be8348a7272742313e7dbe520e6bf89b0090b52848fe56e9f64aab354

Observation affa02d5-02da-4bd1-9626-0ad06068fbf0 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

MixFormer: Linear Transformer with Mixture of Memory Experts Retentive Network: A Successor to Transformer for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.813888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.813888Z digest=sha256:fca0fa696a9d29233bbf44fed9c80b35d0e4ce06f24be56999aff70f6ffe0ea1

Observation a4d0d1ab-6067-46f6-b07d-9b14f63f7418 · outbound

This paper cites Gated linear attention transformers with hardware-efficient training.

MixFormer: Linear Transformer with Mixture of Memory Experts Gated linear attention transformers with hardware-efficient training

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.472604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.818921Z digest=sha256:abed7468820351ec42b38bd92c8f6addc3a131692b7cd1fb380e6d0a60ea3c8c

Observation d97789cb-2202-4d0a-abf9-2067798803ac · outbound

This paper cites Transformers are rnns: fast autoregressive transformers with linear attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Transformers are rnns: fast autoregressive transformers with linear attention

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.456246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.823323Z digest=sha256:c71fed2af7300a8c6a81f3e156c0894234c2f195f2bc0ad0efd5b6064df0a572

Observation 3efb5ef3-2a30-4674-972b-c354604964b1 · outbound

This paper cites Efficient attention: attention with linear complexities.

MixFormer: Linear Transformer with Mixture of Memory Experts Efficient attention: attention with linear complexities

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.439782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.827677Z digest=sha256:ec0ccc83993bbe5f70fa8938f2e71e857eea3ab5f3d8a3dba90a6ac41ba78074

Observation 7af4b9ed-8273-4831-8def-e00b15592c32 · outbound

This paper cites Linear transformers are secretly fast weight programmers.

MixFormer: Linear Transformer with Mixture of Memory Experts Linear transformers are secretly fast weight programmers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.423943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.832401Z digest=sha256:d17eb22e21bf69cd3ebbf6e3805195e9559e73f3c8f67bd63534dd0e238defb2

Observation 349fb060-3e53-4895-95f2-46bc8956a66f · outbound

This paper cites Flatten transformer: vision transformer using focused linear attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Flatten transformer: vision transformer using focused linear attention

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.408559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.837260Z digest=sha256:3cd7570e9e729bcfd306b35d2b8a02b2061a7e2f73af155dcff94f3b84bf2ed2

Observation 902f38e1-f551-4875-817b-b88076078383 · outbound

This paper cites Polaformer: polarity-aware linear attention for vision transformers.

MixFormer: Linear Transformer with Mixture of Memory Experts Polaformer: polarity-aware linear attention for vision transformers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.389436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.841842Z digest=sha256:4dd701faece4537399fdd7908710374996b4aa1df5bf5488d43fa065c86ee368

Observation e04374ce-f3de-4284-ac12-4545da574cf5 · outbound

This paper cites Linear Attention Mechanism: An Efficient Attention for Semantic Segmentation.

MixFormer: Linear Transformer with Mixture of Memory Experts Linear Attention Mechanism: An Efficient Attention for Semantic Segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.846478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.846478Z digest=sha256:c4e7458bdf146487f7019bbf786319ec1ea9311e5657bd83a24164ef1a4bd57b

Observation ec9b2a3f-6171-4bbc-836e-9a14c1f6e354 · outbound

This paper cites Random feature attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Random feature attention

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.370028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.851757Z digest=sha256:63e3d55c3c7764ff6ffeaff7b33dbc892c1fdbd25ad75b7d263720dac82c4473

Observation df083ca3-9ea6-4d5e-8adb-151c05bc334b · outbound

This paper cites Rethinking attention with performers.

MixFormer: Linear Transformer with Mixture of Memory Experts Rethinking attention with performers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.354293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.856505Z digest=sha256:a1e557e7c8016e5dc5849c7e3d564f956d5530d0d2d4b8b005bd36dfd19414b3

Observation 207c6656-18d6-4fc9-ab07-ec9c10cf617b · outbound

This paper cites Mamba: linear-time sequence modeling with selective state spaces.

MixFormer: Linear Transformer with Mixture of Memory Experts Mamba: linear-time sequence modeling with selective state spaces

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.337411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.861100Z digest=sha256:0459d9c5d0a78153f7a5bc4fe32532f3ba262e8879976c426c9661204031da7c

Observation 6bcdf29d-0268-42a2-936e-30e01f8fa144 · outbound

This paper cites Transformers are ssms: generalized models and efficient algorithms through structured state space duality.

MixFormer: Linear Transformer with Mixture of Memory Experts Transformers are ssms: generalized models and efficient algorithms through structured state space duality

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.321854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.865999Z digest=sha256:65fb2ba54c7980854f2533f4511dc15b66e918567735c4b7b6b5be3abc45c90b

Observation b3b9a948-da53-456b-918c-639b4a1da352 · outbound

This paper cites An Attention Free Transformer.

MixFormer: Linear Transformer with Mixture of Memory Experts An Attention Free Transformer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.870460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.870460Z digest=sha256:4c483d0b69533e5ef93e2024849ba2f2a744615d3bd677d87c23ec36ea120258

Observation 366b724a-9899-4afc-89e5-c6030e9a6aa5 · outbound

This paper cites Rwkv: Reinventing rnns for the transformer era.

MixFormer: Linear Transformer with Mixture of Memory Experts Rwkv: Reinventing rnns for the transformer era

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.305242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.875703Z digest=sha256:9bb80abb403d671402e913b26690b8b542ac71b705fa8366566cc62ef9e9af6c

Observation e551ac0a-a7a5-42e3-954b-488100443664 · outbound

This paper cites Group normalization.

MixFormer: Linear Transformer with Mixture of Memory Experts Group normalization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.880665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.880665Z digest=sha256:83cb80e2a4904a8b5a0372b9761c39055caaf24f3f28bcdfb101bc5f6aaa019d

Observation fded8756-0e96-4f2e-b5e6-fe98c74c77b2 · outbound

This paper cites EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction.

MixFormer: Linear Transformer with Mixture of Memory Experts EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.885235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.885235Z digest=sha256:0c3568fd2ce4caea0e103e38e2256e9e896b2409d8c3229e48fff78b0edf28f4

Observation dd78c79f-08f4-42e2-ae0f-67a5679d4a35 · outbound

This paper cites Soft: softmax-free transformer with linear complexity.Proceedings of the 2021 Advances in Neural Information Processing Systems, NeuraIPS, 34:21297–21309, 2021.

MixFormer: Linear Transformer with Mixture of Memory Experts Soft: softmax-free transformer with linear complexity.Proceedings of the 2021 Advances in Neural Information Processing Systems, NeuraIPS, 34:21297–21309, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.274337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.890292Z digest=sha256:ed70e20cc2479ee2abf87722bc9ed271578da8aa208e8e99e42b540f55919a1d

Observation f91b6789-a3d1-4e26-9a39-a2b67165695a · outbound

This paper cites Nystr¨omformer: a nystr ¨om-based algorithm for approximating self-attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Nystr¨omformer: a nystr ¨om-based algorithm for approximating self-attention

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.254217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.895575Z digest=sha256:79053a74aa8312a089e07a3c1dbee7de8f075fc7c4d603e4de97cd1fffa3d355

Observation 594cde54-6c43-4f90-abf2-a79a4cdf328c · outbound

This paper cites Outrageously large neural networks: the sparsely-gated mixture-of-experts layer.

MixFormer: Linear Transformer with Mixture of Memory Experts Outrageously large neural networks: the sparsely-gated mixture-of-experts layer

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.236513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.900238Z digest=sha256:2c5381251026f28bf1cdc0ae05031040d93687af20954055be1655aa8bd3ee07

Observation 11156257-0005-4f0a-aff0-29a3e7c52132 · outbound

This paper cites Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization.

MixFormer: Linear Transformer with Mixture of Memory Experts Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.217429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.904806Z digest=sha256:37986536c727bf5d32c5678f93bcd008e6e481f9e1cb2d04d2df1597ef343e17

Observation 12c2d22a-8cee-4abb-9529-43cf6d45b9cb · outbound

This paper cites Long range arena: a benchmark for efficient transformers.

MixFormer: Linear Transformer with Mixture of Memory Experts Long range arena: a benchmark for efficient transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.198005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.909593Z digest=sha256:98c30bf3569b021109ef21d98052e3ff5828e47acf462e2e207e73b8c862afa8

Observation 985d0c57-9069-4181-bcf6-3803bf82d05e · outbound

This paper cites Listops: a diagnostic dataset for latent tree learning.

MixFormer: Linear Transformer with Mixture of Memory Experts Listops: a diagnostic dataset for latent tree learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.180622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.914154Z digest=sha256:e3c188a46c9bb6cb93076f2f7ae7942b91aaabe5cab32a21c580bd12679de922

Observation b2114500-42d9-4d0a-8f99-32bfeb3f4060 · outbound

This paper cites Learning word vectors for sentiment analysis.

MixFormer: Linear Transformer with Mixture of Memory Experts Learning word vectors for sentiment analysis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.162071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.918732Z digest=sha256:cd73565023ad79c9ebac6692cfef44463e24ba49d64549ded28403bba11d3548

Observation e8449253-7c69-42bd-b3e2-7f360bc81a69 · outbound

This paper cites The acl anthology network corpus.Language Resources and Evaluation, 47(4):919–944, 2013.

MixFormer: Linear Transformer with Mixture of Memory Experts The acl anthology network corpus.Language Resources and Evaluation, 47(4):919–944, 2013

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.146128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.923391Z digest=sha256:ad6921b872b390ef300e574770d3671064c03e4c9d5c7fe71a3e829ec449dc44

Observation a66fb356-0427-48a0-ba50-37fff294600b · outbound

This paper cites Learning long-range spatial dependencies with horizontal gated recurrent units.Proceedings of the 2018 Advances in Neural Information Processing Systems, NeuraIPS, 31, 2018.

MixFormer: Linear Transformer with Mixture of Memory Experts Learning long-range spatial dependencies with horizontal gated recurrent units.Proceedings of the 2018 Advances in Neural Information Processing Systems, NeuraIPS, 31, 2018

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.128493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.927739Z digest=sha256:88e7e60b014835b8cee996b68d9a29024283418e50e5eeb071dfbb8c8591cc2c

Observation 867a0f1f-591f-4288-9ec6-cfd59ae19a5d · outbound

This paper cites Learning multiple layers of features from tiny images.

MixFormer: Linear Transformer with Mixture of Memory Experts Learning multiple layers of features from tiny images

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.932524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.932524Z digest=sha256:56d69d74383520529395843b16ec7b2d963c355b06bafda5df1186b2fb55109c

Observation 12adb45a-901c-4675-b703-276e1affadcb · outbound

This paper cites Emnist: Extending mnist to handwritten letters.

MixFormer: Linear Transformer with Mixture of Memory Experts Emnist: Extending mnist to handwritten letters

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.101239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:02:06.937137Z digest=sha256:243b4bd3b77bf376f8e95267e9725da5636ce47506c0478f6204bb18c1a1660b

Pith citing papers

No inbound Pith citation observations are available.