Pith. sign in

Paper Citation Record · LEDGER

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2506.22169.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22169 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:18:02.043090Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy49
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bbecee42-bc13-46f1-92d9-93486ec9e986 · outbound

This paper cites Antman: Dynamic scaling on gpu clusters for deep learning,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Antman: Dynamic scaling on gpu clusters for deep learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:17:57.168913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:17:57.168913Z digest=sha256:1579a69953b66048b688d9fc2d3927fb70a0b70cba5afebdfe1a196219e587f6

Observation 1185e513-39cc-4cef-9cdb-7eaecf1325c5 · outbound

This paper cites Whale: Efficient giant model training over heterogeneous gpus,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Whale: Efficient giant model training over heterogeneous gpus,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:17:57.209280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:17:57.209280Z digest=sha256:dee67186e07002fd7ce0838147ebc035cdb40a71e74112475ac92b6ebbd51c20

Observation c7e8f312-35ea-4203-b8ea-bf9f81a6929d · outbound

This paper cites Mpmoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Mpmoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:09.420985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:57.265068Z digest=sha256:3eae7c819251133944838e1f51afe54cc6c85f6465d91be40eadb93902d0e154

Observation 76aae5cd-6491-4bda-8d2e-1532eca578ce · outbound

This paper cites Redundancy-free high-performance dynamic GNN training with hier- archical pipeline parallelism,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Redundancy-free high-performance dynamic GNN training with hier- archical pipeline parallelism,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:09.191149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:57.366815Z digest=sha256:58c293e14ab7b28cf2777a19e7a43b856b9a756a0c0472e46d883dc88f721d61

Observation 11ae06a4-6f9f-4b71-b3a3-30d625fc11a2 · outbound

This paper cites Chimera: An analytical optimizing framework for effective compute-intensive operators fusion,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Chimera: An analytical optimizing framework for effective compute-intensive operators fusion,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:08.976798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:57.415899Z digest=sha256:7f2b105ce353769a508bacdcef67450a5b02d0200a573cac8b481874da4f16ed

Observation 40d23f88-78ec-45c4-a560-a1071ee9986e · outbound

This paper cites Dnnfusion: accelerating deep neural networks execution with advanced operator fusion,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Dnnfusion: accelerating deep neural networks execution with advanced operator fusion,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:08.820890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:57.516353Z digest=sha256:f96168c1ed5a5cb27241541ec13a3865514920dddd732116b52cd93255d3a3c5

Observation e259b4ca-baa6-4ae6-8110-70c333860490 · outbound

This paper cites Astitch: enabling a new multi- dimensional optimization space for memory-intensive ML training and inference on modern SIMT architectures,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Astitch: enabling a new multi- dimensional optimization space for memory-intensive ML training and inference on modern SIMT architectures,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:08.657288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:57.599031Z digest=sha256:ecef3e574c5f0323c1c89f9c535a4bb6c1fe1a2efb4c349e275ab7783c9cd064

Observation 89b07b5c-2f64-422d-8bca-4e012988ba5c · outbound

This paper cites Nvidia cublas.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Nvidia cublas

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:08.480013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:57.686091Z digest=sha256:252648b35f10a285e94adcfc6b0d228c970032d5216dade28037f547862e8485

Observation a7463e97-cdba-46a5-8226-b6c0b1eb8fbb · outbound

This paper cites Nvidia cudnn.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Nvidia cudnn

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:08.320639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:57.794966Z digest=sha256:a9002a5d0d122a349ee4efc0334a24b6affb6aefdd40d7546b57949d368edac9

Observation 854fd778-21d7-419f-9b60-aab00d04dcfb · outbound

This paper cites Nvidia cutlass.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Nvidia cutlass

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:08.146636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:57.833125Z digest=sha256:1cede1ac680e192d8ac231f9f66373eea887d24c529cd32c786d68ef5ed0cd2c

Observation ede060d9-800d-4524-8bb7-0330a19fc813 · outbound

This paper cites Nvidia tensorrt.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Nvidia tensorrt

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:07.966648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:57.951859Z digest=sha256:4ce6eed6b508c253392de7874709d3cebf34ad7e2a2f6941fc2b6c7bb83c73c0

Observation 42a04fc8-f71d-47df-9713-a4516b44d7fa · outbound

This paper cites Mpipemoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Mpipemoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:07.873222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:58.032881Z digest=sha256:85a8c09a75494a23fae3df7447c5244b19c986c91159f4e6c1f9e86b74d5a37d

Observation 3d5b6d4f-535c-4e2b-b2b8-d076fe0ae1d1 · outbound

This paper cites Ansor: Generating high-performance tensor programs for deep learning,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Ansor: Generating high-performance tensor programs for deep learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:07.739342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:58.151591Z digest=sha256:d9a28228d4cf32dbda1505e018992218168e743a82e39a6f014b4b4ff6e77be8

Observation 91b5adab-a47a-4d82-8ab8-b8ac17f1de39 · outbound

This paper cites Learning to optimize tensor programs,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Learning to optimize tensor programs,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:07.636632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:58.210833Z digest=sha256:82515e2db2f666201aae1393622aaff07bb14139e9b80e851e876f5070478b25

Observation db245a2a-f1fb-427a-92c9-9fe5077cb374 · outbound

This paper cites Tiramisu: A polyhedral compiler for expressing fast and portable code,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Tiramisu: A polyhedral compiler for expressing fast and portable code,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:07.509586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:58.354573Z digest=sha256:3acd082eed9070a2e967f94363b1a6a83dcee1b9e1eb65eb7796afa004f6f967

Observation 14913f6e-2781-4798-af21-dc23ad56e7cf · outbound

This paper cites Astra: Exploiting predictability to optimize deep learning,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Astra: Exploiting predictability to optimize deep learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:07.418515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:58.407201Z digest=sha256:ac13658fb1ab8eb6b78d2d744b9b0a843f212cf28293ae660f67adfbab267788

Observation c94434b6-7232-405f-baa9-59cba381aaac · outbound

This paper cites TASO: optimizing deep learning computation with automatic genera- tion of graph substitutions,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators TASO: optimizing deep learning computation with automatic genera- tion of graph substitutions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:07.317022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:58.533271Z digest=sha256:7aedafaf55be532654ea61ffbc1dbdc0844c7ae436014b1d8f65c497aa6b9a6f

Observation f5de12cf-1279-4361-b77d-755584ea987a · outbound

This paper cites Scalable kernel fusion for memory-bound GPU applications,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Scalable kernel fusion for memory-bound GPU applications,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:07.197903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:58.593356Z digest=sha256:61df505a9a07ce77acf5e4269673a56f4c7c5867644c223ece56750c179f5022

Observation d9eef3ca-6525-4084-b69c-59a27eeebd11 · outbound

This paper cites AKG: automatic kernel generation for neural processing units using polyhedral transformations,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators AKG: automatic kernel generation for neural processing units using polyhedral transformations,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:07.097573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:58.722719Z digest=sha256:c7ae2401660ea772af43ab8ffee5fe5bd2304cc5b01abc6f120c37e3ce8604db

Observation 38be9a61-80d0-4911-a323-514c7c633aba · outbound

This paper cites Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:17:58.793850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:17:58.793850Z digest=sha256:241349ca1416d0bdadc2e90ad31ba50c05c35bc7f25da7aa5e1dec7e3cc1767e

Observation 457af5e1-b481-43ce-89ec-6851188d1565 · outbound

This paper cites Tensorflow xla.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Tensorflow xla

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:06.981353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:58.865681Z digest=sha256:c98676fba2d119e822e6c34a7a363499e0dbbb32827cdd8704ee20a3f89d9ba2

Observation 7afa9ddf-2509-4e84-b5ad-b541e46d58f9 · outbound

This paper cites Tvm: An automated end-to-end optimizing compiler for deep learning,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Tvm: An automated end-to-end optimizing compiler for deep learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:06.886879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:58.918888Z digest=sha256:c0434669fb1750d7fd00cf7565e6ddd98c09bbffce800b7a1755486df3bb6fee

Observation 5424f76d-c5aa-42d5-b1b3-97d91c422677 · outbound

This paper cites Attention is all you need,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Attention is all you need,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:06.779195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:58.968018Z digest=sha256:3efe9c1dc49a40fd5548cd79619dfec4c91dd8e5bc328634651985dbc07f8737

Observation 84e453b0-3c1d-4654-aae4-9982a97cf419 · outbound

This paper cites Xgboost: A scalable tree boosting system,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Xgboost: A scalable tree boosting system,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:17:59.047371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:17:59.047371Z digest=sha256:2e712e770c909ffca557f297572602da982d0fe8e8a81e23bf331bd676eb793e

Observation 9e7916d3-35bf-449d-b27a-866054634b0a · outbound

This paper cites Bolt: Bridg- ing the gap between auto-tuners and hardware-native performance,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Bolt: Bridg- ing the gap between auto-tuners and hardware-native performance,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:06.663284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:59.129464Z digest=sha256:284cadee9a86ad75bfcccf42be952e330c084c027e2e544caadf7bf1ad2d424c

Observation 050e6687-95bb-44a7-9320-6e67ea74da55 · outbound

This paper cites Pytorch: An imperative style, high- performance deep learning library,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Pytorch: An imperative style, high- performance deep learning library,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:06.542281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:59.215230Z digest=sha256:e235181879df7d2e4246e39b8bf78b85b243459b0eefc5c0696a8b8ede4a1496

Observation c85a5a73-1353-4154-a185-c9d99004c0dc · outbound

This paper cites Bytetransformer: A high-performance transformer boosted for variable-length inputs,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Bytetransformer: A high-performance transformer boosted for variable-length inputs,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:06.399045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:59.305175Z digest=sha256:a563d4d308a32652fbd617a4e30b8fc5fc18a1a7076558d5ed61ca5cdd48f9f7

Observation 52519be3-2c40-4494-9246-47f7201e5034 · outbound

This paper cites Raptor-t: A fused and memory-efficient sparse transformer for long and variable-length sequences.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Raptor-t: A fused and memory-efficient sparse transformer for long and variable-length sequences

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:06.244303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:59.403654Z digest=sha256:42de6559d2eff85c18ff3c1e82c4dbfdef363cd593b51b1b4c93a7ac9e806a7f

Observation a8893b13-17ce-4513-b995-54b4b8c4b4f1 · outbound

This paper cites FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:17:59.486067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:17:59.486067Z digest=sha256:7575678dace60e7e8bacdf5ece6e9d08415e8dcc74d44dd7814bcf1ec0b79f35

Observation 80b96ac1-79d9-4651-aad5-108bd532935a · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:06.091515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:59.583398Z digest=sha256:a816e5b0fcd15a9e4b4bd3726993e1a31e5f584c380062b280a2950d6b7e57ff

Observation 9101e053-b75b-409a-a3e4-5492ca3005d7 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:17:59.630141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:17:59.630141Z digest=sha256:63adeeda23803cb33bdcc03c126dea01c25063fc4696456e6dfaab2be9022f89

Observation d6c552ae-5e74-433f-915f-bd4584307700 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:17:59.732866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:17:59.732866Z digest=sha256:96101d0eee03b1544ebc9bb3019f57de3cf30ad929df11dd68f47a67d6d5a325

Observation 7fb5eb6b-639c-4275-b1a8-fa9b663f7e1c · outbound

This paper cites Hidet: Task-mapping programming paradigm for deep learning tensor programs,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Hidet: Task-mapping programming paradigm for deep learning tensor programs,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:05.913727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:59.806635Z digest=sha256:dfa36e58027e211efb570a8a1e4202d7e755f2f77cf36c6a39193fcd300c6f8a

Observation 130c99b9-848a-486b-85f1-cd4af975f668 · outbound

This paper cites Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:05.733116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:17:59.979562Z digest=sha256:74d7718ceed9a80e5f3658aa14309a679036282324a8ce8fe69cd1cd8086d78b

Observation dddcb304-b24e-4fa5-9785-ddf91d1a97cb · outbound

This paper cites Flextensor: An automatic schedule exploration and optimization framework for tensor computation on heterogeneous system,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Flextensor: An automatic schedule exploration and optimization framework for tensor computation on heterogeneous system,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:05.542898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:00.054589Z digest=sha256:48c342d46135c2e9884fc3869430d9b1980de1d678b14afbc63858e747830fff

Observation ec7a60cc-8591-444a-9d5a-12055042f1e7 · outbound

This paper cites AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstraction,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstraction,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:05.391723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:00.135867Z digest=sha256:78d2f5bea325499ef7db2fbe7e9dd51517884eeb232085dcaa1d565328b932ee

Observation 30278ef8-a73e-4540-81af-cc31728684d1 · outbound

This paper cites Atomic dataflow based graph-level workload orchestration for scalable DNN accelerators,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Atomic dataflow based graph-level workload orchestration for scalable DNN accelerators,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:05.211211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:00.213362Z digest=sha256:d657f0ecd8fc823d684086653f8faa2915af287ff68065e9067d84b5efe2009d

Observation 27a2a5c7-9b64-443a-89ee-aabfae0d6714 · outbound

This paper cites ROLLER: fast and efficient tensor compilation for deep learning,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators ROLLER: fast and efficient tensor compilation for deep learning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:05.031888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:00.316360Z digest=sha256:ab23dcafb2fdd3ded21a334af87af8b91bd7ddcb94f23436813e5222311f37ca

Observation 2fdbc928-81b9-4ac9-9368-a2d7a8301d09 · outbound

This paper cites Triton: an intermediate language and compiler for tiled neural network computations,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Triton: an intermediate language and compiler for tiled neural network computations,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:04.848685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:00.357498Z digest=sha256:39b78a86dac1d0b800b4d6b16acc560f6572ee8ca332b17331a2d8a4a02b3fa5

Observation 08894df8-1172-4a27-b829-9d08d8d64032 · outbound

This paper cites Nvidia, parallel thread execution isa.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Nvidia, parallel thread execution isa

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:04.710568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:00.444478Z digest=sha256:02608c79e65c975678c65b7f0690d27497343db837909dc76930410e1bfb04d8

Observation 62cbf713-64c8-4336-8933-2d8a24b8989a · outbound

This paper cites Tensorflow: A system for large-scale machine learning,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Tensorflow: A system for large-scale machine learning,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:04.524401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:00.502303Z digest=sha256:97550e41738816bec3f447ed8bb4251b2f8df64f5d57392cad60020eee7e3795

Observation c63e8000-3b1c-4edf-a171-006af5dcddbb · outbound

This paper cites Available: https://onnx.ai/.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Available: https://onnx.ai/

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:04.413135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:00.572160Z digest=sha256:94ed5d2e57389f6e7894fb69b10836b70c12cbe99eab9d75ee8af0354d89559b

Observation 39f3acaf-587b-43bb-8c80-425fa8f42d4d · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:00.661573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:00.661573Z digest=sha256:5cd11432f4e3da601c03813a33aba439a049f16616fac30ded3e894bcde0ee7a

Observation efe2bd7b-2c42-4c2c-9ac6-138960b9a649 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators An image is worth 16x16 words: Trans- formers for image recognition at scale,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:04.281331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:00.750686Z digest=sha256:f7db92b4d57faa61b12b4c991079045dc5f3b5e643da0b1dbc638809722ef5ea

Observation 79551829-ecc3-4148-b775-011159bdbf85 · outbound

This paper cites Mlp-mixer: An all-mlp architecture for vision,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Mlp-mixer: An all-mlp architecture for vision,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:04.147790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:00.840062Z digest=sha256:4f2eb6b4882fcc7c326db33c578236ef55033731bdc5110b78bcac6921fb71e9

Observation 7c05d49d-2121-4510-ae70-e5ae4570359f · outbound

This paper cites Relay: A High-Level Compiler for Deep Learning.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Relay: A High-Level Compiler for Deep Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:00.935149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:00.935149Z digest=sha256:7575ada18d964f5bc165a796d570332a8f3e65873cc33558f4257a9ac582f09e

Observation 0e018dfd-7f5f-40ca-bc6f-70790c7bd1ce · outbound

This paper cites Intel oneapi deep neural network library.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Intel oneapi deep neural network library

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:03.968816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:01.019537Z digest=sha256:dbab04c34416e6608b40082c9e5653395aeacd9740765ec247759d740d99aad0

Observation 77abbbf7-6fe3-4324-be58-7417b2784e14 · outbound

This paper cites Intel oneapi math kernel library.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Intel oneapi math kernel library

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:03.782986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:01.121244Z digest=sha256:1a77d8751f39e3be7bf5b139eaac10cdd6a1f7371ae9c03413c457f894e200a5

Observation 39b837f5-2f97-48a1-b8fd-e28022d2d695 · outbound

This paper cites Rammer: Enabling holistic deep learning compiler optimizations with rtasks,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Rammer: Enabling holistic deep learning compiler optimizations with rtasks,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:03.657856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:01.198588Z digest=sha256:ddc08b30f2d3435a80ae40a710ae9bf4c4bed0134b2ca5ca895790ce95dde9f5

Observation 972a4a4e-d287-4515-8b4d-97565b441020 · outbound

This paper cites Accelerating deep learning inference with cross-layer data reuse on gpus,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Accelerating deep learning inference with cross-layer data reuse on gpus,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:03.481663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:01.301835Z digest=sha256:c144f037cd8d394f38504eb3702ba4da647ba6a1ba59b4d4333e5df863112788

Observation 27e8c3fb-6275-463b-aad3-ba5f8a7e2c85 · outbound

This paper cites Efficient GPU spatial-temporal multitasking,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Efficient GPU spatial-temporal multitasking,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:03.325638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:01.414094Z digest=sha256:b9c123f80156c43e6af10a15b18821e63d93c0462b1df5adb55970dea87da364

Observation 80618246-eca4-47aa-8edc-703189e895f5 · outbound

This paper cites On optimizing machine learning workloads via kernel fusion,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators On optimizing machine learning workloads via kernel fusion,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:03.186137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:01.537926Z digest=sha256:af27035959a50a47f4989165973fe92fd70c4b7662fb194fb45f3afc8113d877

Observation fcdd5e74-8860-4a8e-a15c-3728e6ee456e · outbound

This paper cites Learning to optimize halide with tree search and random programs,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Learning to optimize halide with tree search and random programs,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:03.032273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:01.643595Z digest=sha256:786b9b674f3998b2fe6f9c1a8f7ef6aa4ff59bb8edd589788b8caa3bccedd1da

Observation 3acc49c2-e698-4f4d-a533-d67a3ee08957 · outbound

This paper cites Nimble: Efficiently compiling dynamic neural networks for model inference,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Nimble: Efficiently compiling dynamic neural networks for model inference,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:02.874308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:01.732170Z digest=sha256:f4755995f37e82cf231426e84fd478a6c28842894d84529e9ccd96a8537c5766

Observation 40723419-b400-4c46-8dad-ebf01ffe234b · outbound

This paper cites DISC: A dynamic shape compiler for machine learning workloads,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators DISC: A dynamic shape compiler for machine learning workloads,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:02.743768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:01.793959Z digest=sha256:9c563bbb436f18f217034d0732479b95850453cc1950ef397d8b92a975234dac

Observation f6faa016-543a-48fc-a467-ed5720274a01 · outbound

This paper cites Automatic generation of high-performance quantized machine learning kernels,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Automatic generation of high-performance quantized machine learning kernels,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:02.607381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:01.864329Z digest=sha256:8bcb6b30ad19c59e09626afae73c0aed98863bc985b94310c177f610ba84ee2f

Observation 5827d460-ad2e-4dfb-8d17-8b7fb5245746 · outbound

This paper cites A code generator for high-performance tensor contractions on gpus,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators A code generator for high-performance tensor contractions on gpus,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:02.442784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:01.936505Z digest=sha256:1d76cc24ed7ffa0bd8ce074522d4611dab31b668fb1e8693a58b552a5bb36dc0

Observation 671f0ec3-4e56-4dc6-9149-51edd5237477 · outbound

This paper cites Neoflow: A flexible framework for enabling efficient compilation for high performance DNN training,.

MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators Neoflow: A flexible framework for enabling efficient compilation for high performance DNN training,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:18:02.297777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:18:02.043090Z digest=sha256:d75594349351812261b49a7530445df9d80da9995acfa0bc7e342c56bbb787cd

Pith citing papers

No inbound Pith citation observations are available.