Pith. sign in

Paper Citation Record · LEDGER

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures

As of 10 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2608.03537.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03537 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:14:18.740009Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de83941b-3105-4d33-8c4f-3e8f4b789b38 · outbound

This paper cites TensorFlow: A system for Large-Scale machine learning,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures TensorFlow: A system for Large-Scale machine learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.269781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:15.150954Z digest=sha256:552c08e8b62e40355b2c97a53b9391a68d46caf9ba5332db1095dc21bca475f3

Observation f8a1a920-6bfd-4beb-acb7-5ba9610ede9a · outbound

This paper cites Learning to op- timize halide with tree search and random programs,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Learning to op- timize halide with tree search and random programs,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.260061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:15.176970Z digest=sha256:fe1ad10f02e5fa1e43c60a0a00b2b010863ab73b175691dd8daeda1e0da5794b

Observation dbbd22f3-545c-4526-8c50-63f458e096f7 · outbound

This paper cites Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:15.249607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:15.249607Z digest=sha256:c0c04e838d7fe89babb17d5077e4e86a0cb7e3ad6583ac37b044826c4e63735a

Observation 8a731703-094b-4bfc-b007-dbb1f94e2176 · outbound

This paper cites TVM: An automated end-to-end optimizing compiler for deep learning,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures TVM: An automated end-to-end optimizing compiler for deep learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.245075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:15.316409Z digest=sha256:0d366a506d32301495bd5199e865acde70ee92dcd2d4b8b00aea47cdbd83c716

Observation 58231797-7d19-4de7-a11b-dcebafd6cc81 · outbound

This paper cites Learning to optimize tensor programs,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Learning to optimize tensor programs,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.235551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:15.395958Z digest=sha256:aa8fa73af21c785fffa88b460d110db1321574a6e9c3a46649b70725069e4451

Observation d1a62d32-53c3-4fb2-b398-29dad4f513d3 · outbound

This paper cites Evt: Accelerating deep learning training with epilogue visitor tree,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Evt: Accelerating deep learning training with epilogue visitor tree,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.226744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:15.470126Z digest=sha256:cac53a96c5602b42a29600e720e9f9e8b8e4f17692c2d622d7c1610a50cc51d4

Observation 95ca7f43-86fd-4cc4-ac48-9c0896e07727 · outbound

This paper cites cuDNN: Efficient Primitives for Deep Learning.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures cuDNN: Efficient Primitives for Deep Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:15.535544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:15.535544Z digest=sha256:6a757d2ba0c9925c2dfd1b94d03830f27439b714355e48b7f9cb89f8b6ed33bd

Observation 51f1f166-1332-4459-969e-2676d49c715e · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Flashattention-2: Faster attention with better parallelism and work partitioning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.217549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:15.581087Z digest=sha256:8938c8d1250d1180bd9e6ced2330529a63fe6516a3c4fbb5d73cffba213672ab

Observation 6af0f6c4-efec-4282-818d-dde8d198105e · outbound

This paper cites Analyzing the impact of kernel fusion on gpu tensor operation performance: A systematic performance study,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Analyzing the impact of kernel fusion on gpu tensor operation performance: A systematic performance study,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.207763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:15.658022Z digest=sha256:7ef3a4aced4c953c4e4f8d529aac361dd4c02aecdcc4521e64490b5ecd529417

Observation 1ebde18d-88a7-4062-ad49-e34abe702834 · outbound

This paper cites CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:15.729983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:15.729983Z digest=sha256:a131d7962d0510b41a1540fb66268805a90f6fdce9084bbabae82edeb8c59da0

Observation 6078042a-77b1-4e35-841c-37f57c13b633 · outbound

This paper cites Mixed-input matrix multiplication performance optimiza- tions,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Mixed-input matrix multiplication performance optimiza- tions,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.197930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:15.790178Z digest=sha256:25fd49e187e832d7be8008139a2aa6876308d75cb8a708efc8b3f880a430cc13

Observation 7fa5ce43-b93d-4440-af7d-cc99e82c7e42 · outbound

This paper cites Fireiron: A data-movement-aware scheduling language for gpus,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Fireiron: A data-movement-aware scheduling language for gpus,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.188148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:15.843159Z digest=sha256:5614f9cd84d025fff21497f3dc8fbaa8931f92f043cb323547d7d655ce334e71

Observation fb018d17-0471-4be0-9f06-6950717cd1fc · outbound

This paper cites Making deep learning go brrrr from first principles,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Making deep learning go brrrr from first principles,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.178532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:15.881797Z digest=sha256:d15e1997f887df7352719fff26cc3ed38a411dee8d0af15c22b517c14187529c

Observation 9d5a46cd-8399-4fc1-8dd7-111352935560 · outbound

This paper cites Data movement is all you need: A case study on optimizing transformers,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Data movement is all you need: A case study on optimizing transformers,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.169517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:15.944166Z digest=sha256:1548571610395994c660a7c37c0da8defcb825f9cb02abd42b81ca7f95ad27fe

Observation 6697361a-1a6d-40cb-9f81-84b8fbb683c8 · outbound

This paper cites In-datacenter performance analysis of a tensor processing unit,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures In-datacenter performance analysis of a tensor processing unit,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:16.004122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:16.004122Z digest=sha256:1400197df6e37085d446baa3f11b546b88bd6ed8fd0348e3eef63df3f52d257e

Observation 05a01e35-2706-42a8-b9ab-8ee24c562ece · outbound

This paper cites onednn graph compiler: A hybrid approach for high-performance deep learning compilation,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures onednn graph compiler: A hybrid approach for high-performance deep learning compilation,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:16.083133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:16.083133Z digest=sha256:a235f492db6fd332eb1dec22964b3665fb5c8bc73c2d42525fb83dd4380e22e2

Observation 23b93095-fb03-43d5-aac3-1699d63f0319 · outbound

This paper cites Deep Learning Recommendation Model for Personalization and Recommendation Systems.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Deep Learning Recommendation Model for Personalization and Recommendation Systems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:16.231025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:16.231025Z digest=sha256:6617a700970e71dd77115117e433e7cbab98f0c6245ebcb289b531d5fa80afc8

Observation f1e3e51c-3f46-4c85-86e8-8cbfa62bab2d · outbound

This paper cites Dnnfusion: accelerating deep neural networks execution with advanced operator fusion,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Dnnfusion: accelerating deep neural networks execution with advanced operator fusion,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.149447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:16.305681Z digest=sha256:ec45ecad789f4520ef26ce93c4fbdfd4c38825ac5ac290bff1b6e3d795d3f99e

Observation f1901e4c-4a20-4c97-b9b9-409eb397507a · outbound

This paper cites Cutlass: Fast linear algebra in cuda c++,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Cutlass: Fast linear algebra in cuda c++,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.139488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:16.385907Z digest=sha256:f750dc1805e3a758905d1c01f168576b05c4f8c47754ac919a40f24283843738

Observation ca35fd91-cbdb-48fb-a424-ec2e0d94829d · outbound

This paper cites NVIDIA TensorRT: An sdk for high-performance deep learn- ing inference,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures NVIDIA TensorRT: An sdk for high-performance deep learn- ing inference,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.129913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:16.440457Z digest=sha256:72c02d5b9b0b42af86ab96526384d030108b677f4cd81a27a4dd0cbb29fca8d4

Observation 89097831-03eb-41e5-b98f-24a82aea9918 · outbound

This paper cites cuBLAS: The nvidia cuda basic linear algebra subroutines library,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures cuBLAS: The nvidia cuda basic linear algebra subroutines library,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.120105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:16.547526Z digest=sha256:e50a2571a1ef040372c6bd89b6402d281443f63822f51603d174cb8a8bc70b1a

Observation 5a209196-fa8c-452d-b4d0-d75d7bf85ba5 · outbound

This paper cites Cuda c++ programming guide,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Cuda c++ programming guide,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.109012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:16.641554Z digest=sha256:39422b4473a74c910a6e9bc4959f836eb37e78311ec6759de4565d71a5eaaad2

Observation 686478c4-0146-4fbc-b377-e36874ff1810 · outbound

This paper cites Triton: An open-source programming language for writing highly efficient gpu code,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Triton: An open-source programming language for writing highly efficient gpu code,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.098694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:16.692826Z digest=sha256:a99acd223760bd076dc4341aba3e8fdfa2783c1f1a82f3d156aefae4d99833b5

Observation 0a9ee1b1-8a15-42b4-8ec1-b9bbe5dc9b10 · outbound

This paper cites Automatic kernel fusion for image processing dsls,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Automatic kernel fusion for image processing dsls,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.089869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:16.771487Z digest=sha256:68ee926193834be8cbfad9c32944fed61c80b7998c49183b0f05d76bd624441d

Observation 312e487d-0668-4644-8a4f-af5fe7ce1592 · outbound

This paper cites Tensor program optimization with probabilistic programs,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Tensor program optimization with probabilistic programs,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.079505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:16.872430Z digest=sha256:e1c596ebae56fdc5f7a020690b25301cc70456664b0499b85c4aa935d8885ef1

Observation ffe88e2c-271f-4114-aea2-50edce9eb9ac · outbound

This paper cites Astra: Exploiting predictability to optimize deep learning,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Astra: Exploiting predictability to optimize deep learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.068888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:16.938136Z digest=sha256:d926eaa5c7322fd8e6aaf2551b9a199d0229cb0f29c11eeba63d81494eb13e87

Observation 863ded5a-d359-412e-9574-d3d5d3b30703 · outbound

This paper cites XLA: Optimizing compiler for machine learning,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures XLA: Optimizing compiler for machine learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.059093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:17.003096Z digest=sha256:bdac3c8ea058bea543e04876525a9dc92df9c2d1714e1326a1b7e4e6b2db3b1a

Observation 4fa07624-d39a-45ce-9125-a17a8d8aeea1 · outbound

This paper cites Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:17.006441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:17.006441Z digest=sha256:f2eb84e08d37b0bac478b4aed1addecca45fefdb9d149282e4e694fa06fa86ac

Observation a844f1b8-d2d4-49c5-95a5-de50376987a9 · outbound

This paper cites Attention is all you need,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Attention is all you need,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:17.033024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:17.033024Z digest=sha256:8d3e3564fb3061f40f718a9d5a3afa347bdbf67686c2eb7bbab2abad27bb1ad0

Observation f6f83c22-51b8-4517-8890-71eb6f4cf4a1 · outbound

This paper cites Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:17.119250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:17.119250Z digest=sha256:77dcf6a53f1992cf456b8984c86d695f8374d1e377a6fee3f834e388ece646aa

Observation cdffdb4f-b987-41bd-ac8e-7b6ce2e49f4d · outbound

This paper cites Mirage: A{Multi-Level}superoptimizer for tensor programs,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Mirage: A{Multi-Level}superoptimizer for tensor programs,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.042563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:17.204354Z digest=sha256:638fe15b2c3386341434eb2d8a7818e35e67e4bf7de98bdbc78651fb41a123fc

Observation 36b85fd2-7fc5-45ab-b66d-1afe3a4597ef · outbound

This paper cites {PluS}: Highly efficient and expandable{ML}compiler with pluggable graph schedules,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures {PluS}: Highly efficient and expandable{ML}compiler with pluggable graph schedules,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.033790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:17.291030Z digest=sha256:43432608b7be2aead5565ab58c2251b18945a229d2e99590c5c8ae5604a2df0d

Observation 928d8a25-e863-44a8-90f3-64c05f261e48 · outbound

This paper cites Bolt: Bridg- ing the gap between auto-tuners and hardware-native performance,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Bolt: Bridg- ing the gap between auto-tuners and hardware-native performance,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.024271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:17.409790Z digest=sha256:4a4541dc855d6ef0d2f4b05bc5237f7f66a2d034a1ba389d0e5ed8f9f41d7d48

Observation 5f0a2b04-2e16-4526-9ce1-76b6c58dd2a8 · outbound

This paper cites Demystifying tensor cores to optimize half-precision matrix multiply,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Demystifying tensor cores to optimize half-precision matrix multiply,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.015551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:17.590768Z digest=sha256:1a834b142b3e3fa8b80307b76baf47b497551b3a4a477411527044df8db150c7

Observation 2129f888-f11e-474b-b1e0-4ac00c7a7479 · outbound

This paper cites Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:14:19.222523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:17.683649Z digest=sha256:efada18493b6d4bd4c2ceb20c742c65b6ce34f266f89e9310102d9535e1928e3

Observation eac4fa19-6532-4ecc-9fc5-12b560288995 · outbound

This paper cites Mcfuser: High- performance and rapid fusion of memory-bound compute-intensive operators,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Mcfuser: High- performance and rapid fusion of memory-bound compute-intensive operators,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.006601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:17.813122Z digest=sha256:f888a3c53a3d9c7a73f4cb940958fd7d6df562ca3c9f25e65a1ae3b499f28f52

Observation 7286c7c0-384d-4940-984f-f803229c4991 · outbound

This paper cites Apollo: Automatic partition-based operator fusion through layer by layer optimization,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Apollo: Automatic partition-based operator fusion through layer by layer optimization,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:19.997973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:17.906380Z digest=sha256:4043492130a8553810683c636ab28608d53dd272e083e3e0e20991c4eddc76f3

Observation 8988c86d-05de-410c-92a2-c50d5273d328 · outbound

This paper cites Operator fusion scheduling optimization for tvm deep learning compilers,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Operator fusion scheduling optimization for tvm deep learning compilers,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:19.988750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:18.030327Z digest=sha256:a611cb8c6de1075e4d9d06a32c8d140bc12e778dd4835f4ea5ed623df3129553

Observation c3151d97-b989-4776-adf9-6845a8625ce6 · outbound

This paper cites Ansor: Generating{High-Performance}tensor programs for deep learning,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Ansor: Generating{High-Performance}tensor programs for deep learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:19.979920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:18.157773Z digest=sha256:961e432c95d08294c472ffadbefc4a3530b122138591215a9f5e6f8445c40cf8

Observation df1ebaf6-9996-4733-8e75-88d7bbf85b24 · outbound

This paper cites Chimera: An analytical optimizing framework for effective 12 compute-intensive operators fusion,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Chimera: An analytical optimizing framework for effective 12 compute-intensive operators fusion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:19.970526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:18.299365Z digest=sha256:bf1f76c33821f71589a4c4cf4eb35565d436ca6a67e1b544806f9044f5d26340

Observation 1fd8d8ba-539e-47e2-8964-c73a52e0752b · outbound

This paper cites Astitch: enabling a new multi-dimensional optimization space for memory-intensive ml training and inference on modern simt architectures,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Astitch: enabling a new multi-dimensional optimization space for memory-intensive ml training and inference on modern simt architectures,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:19.768992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:18.390911Z digest=sha256:efa973a9f26854825eb1be5e8f8e523cacc3e7abc497673c623365b1c8ebd6a3

Observation 6c2a8ee3-907b-4f38-bae2-b48a41f44e22 · outbound

This paper cites FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:14:18.939842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:18.523453Z digest=sha256:1a55cc6f4631a0d36a1889ddf559730ba4754bc09e2f8f589b9041d5efc5372b

Observation 64845e81-c353-4684-8de4-7f7da5983a75 · outbound

This paper cites Deep interest network for click-through rate prediction,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Deep interest network for click-through rate prediction,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:19.549972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T17:14:18.740009Z digest=sha256:17e64ef72580466798372e2a1f76f2f15e0a708917e7c9f6831548ca2a45564d

Pith citing papers

No inbound Pith citation observations are available.