Pith. sign in

Paper Citation Record · LEDGER

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers

As of 9 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2605.25655.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.25655 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T20:33:14.055182Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact15
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d2edf1c-7b5c-4f01-a649-3537258fff8e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers LLaMA: Open and Efficient Foundation Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.248919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:37663d138a5a1e070f934924cfa5e77263700458d3ddaa7ac7b235a03cbc979c

Observation 95e28b52-88d2-4668-90fd-4cf9793b9d6f · outbound

This paper cites Qwen Technical Report.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Qwen Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.251341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:a0a058fcf155dcb294f8a855eb430bea8626ff0eb6572191f67d250805cb33b0

Observation 1b25a901-7741-4fb2-8349-057b0166afb2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.239994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:fcc236e68da932022a55cfc5342a2cce123ff9628d5428c758091dce4dd4c99b

Observation b24729db-6e86-48c9-b3d3-1f65efdda61e · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Efficient memory management for large language model serving with pagedattention,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:1441255a395b884b5d9c91b38476c16658a4d4b7331de5d8bd4f960062d06190

Observation e28ca0df-cba0-4d5a-bac2-94ae370e5abf · outbound

This paper cites TensorRT-LLM,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers TensorRT-LLM,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:b30f2e98b9246af58b79c59b6400daec7827cc6f2c94df665ac130c7b8647844

Observation 657778f7-9fef-4f4c-89ee-d74f0b859b80 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:3e5147c922263fbb38da66c981f1d5db31e65dfb071381e1d185e5cfb5ff7900

Observation a82dc052-3a63-447a-b7f9-a44ca9487c34 · outbound

This paper cites Large- scale parallelization and optimization of lattice qcd on tianhe new generation supercomputer,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Large- scale parallelization and optimization of lattice qcd on tianhe new generation supercomputer,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:77b7dacc1822b63552a49a0d2a998e9d8822d42184cf462575904f2931eb2771

Observation 7e9a3747-e92f-4425-857a-33790a644fc7 · outbound

This paper cites Mt-3000: a heterogeneous multi-zone processor for hpc,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Mt-3000: a heterogeneous multi-zone processor for hpc,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:70db610a610651684d41d2a082ed84952b7578851147d9b9cf42f29fc38f601e

Observation fa0119f3-4f5a-4d2c-b148-846aa8e24575 · outbound

This paper cites Performance analysis of cuda, openacc and openmp programming models on tesla v100 gpu,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Performance analysis of cuda, openacc and openmp programming models on tesla v100 gpu,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:1895d290b19e6ed055a2045cda7bdf398788b2bafd6746b62dd05ebd34fdba04

Observation 2ae047b4-1459-4805-b46b-f74eb8462b31 · outbound

This paper cites Mpi: a standard message passing interface,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Mpi: a standard message passing interface,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:9730320912b9a7f2b0e14c8c59ae7590396e84035e318a6bb28d76348984af0d

Observation d2a6cd98-85fd-4ac9-8cfd-58928f2157a2 · outbound

This paper cites Attention is all you need,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Attention is all you need,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:ead7fc61ceae05f7f151eab09ff194599d1de048e4361a6212a706b30fafddc5

Observation 7d4df453-538a-457d-ad6a-318341307ef8 · outbound

This paper cites BERT: A Review of Applications in Natural Language Processing and Understanding.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers BERT: A Review of Applications in Natural Language Processing and Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:33:58.246290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:f2f489303370290bbe3fa8a355dfd3aebe6d5353cdfc956f994fed5f82c97aa4

Observation 0c4244fa-f83a-492e-9ffd-eeb09a8c6d1a · outbound

This paper cites Language models are unsupervised multitask learners,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Language models are unsupervised multitask learners,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:d32a935702b657d21b4a466f971983d6432fe43f04eff2f3f56ea93f3cbd4bcd

Observation ab7d1c02-3047-42c2-8ab1-b67ecc721b36 · outbound

This paper cites Language mod- els are few-shot learners,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Language mod- els are few-shot learners,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:3208e3dc4bb38d7c8ac240065b63c380306ebd2c75ad5bcd850e93a70fc227c8

Observation 6aa8fd6e-d620-4b6c-bb35-add9719f99f9 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.243502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:af1b1340ff08e3b3c2f535fd748eefed63455cbe28d055a07dc5b364db32c6f7

Observation 1dea15d5-e615-4a35-a884-d6ba80c9ff79 · outbound

This paper cites The Llama 3 Herd of Models.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers The Llama 3 Herd of Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.246240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:60c147ede4476d999c5d2f4a1af464877e47776cccfdbe87d29a083e7db497ec

Observation 21f14377-3571-4e4f-9ba5-b10509f94842 · outbound

This paper cites DeepSeek-V3 Technical Report.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers DeepSeek-V3 Technical Report

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.272198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:e6fa5d8fc854a967ecc92a79b53046036c18e9453d48b92a3c165f2fa189a90e

Observation 8ebbd727-232e-4a78-bfff-41660976d5a1 · outbound

This paper cites Qwen2 Technical Report.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Qwen2 Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.275350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:0d13733b689539ee0e9e61654c0b2166a480c47e59bc06a6fb9660d92174c97d

Observation b1f198ea-f712-4f87-bab2-ab647243519b · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.264414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:37687edd625de1ddcfdf924925fcc12dfe738a1496c2af8672186b347fca5d11

Observation ddbe6671-80a4-4826-a793-95ae20a5bb46 · outbound

This paper cites Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:68e9f739021d778a584871e3134dc2bf4dbe331d6a30d990a8ef35a18510f637

Observation 40a76898-7317-486d-9e37-429289e25055 · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Sparsegpt: Massive language models can be accurately pruned in one-shot,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:e3efe2b38db62db83f6e61c2d6fca6b7a37f746ae1b5b47b8c9ee0bbacaf36ca

Observation f327fb1e-5677-4a0f-910b-04398689bbb4 · outbound

This paper cites Llm-pruner: On the structural pruning of large language models,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Llm-pruner: On the structural pruning of large language models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:f42ac8d5c74ed67f2b653685e88810fc11f93e1c38451921d16a1b07ef4ce08a

Observation f3dc798f-e95f-4069-b52d-59445f591d66 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers MiniLLM: On-Policy Distillation of Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.259894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:737b9950d89587aa10404cf1e215566f548febb06996b6d1c368471384524848

Observation 9f542ca8-bdeb-465e-ba21-48a3f60379c6 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:f18330030745c48da1e4c056d201811423ab178495af5db799ec85c64551a641

Observation 888d6e6f-091a-4a3a-85ba-1992102554e9 · outbound

This paper cites Longformer: The Long-Document Transformer.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Longformer: The Long-Document Transformer

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.266584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:2d263c01e92cae956e3573447116a07f52776753127794aa93199da373869996

Observation 542292c2-31d1-45dd-bcbd-b8c389648252 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Linformer: Self-Attention with Linear Complexity

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.269898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:9a94c199689918d39b23839acebcb790c54cb0e46870d4b5105d13deae440aa3

Observation d95be2c3-ca5a-46fc-acc9-0b2073717876 · outbound

This paper cites Reformer: The Efficient Transformer.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Reformer: The Efficient Transformer

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.272420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:5b7ff172a4af94047fa68bbae91697f7eb5ac7245ca47c3695c0a7ca305f0adb

Observation 45e8b6c0-c945-460d-948b-a0152440e980 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.257118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:77068911765af271fab5c77427004f573ce20f3e66d6d8604498945bfea75549

Observation c3ca5121-d7de-485e-b279-63a70b105c69 · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:e6bd9fcee34ebc6bc73a85219ddca09f4a01b0eb7fa5d069e833f88bbbc29bca

Observation e3237c88-7dea-4707-bf75-c063fed24f3e · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Efficient Streaming Language Models with Attention Sinks

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:33:58.255008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:95b685a7d75f78ee5532c0e310b799c41cdd8d84866eaa2791da4937eb73ed7a

Observation f1090996-0c2c-44cb-bd2f-5cf32274a2a9 · outbound

This paper cites Orca: A distributed serving system for{Transformer-Based}generative models,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Orca: A distributed serving system for{Transformer-Based}generative models,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:f6d0787abed5d1193d8c0c45f81a754ac24451395d0347d00762d764f79d24e3

Observation 233a8c46-11e2-40e9-9cd8-bc11ec110e61 · outbound

This paper cites Text Generation Inference,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Text Generation Inference,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:f81ac2979fe76f275541b7c5802e1e6c458d4ba7d74e55c0789ae8cc56f0c805

Observation 7bad1ad4-8492-4b17-8cd0-4099a8cf1dc4 · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Flexgen: High-throughput generative inference of large language models with a single gpu,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:369dcf5254f1a216e6135f02dce1e2b23b98c95f88640e601219e2b21fb8df9f

Observation 5361c5e1-7e74-4f72-afda-c3f52ffbf880 · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:928ad5c4d642f8c9752c360a51dfc7338f6ddd05899348d5cfd66d78b705f6da

Observation 889c0a49-7b04-4a19-8809-a10db6997573 · outbound

This paper cites Efficient processing of deep neural networks: A tutorial and survey,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Efficient processing of deep neural networks: A tutorial and survey,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:586a2533783e353a7c41ebf75975a24a769ba719953b2019e5234be6daed82a7

Observation 894cf718-d9e1-42dc-bc74-dea31e4d690d · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Roofline: an insightful visual performance model for multicore architectures,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:6f059fb5f78baef98e66859a705118894dc23e890c6bd87e28b13e8dabbd2c78

Observation 41e96e60-b30d-4773-abe3-6f9f0800e040 · outbound

This paper cites Optimizing general matrix multiplications on modern multi-core dsps,.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers Optimizing general matrix multiplications on modern multi-core dsps,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:7261d3c8a11aba3c2f6c2621c144ecdf92ca1506446abc9de5f09b93bb9d0f95

Observation 6629f8c6-9f15-4ee2-a633-2beda16cb7db · outbound

This paper cites He served as the chief scientist of China National High Technology Program on high perfor- mance computing for 20 years.

Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers He served as the chief scientist of China National High Technology Program on high perfor- mance computing for 20 years

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T20:33:14.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T20:33:14.055182Z digest=sha256:2b1576f222f075729ccb33c69809e3f205c94e74f89f646c76e79c7f526a7bd9

Pith citing papers

No inbound Pith citation observations are available.