Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-10T11:11:55.188720Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2607.08215.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-10T11:11:55.188720Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 897c8722-3ab0-4963-9a71-eba72526e343 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e66b49e2-54d3-4c08-875f-b7759226bc93 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 80283ae9-d170-43b4-b891-24b98ccdd472 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fbbf6edf-a691-4e0f-a2dc-2d08c5b65790 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac465b48-fb90-404e-895c-b07fd8522086 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend DeepSeek-V3 Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 075d6415-14e4-4c52-9554-8eacd184263e · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4b94852-e38b-4d51-bc2a-87c2046b499b · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b9970f5c-570b-4a96-a4b1-97c1a9088a8a · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Le, Yonghui Wu, and Zhifeng Chen
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0babd0a7-68ce-4a4c-a87a-51bc8ee76706 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend CANN: Compute architecture for neural networks — documentation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 52f93f76-9e2c-49a1-a96f-f8c95a9f0ff4 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Mixtral of Experts
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cbb4e218-af68-4820-aced-9213ad750dcb · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ff08e0d9-58a2-4217-8ea1-a6155940bf6c · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2383e7e1-9dca-4091-bd4d-ec137da3ee53 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Fast Inference from Transformers via Speculative Decoding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af5ff6bf-dfc2-4cae-8b5e-f6d934e557ed · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 474e0dc6-8c1c-429f-b6ea-8ab81a78b075 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend DaVinci: A scalable architecture for neural network computing
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc4c22cc-3733-4152-bbcb-16b3bd90348b · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Ascend: A scalable and unified architecture for ubiquitous deep neural network computing — industry track paper
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ef1077cc-dc58-4c02-98c2-a8748651db36 · outbound
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c805e96-a852-43d1-a064-730d043576f1 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e2ac552-fbc7-4114-909c-a3122c01d833 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Visual Instruction Tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 520683a6-f800-4144-b9bb-cdfd752c4fa7 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend PyTorch: An imperative style, high- performance deep learning library
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7d6012de-f9a1-497b-81d4-145711defedc · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Learning Transferable Visual Models From Natural Language Supervision
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation be493855-4b23-41f1-a0f3-057dc266a059 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8914cb5e-0ea8-437e-822a-44a85d158b13 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 30170d5b-f11a-479d-a145-a5c690ad7408 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Gomez, Łukasz Kaiser, and Illia Polosukhin
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e20f6114-626b-4f78-a0c9-7d43a4ab4a73 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend vLLM Ascend plugin (vllm-ascend) documentation.https://vllm-ascend
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d2aee196-2cde-4ef4-85af-ce384a7e2106 · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 619f3c47-2f2b-44a3-8616-3e195866da7f · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b52fa409-2221-4a6d-81cc-6993095aa59c · outbound
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Orca: A distributed serving system for Transformer-based generative models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.