Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:14:34.248578Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2507.09010.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:14:34.248578Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 125dcfd3-a8ce-415c-aff3-bf46e1c0c658 · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a46ac35-ff27-42bf-88d0-41f738ec937c · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Architecting an energy-efficient dram system for gpus,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a25e08c6-6864-44fe-a9c9-88abd2013d9d · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Sparc: Token similarity-aware sparse attention transformer accelerator via row- wise clustering,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 34755e6b-31d8-4b32-94d9-815620da707c · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3a0cdbe1-cfcc-48f5-b0d1-54c28458854c · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Transformers are ssms: generalized models and efficient algorithms through structured state space duality,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66c0df4d-6f78-4f89-ba81-79b91d88541e · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3784dc49-c9c0-4a86-8470-cbfb4a3b2835 · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8bc71ae-75fa-4932-9ac9-b44e22b27322 · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Ita: An energy-efficient attention and softmax accelerator for quantized transformers,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6ae7e611-7543-4731-9974-a386988a59ee · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference A 95.6-tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fdb2cc5d-6404-4b53-92de-dc385ed4ba1b · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference A modular digital vlsi flow for high-productivity soc design,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b8d9e08d-f57f-4dab-931e-5d180b80849e · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0387f486-e94f-49c8-8ebe-fc057f5bc4cd · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation def74e63-7de3-44bf-a8cb-f29f4cb1ce2f · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Bucket getter: A bucket-based processing engine for low-bit block floating point (bfp) dnns,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 16c20b43-0be1-4b68-9fb2-77ac679ae042 · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Pointer sentinel mixture models,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e4aa70-d83f-4540-996b-5287005ee479 · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference A 127.8 tops/w arbitrarily quantized 1-to-8b scalable-precision accelerator for general-purpose deep learning with reduction of storage, logic and latency waste,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 68ca6f1a-0c32-4cd3-aa27-8741e8d6070b · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation defed1f5-5f58-4e28-b12b-ed08d7d6e9d7 · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Nvidia jetson orin nano,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 70b8b192-fcd8-4095-a1ae-e25967db41e1 · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Fine-grained dram: Energy-efficient dram for extreme bandwidth systems,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20372408-fd1a-45c2-a80b-55319903ab74 · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Fact: Ffn-attention co-optimized transformer architecture with eager correlation prediction,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 94c2d3bf-c2c2-4fa2-bfb2-d7bc4e72f9ce · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Mecla: Memory-compute-efficient llm accelerator with scaling sub-matrix partition,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 260f865f-bf22-41ba-941c-9f2c8516265b · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Microscaling Data Formats for Deep Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc362b9-8a4e-4af1-ba6e-438221315542 · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Roformer: En- hanced transformer with rotary position embedding,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07c632fd-f886-4442-81de-7ebd830008e7 · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Retentive Network: A Successor to Transformer for Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffd7964c-373e-45db-95dc-293c43766d5f · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f1cfe7-9c9c-4ac1-b62e-a6cc49823000 · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Sole: Hardware-software co-design of softmax and layernorm for efficient transformer inference,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2079e212-5704-4a29-b94c-6110d504a9e2 · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Smoothquant: Accurate and efficient post-training quantization for large language models,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb639e50-da2a-4437-85b3-05ef8531e8da · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Llmcompass: Enabling efficient hardware design for large language model inference,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation be55cd61-3724-4471-a10a-d6df02d4d13d · outbound
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference Atom: Low-bit quantization for efficient and accurate llm serving,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.