Pith. sign in

Paper Citation Record · LEDGER

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding

As of 9 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.05303.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05303 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:48:05.592441Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact5
  • verified fuzzy25
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed695317-8d7c-41bc-837b-f2d21b5eca00 · outbound

This paper cites MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.370776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.370776Z digest=sha256:0b5582cfa6f59707ce71176a59df5f9289df9b1fac2cdf9161598f7e3269ac70

Observation f3f5d51a-8acc-4b44-ae49-dbfe0854c943 · outbound

This paper cites EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.701101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.374909Z digest=sha256:b7fb4cbbe07bf21f89c0e2e0936ef3fa41c474a29d494a5520220972c261ea7d

Observation ae263b66-9294-4466-8a4f-00936893b512 · outbound

This paper cites SMoLPU: 122.1µJ/Token Sparse MoE- Based Speculative Decoding Language Processing Unit with Adaptive- Offload NPU-CIM Core,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SMoLPU: 122.1µJ/Token Sparse MoE- Based Speculative Decoding Language Processing Unit with Adaptive- Offload NPU-CIM Core,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.693024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.378530Z digest=sha256:d8d25e177f288d9f9e35efebe1358221987e44eed468dc7e507c969ef5ddff36

Observation 5b2ecdfb-c0dc-4ed0-a45a-9bcbf43469f4 · outbound

This paper cites GPT-4 Technical Report.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.381623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.381623Z digest=sha256:a838814de9b50938694124efd9f52b657c0cc71c7f4a1aa20989949536577045

Observation db22c0f0-b175-463e-9034-0cf8ee242061 · outbound

This paper cites The Llama 3 Herd of Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.385130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.385130Z digest=sha256:e80e2497b6649b8ab3be85f47a0687b832bca7ff219cf8e6c3618da08e5efb7d

Observation 051c5fe7-3177-483f-a8ef-098e51f1b1e2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.388512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.388512Z digest=sha256:b9f12dfc9e2ac18907e0398dc91cc83dbc2e2edcb877eca8c37a1c990acdb70f

Observation 4e9b8a5e-85bd-4193-bf90-dce9372c11d4 · outbound

This paper cites Squeezed atten- tion: Accelerating long context length llm inference,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Squeezed atten- tion: Accelerating long context length llm inference,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.392025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.392025Z digest=sha256:5b9275cd247b8f099556b60bba1189a1252f4d4ef254db59abae763a16eb0b80

Observation 809ce551-d722-49fb-b103-277daee70d60 · outbound

This paper cites ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.685342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.394901Z digest=sha256:01645658b32e20bb27a78e81114854ab477b92dac3f92932815b0e8edd7b9b34

Observation ac65b07a-5a20-485d-aca6-501ffe8b44d5 · outbound

This paper cites 23.7 BROCA: A 52.4-to-559.2mW Mobile Social Agent System-on-Chip with Adaptive Bit-Truncate Unit and Acoustic-Cluster Bit Grouping,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 23.7 BROCA: A 52.4-to-559.2mW Mobile Social Agent System-on-Chip with Adaptive Bit-Truncate Unit and Acoustic-Cluster Bit Grouping,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.676879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.398249Z digest=sha256:39d8cfc1cb66816c3126f77212ddc5c4e886355200a7bf932f840487379011e1

Observation ba28f471-836c-4fd3-a67c-ef42ce85d1c0 · outbound

This paper cites Fast on-device LLM inference with npus,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast on-device LLM inference with npus,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.667360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.401099Z digest=sha256:5033b1589cc0ce2710473d023cc9702cd34af4a736dc8b8b94728894b843e5a1

Observation 511be75b-39ab-4849-b3e2-28be5a08d4d8 · outbound

This paper cites C-Transformer: An Energy-Efficient Homogeneous DNN- Transformer/SNN-Transformer Processor for Large Language Models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding C-Transformer: An Energy-Efficient Homogeneous DNN- Transformer/SNN-Transformer Processor for Large Language Models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.658221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.404815Z digest=sha256:076b22279af6fbc454b5bd40fa183090fe726d8062d637e15738e98707394cc9

Observation c4fb3c66-d9e3-42e2-8ef6-7be519996abf · outbound

This paper cites MECLA: Memory-Compute-Efficient LLM Accelerator with Scaling Sub-matrix Partition,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MECLA: Memory-Compute-Efficient LLM Accelerator with Scaling Sub-matrix Partition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.648874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.407945Z digest=sha256:30b5fa060a034826b1fb5abe66789c81b4822ff055e1261032470e2a87eb2a7f

Observation 1838530b-5a93-4da0-8371-ce4af794968f · outbound

This paper cites Granite 3.0 Language Models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Granite 3.0 Language Models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.639501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.411220Z digest=sha256:00250189e7d2e14fb3f7d6d57a91199a5a4ed09400d3a4b214a2bbbf6198330e

Observation 4abd0f5e-0d08-47be-a10b-d0e55a78211a · outbound

This paper cites Qwen3 Technical Report.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Qwen3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.417780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.417780Z digest=sha256:136d19cdeeeb9a99823a6a7ea9b47d6902ed5e5ed64474920c3e60acb019f760

Observation a5d2494a-4a8b-45d7-94f5-a145dd962520 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.420966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.420966Z digest=sha256:4de4b96273cef8ed1200d6dde83a1ce6d3db25cc949bfb4cb6a5219d62a329d6

Observation b6b9e935-9268-43a5-bac9-02781e31ef48 · outbound

This paper cites GlaM: Efficient scaling of language models with mixture-of-experts,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding GlaM: Efficient scaling of language models with mixture-of-experts,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.621676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.424336Z digest=sha256:d71ad1d1a79cf06801c541d1ff4ab04c893b8f9cb8514eb0c79ce82502d2f855

Observation 664ae81a-5b4f-4a52-919c-51ae0d8ffa87 · outbound

This paper cites Mixtral of Experts.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Mixtral of Experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.427000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.427000Z digest=sha256:699f9ce5a2f694e86dbc777c439f96a56fecd819193ebcab1185766e8cd61e43

Observation 7a846a52-6e18-4057-9cca-fb3633c892d5 · outbound

This paper cites LLaMA-MoE: Building mixture-of-experts from llama with continual pre-training,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding LLaMA-MoE: Building mixture-of-experts from llama with continual pre-training,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.612630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.430825Z digest=sha256:7ee265034203b5753c574e1522cf5061e6b576712666040f661ed21e8dc21e1c

Observation 88b87ab0-4e6a-4802-9d8f-310c81853211 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.433834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.433834Z digest=sha256:729c19b9873823c15f2d6f27f1da551c0d9aced393990330ea0fd520d1c59492

Observation 1823bac8-8314-4eac-823e-f9ef4c1401c8 · outbound

This paper cites Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.437417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.437417Z digest=sha256:556b07d77c934ae958cef61431dbb4f9f8fe8cc420b45b7b8ac0ce0d15ff08ad

Observation 8c9a80ed-3034-4919-b947-ed19cdf9a461 · outbound

This paper cites Fast inference from transform- ers via speculative decoding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast inference from transform- ers via speculative decoding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.440658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.440658Z digest=sha256:5c5d1abecc3fbfb6e41eed024b5202a2ce8c6267f2a64cc728f2928a2ce5f696

Observation 961ebce8-44ad-4258-a45a-2eba2a942fc6 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sampling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.444304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.444304Z digest=sha256:40a20521fb0ddcec32b417c07608cd48dcd3ce2a2d1d988f127addc88209e9ef

Observation e6f76a28-735c-4b5f-9efa-b3e29193f342 · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.447443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.447443Z digest=sha256:7da0622dc4b8c6d53c5c61e4aac9d11b87e25131f8f59995a60f46f23e4403a6

Observation 7f16ba3b-926f-441d-91b0-9aa9211f3dee · outbound

This paper cites Speculative decoding with big little decoder,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Speculative decoding with big little decoder,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.597995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.450810Z digest=sha256:8d7ab810f9b54efa7f81edc07fb0e10422e585e3651e8d19758d7e57e5e01038

Observation eed58b92-de0e-40ab-9dd9-621f04b8eca5 · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.453468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.453468Z digest=sha256:bcf476f3863365925de36018c9c82f9623e3cd26414acd459c8e127d0c109bd3

Observation 85835bdf-8a21-4e13-bce0-e53038b268d5 · outbound

This paper cites ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.456995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.456995Z digest=sha256:48aaf9c0259c06d9119c03a1d47f1a6e62395783fd41b4e9f95cae20f171ccde

Observation fb2305c1-23da-42da-8dc2-8b7077efd693 · outbound

This paper cites EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.588032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.460225Z digest=sha256:7766b0e211cd348065e200f51ea63e8d9b98da72b88967ee30c2da5ae8fc3eb8

Observation cc9a7068-a057-40fd-b7fe-df79c43812e4 · outbound

This paper cites SpecMemo: Speculative Decoding is in Your Pocket.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecMemo: Speculative Decoding is in Your Pocket

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:48:07.063112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.463273Z digest=sha256:626398130b38d0ebdb332fd1554bdd398ec5cd1805a4125554ff5d62148d16e0

Observation 82ee0902-2a62-4b5e-90af-0586845a430d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.466359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.466359Z digest=sha256:75f485974c0d783c4bd20c7f3f7b918b4169280c1e46c1b5698d0b0732efdea9

Observation aa1bf1a1-79ba-4f9a-860e-05bd8e02feb6 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.469674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.469674Z digest=sha256:52434b84e3473d03662817ef90a5a95e0a33319a9f4bf9d4046d6df1f96905bb

Observation abf13411-7af7-4b04-82e5-104d3b16802a · outbound

This paper cites MoESD: Unveil Speculative Decoding’s Potential for Accelerating Sparse MoE,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoESD: Unveil Speculative Decoding’s Potential for Accelerating Sparse MoE,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.472875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.472875Z digest=sha256:0075bc7af2865544987286c8be3f364f100f10b1c9b1252d7a8b0e29cc5f8bc6

Observation dfe7d41f-1e8f-4dcc-9ad0-2d43a8c8b853 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.475775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.475775Z digest=sha256:13dea4d3daca687d6228c7c1843f24d3fbaf73e8e526f10e74e4a21b515064e2

Observation fed56ccb-64d4-4da0-84f1-215bb62753bd · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding OLMoE: Open Mixture-of-Experts Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.478932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.478932Z digest=sha256:d27d9034e4644fe32a1a98f21f93ce65711bc95c49a048828365a0a32cfa9f86

Observation a17e14cd-43c7-43ad-a5bd-340afe275ac9 · outbound

This paper cites Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.482175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.482175Z digest=sha256:a516c33d6926c88e0e00670a38f51f02beb7602e8c2fad4f49847efa43a53db1

Observation ea6b6aec-b0aa-4b41-9583-7c5cc053537a · outbound

This paper cites Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.578776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.485453Z digest=sha256:12f9935ae29eb790760238f39e081a08d07660df5d6a5f59abd3a5b581654d77

Observation 0251c8c5-b9f9-48cc-ab79-2af095ca6c76 · outbound

This paper cites Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.568814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.488860Z digest=sha256:037fc2efd599ee6969af33407a76a258b0266389b69b10b7a5e3d79af43eea45

Observation e67b1ddc-3bec-4b49-9202-303e48de200a · outbound

This paper cites 20.8 Space-Mate: A 303.5mW Real-Time Sparse Mixture-of-Experts- Based NeRF-SLAM Processor for Mobile Spatial Computing,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 20.8 Space-Mate: A 303.5mW Real-Time Sparse Mixture-of-Experts- Based NeRF-SLAM Processor for Mobile Spatial Computing,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.559672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.491756Z digest=sha256:fed3e54dbcc23be87436d373ab2ada4e9bca769e1376c8e97d73dfb7564c4ddf

Observation 8c15e677-9f27-46b1-bb99-04049c5179a9 · outbound

This paper cites SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification,

Reference 38

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:06.818755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.495329Z digest=sha256:54cce4cf0e1cbdc2a27c22faef56ccc262fe689da3456dc5dc36cdde23a0047a

Observation 1de07ba8-0bac-4dc7-a0db-c8acf3181940 · outbound

This paper cites SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.498093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.498093Z digest=sha256:379efe3b679d972d9a8f2c39a36dfa6c3b80e8f3c9a57603d81028d7d3919ef3

Observation d7160d2e-1d4a-4abe-a1ef-4aeb4170b491 · outbound

This paper cites Fast best-of-n decoding via speculative rejection,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast best-of-n decoding via speculative rejection,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.550479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.501944Z digest=sha256:c3af19566d7aa74011d8081f72c3e23d0fa76d624536992e7105aa373a0f7ea3

Observation 2055f79f-efb3-4781-b2ee-6dea2c53cc7a · outbound

This paper cites Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.504790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.504790Z digest=sha256:fdaafa5f5ade0995468240232d8bf14616e26f19df3bd4145e2683e995f9e41b

Observation 5934ef46-3f3a-4b75-92bd-25a4b590bc57 · outbound

This paper cites Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.541433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.508383Z digest=sha256:f8810c0f0a7b55ceb18cf72351ba391e3e0b0e29b2ce0e8176bf15364fefb00c

Observation 8385853b-6233-4470-88db-f6eccc8cde97 · outbound

This paper cites Moe-i2: Compressing mixture of experts models through inter-expert pruning and intra-expert low-rank decom- position,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Moe-i2: Compressing mixture of experts models through inter-expert pruning and intra-expert low-rank decom- position,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.532354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.511406Z digest=sha256:d9c0a8e5f77e3f37def673c5a018ce8ce060035a84dffcff4880a5b3cbc1d4b0

Observation 5ed985b9-0324-446d-8bd4-5eb1f49936ce · outbound

This paper cites Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.515013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.515013Z digest=sha256:42458ed4ed91a3b8344af5e01ff23b503d5a277144e3517269b2b92fb3f77054

Observation c2592ca9-03d6-4912-86e9-b7e0c254c2f5 · outbound

This paper cites Self-speculative decoding for on-device moe acceleration,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Self-speculative decoding for on-device moe acceleration,

Reference 45

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:06.627756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.518084Z digest=sha256:fd0e2aa73c46c2d6073fcb1935a3b9485290392a3019481330df2b836924a7f2

Observation 5b962831-68d3-4223-a6e7-fc29abcb80b6 · outbound

This paper cites MoE-Spec: Ex- pert Budgeting for Efficient Speculative Decoding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoE-Spec: Ex- pert Budgeting for Efficient Speculative Decoding,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.521241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.521241Z digest=sha256:452dfbfb4a5650f391637e765361d4a914b126dc3d735f0cdc767c8dba3c29db

Observation 671a1651-b0a5-4294-a807-7dd54cf9ccbf · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.524086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.524086Z digest=sha256:9e0ecf5d0fc9615b328f6f95504e001af20dc0b5b90370c0cfe11f3c8059d8f6

Observation f549ddb7-a3b5-4147-a11c-f38f85d2d55e · outbound

This paper cites Emerging properties in self-supervised vision transformers,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Emerging properties in self-supervised vision transformers,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.527731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.527731Z digest=sha256:45b147686b784cb0352d63a7ad8eecb99e74b8704381dd5c9a7ebe1e6f3e54b6

Observation a764ec8c-961b-4968-96a1-09507f4cd167 · outbound

This paper cites Dense passage retrieval for open-domain question answering,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Dense passage retrieval for open-domain question answering,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.530625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.530625Z digest=sha256:7de9c6bce14d5df7d8fe934de9cb8709bbe5975b34a3805d63423e35354254b0

Observation 00f90a41-8bbb-41cb-a3f2-5cb76184f4b1 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.534196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.534196Z digest=sha256:60456df758ac5703752f53a1dbee347ea608545e8b262b012e6a7f7e2cc2c11f

Observation 825ed448-8683-47ed-81eb-7260bb47b5e6 · outbound

This paper cites HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models (Student Ab- stract),.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models (Student Ab- stract),

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.507759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.537368Z digest=sha256:dc9fcf72902f05f015af636586dd64f0ec77623e31287a288b088c15040f8fb3

Observation 50900402-f295-4453-b4d4-e4d01cb9b06d · outbound

This paper cites Atp-llava: Adaptive token pruning for large vision language models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Atp-llava: Adaptive token pruning for large vision language models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.498714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.540750Z digest=sha256:4aeab4d915be22e3586fcdb2086820f816870752fb4f569928668a36041a02dd

Observation 6bc4df81-91b6-4a07-8c93-e95bf1956905 · outbound

This paper cites all-minilm-l6-v2: Sentence transformers model,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding all-minilm-l6-v2: Sentence transformers model,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.489593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.543761Z digest=sha256:6599e58693f1439f6b7270a3fc528dd6c87b2ef4e356de3ac8c459c800e6a6a1

Observation 29504228-7919-4bce-8800-7b54eac0093d · outbound

This paper cites Pre-Gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Pre-Gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference,

Reference 54

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:06.180319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.547508Z digest=sha256:4a1f0b87b31154ced9f16883a04922ca454016cb20a93a19872bb9a74bb79f4e

Observation c2805c7d-735b-49ee-baec-2298e994a166 · outbound

This paper cites Sigma: A sparse and irregular gemm ac- celerator with flexible interconnects for dnn training,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Sigma: A sparse and irregular gemm ac- celerator with flexible interconnects for dnn training,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.550292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.550292Z digest=sha256:96d21f7066a6185e6bc42798f783819828c6f40e899567c3760e062b67afeda2

Observation 81e16d87-8484-431b-a7dc-94a1f2e2fd3d · outbound

This paper cites 2022, rev.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 2022, rev

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.475928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.554023Z digest=sha256:7bd4350ad5bef04368f12aeea9c86cbbdb6ea8f629fd7cf90e00df1100ee9793

Observation d188d3c3-0f45-45c3-9001-e3110bebfdb0 · outbound

This paper cites SCNN: An accelerator for compressed-sparse convolutional neural networks,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SCNN: An accelerator for compressed-sparse convolutional neural networks,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.467223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.556825Z digest=sha256:cc3df3021864cdabfd9993cf88fb18850eb690ea75b06136ef9c4836367e768e

Observation d61f043c-d1fe-4200-a1e9-596323a909b5 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.560519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.560519Z digest=sha256:65d42a24d86c77abd0e24dd605aa1461760c2c7eb4afd76c6b2758d34ba21bc4

Observation 49e0ac6f-c82f-421f-870a-ebb14c9e39f8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Training Verifiers to Solve Math Word Problems

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.563489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.563489Z digest=sha256:2f732abefe9514164acdc1c6d645cb10345fe2dc4cc49142b302b97ad8e8a95e

Observation ff4a0b63-6cce-403f-be61-5e4fa02f69d8 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Measuring Massive Multitask Language Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.566935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.566935Z digest=sha256:152a86127b5fcc8866ce45ffd27126012c213fc416da957568a8b92810620708

Observation 78b6eddc-74c7-4cde-9c81-9e55b7cab3fd · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.569995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.569995Z digest=sha256:0267ee8961afe255af8dab0f16a87c4260d0b5aad96931cda4e8337649a90616

Observation 3a9f32e8-afbd-456d-870b-e8a4b73faf5d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.573207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.573207Z digest=sha256:4056b7c291fcbed4a0f1c1888a119828bac83ceacc81d75ba59a9454f55f95d9

Observation 69a296cd-70d0-443a-9e69-0552bbec13ea · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Winogrande: An adversarial winograd schema challenge at scale,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.576315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.576315Z digest=sha256:9e7f4fd0b118cb63473f6197f7d3509424ce2c256c5ff8b2961b109b212036ce

Observation 88be2782-5a1d-43a9-b183-231a003128c7 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Piqa: Reasoning about physical commonsense in natural language,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.579324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.579324Z digest=sha256:657ff0b0d0708369394676f276850202c0f0fe795419cde3196b2dd863da84e8

Observation 5a9cb099-3d82-4296-8501-7d41cfc99c69 · outbound

This paper cites Pointer Sentinel Mixture Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Pointer Sentinel Mixture Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.582063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.582063Z digest=sha256:82caf33ae5affd42658318b7a55abfe1c4454ed887116bd4009ec7ad35579f9f

Observation 3c6fcc8c-25af-4813-b23b-866bc30e9d7f · outbound

This paper cites MLPerf Inference interactive benchmark,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MLPerf Inference interactive benchmark,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.443859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.585984Z digest=sha256:db7042b44a61cda271f5c57ce084ce76dc015155625c26f6c0a6153a418b9a1a

Observation dc73aff7-80fb-4e2a-8b99-4cc40d8d1e95 · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Spatten: Efficient sparse attention architecture with cascade token and head pruning,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.589074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.589074Z digest=sha256:a630798b906dc3432a913e9b50edc4624e6207f6e0131ec719819a731aa39251

Observation 0dbda71c-3190-4436-87ef-c3b05591e032 · outbound

This paper cites FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation Prediction,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation Prediction,

Reference 68

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:05.921856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.592441Z digest=sha256:c70614876e8fe8f25d8a7d70b4db65056987fd00bf4cc00e22b0aa77d9a1e66e

Observation 55ad2281-72cf-4765-8890-eabf11a64b56 · outbound

This paper cites Available: https://github.com/ibm-granite/granite-3.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Available: https://github.com/ibm-granite/granite-3

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.630644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.414109Z digest=sha256:abee62f94dbde9403241279f957350c14aa1a8d8433cbff706e36565c210f8fa

Pith citing papers

No inbound Pith citation observations are available.