Pith. sign in

Paper Citation Record · LEDGER

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding

As of 9 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.05303.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05303 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:48:05.592441Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact5
  • verified fuzzy25
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed695317-8d7c-41bc-837b-f2d21b5eca00 · outbound

This paper cites MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.370776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.370776Z digest=sha256:2814b02680a097ae935583f1b536494850c39b2f1feb6b22422b1fa9597e3071

Observation f3f5d51a-8acc-4b44-ae49-dbfe0854c943 · outbound

This paper cites EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.701101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.374909Z digest=sha256:ea60de39317e1e8b3fc7ddd813e3857197b34aabe9b6283bd89a5801689d0ee7

Observation ae263b66-9294-4466-8a4f-00936893b512 · outbound

This paper cites SMoLPU: 122.1µJ/Token Sparse MoE- Based Speculative Decoding Language Processing Unit with Adaptive- Offload NPU-CIM Core,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SMoLPU: 122.1µJ/Token Sparse MoE- Based Speculative Decoding Language Processing Unit with Adaptive- Offload NPU-CIM Core,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.693024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.378530Z digest=sha256:56d47d5aeea89242986090e6cb421087cdd72452539135b0e921a71dcfa514a7

Observation 5b2ecdfb-c0dc-4ed0-a45a-9bcbf43469f4 · outbound

This paper cites GPT-4 Technical Report.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.381623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.381623Z digest=sha256:5130435426f87518e055fe7b4459c8ce747a4af8f455fc87bbc5ba5ad021b91c

Observation db22c0f0-b175-463e-9034-0cf8ee242061 · outbound

This paper cites The Llama 3 Herd of Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.385130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.385130Z digest=sha256:87fb70adb36c3fd0bb4f04aef9e8f8d430e8afef9486a9b4c3618335208b8ba0

Observation 051c5fe7-3177-483f-a8ef-098e51f1b1e2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.388512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.388512Z digest=sha256:75c936fb4e3ff5e2f7d1efa99ccb1b6c8fa29c89223bf29e40f8c0be4c894d80

Observation 4e9b8a5e-85bd-4193-bf90-dce9372c11d4 · outbound

This paper cites Squeezed atten- tion: Accelerating long context length llm inference,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Squeezed atten- tion: Accelerating long context length llm inference,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.392025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.392025Z digest=sha256:08e99e969f3c1306be6b6efde3f346610acb74fd468ba22b3342aecb300e688e

Observation 809ce551-d722-49fb-b103-277daee70d60 · outbound

This paper cites ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.685342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.394901Z digest=sha256:bdb009976f2f822897951e14f7e2574b37f7bea38521819f7ad55a0cdcf27806

Observation ac65b07a-5a20-485d-aca6-501ffe8b44d5 · outbound

This paper cites 23.7 BROCA: A 52.4-to-559.2mW Mobile Social Agent System-on-Chip with Adaptive Bit-Truncate Unit and Acoustic-Cluster Bit Grouping,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 23.7 BROCA: A 52.4-to-559.2mW Mobile Social Agent System-on-Chip with Adaptive Bit-Truncate Unit and Acoustic-Cluster Bit Grouping,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.676879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.398249Z digest=sha256:d81596d8acf6aad31062319ddbbcd57c04b423eb4dfb37a7b567f54e949af570

Observation ba28f471-836c-4fd3-a67c-ef42ce85d1c0 · outbound

This paper cites Fast on-device LLM inference with npus,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast on-device LLM inference with npus,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.667360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.401099Z digest=sha256:024669c0c5b98a7b51699e8ebda76840696123c9922ac982ed23ed6e1a84ffb0

Observation 511be75b-39ab-4849-b3e2-28be5a08d4d8 · outbound

This paper cites C-Transformer: An Energy-Efficient Homogeneous DNN- Transformer/SNN-Transformer Processor for Large Language Models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding C-Transformer: An Energy-Efficient Homogeneous DNN- Transformer/SNN-Transformer Processor for Large Language Models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.658221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.404815Z digest=sha256:73624a660779c97e7ce11cbfe4a9498c628031bbdf75e6f751df41943a36330b

Observation c4fb3c66-d9e3-42e2-8ef6-7be519996abf · outbound

This paper cites MECLA: Memory-Compute-Efficient LLM Accelerator with Scaling Sub-matrix Partition,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MECLA: Memory-Compute-Efficient LLM Accelerator with Scaling Sub-matrix Partition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.648874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.407945Z digest=sha256:431e5a4358ff933f9e9f1c2e97ab14c7a1611bc33a2237bb84b1249904db82b8

Observation 1838530b-5a93-4da0-8371-ce4af794968f · outbound

This paper cites Granite 3.0 Language Models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Granite 3.0 Language Models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.639501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.411220Z digest=sha256:9e2ca612a831c1d118be7df1ab87083f0fdeca96913e1425c416f12caea26b7a

Observation 4abd0f5e-0d08-47be-a10b-d0e55a78211a · outbound

This paper cites Qwen3 Technical Report.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Qwen3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.417780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.417780Z digest=sha256:5b452251c3fd2b98742eb9c661c824c82fc2f8f09fe049239c46f88bd825ad46

Observation a5d2494a-4a8b-45d7-94f5-a145dd962520 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.420966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.420966Z digest=sha256:855a0d49dd116a03c3e96be3ffabae0140f5219629e87e16a0facdd0337c0aa7

Observation b6b9e935-9268-43a5-bac9-02781e31ef48 · outbound

This paper cites GlaM: Efficient scaling of language models with mixture-of-experts,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding GlaM: Efficient scaling of language models with mixture-of-experts,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.621676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.424336Z digest=sha256:02bda9152b703511c06d6e040147501415b205261c2818af961b58714640145a

Observation 664ae81a-5b4f-4a52-919c-51ae0d8ffa87 · outbound

This paper cites Mixtral of Experts.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Mixtral of Experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.427000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.427000Z digest=sha256:6fa903df753e20ffe5ed0a2a5823f46b04c366f1b3a6024343298b7a4f1d2c6a

Observation 7a846a52-6e18-4057-9cca-fb3633c892d5 · outbound

This paper cites LLaMA-MoE: Building mixture-of-experts from llama with continual pre-training,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding LLaMA-MoE: Building mixture-of-experts from llama with continual pre-training,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.612630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.430825Z digest=sha256:7bf49b15c58ca90c5fbe321e5463920ddea5b136c91b9be7ac772da3b2254789

Observation 88b87ab0-4e6a-4802-9d8f-310c81853211 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.433834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.433834Z digest=sha256:196928b926c1928deeec7cd870807c7f50c52c3b26aa32ab74ebf8fbde6540dd

Observation 1823bac8-8314-4eac-823e-f9ef4c1401c8 · outbound

This paper cites Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.437417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.437417Z digest=sha256:104c9063e98e7dbe80c089544510d21fa46397c9d24dddd28caf963ac64f7160

Observation 8c9a80ed-3034-4919-b947-ed19cdf9a461 · outbound

This paper cites Fast inference from transform- ers via speculative decoding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast inference from transform- ers via speculative decoding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.440658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.440658Z digest=sha256:b03d089f008a62f66060946f305630eb2b2e621d29e86854817660dcecddce20

Observation 961ebce8-44ad-4258-a45a-2eba2a942fc6 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sampling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.444304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.444304Z digest=sha256:9c712f48ae30a3a73cf3f47b214c0993590ddb0eacb0cfd5f534d5c66239f168

Observation e6f76a28-735c-4b5f-9efa-b3e29193f342 · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.447443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.447443Z digest=sha256:3f44bf613e0a815d6f256f68d438af6760f8871211495e10db1cb0cc0fe52c13

Observation 7f16ba3b-926f-441d-91b0-9aa9211f3dee · outbound

This paper cites Speculative decoding with big little decoder,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Speculative decoding with big little decoder,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.597995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.450810Z digest=sha256:50721dd5c9390ee5ec8c85ff5a6e181ca7e3f4f18a6a30c9bdb39a3bacf60aa2

Observation eed58b92-de0e-40ab-9dd9-621f04b8eca5 · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.453468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.453468Z digest=sha256:94f417bc412bac4ecf9cccb8438bc8174d3b92de2d9b9563aff333c9be3ad7b8

Observation 85835bdf-8a21-4e13-bce0-e53038b268d5 · outbound

This paper cites ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.456995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.456995Z digest=sha256:08db35de11b85f1b1be1fce926dbd2ad1e35d51dc801fedacab0fc4f246196ee

Observation fb2305c1-23da-42da-8dc2-8b7077efd693 · outbound

This paper cites EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.588032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.460225Z digest=sha256:2ce7a5b9c7329878aed43cc52626f7001f0f60372e23414dc2843648d5fde814

Observation cc9a7068-a057-40fd-b7fe-df79c43812e4 · outbound

This paper cites SpecMemo: Speculative Decoding is in Your Pocket.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecMemo: Speculative Decoding is in Your Pocket

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:48:07.063112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.463273Z digest=sha256:14a8c1d7953b5f791442105db4d9610a080cbed4086f003400c322bc7d6b24c7

Observation 82ee0902-2a62-4b5e-90af-0586845a430d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.466359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.466359Z digest=sha256:f81a8d6061dd1f46d71158a515c2755283f98a1a76cd20d0ee760dd3e0b297d4

Observation aa1bf1a1-79ba-4f9a-860e-05bd8e02feb6 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.469674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.469674Z digest=sha256:10aad0ec02f70101630e0a89179f681fab7605e46a85f72fbcdfc311ee3faf40

Observation abf13411-7af7-4b04-82e5-104d3b16802a · outbound

This paper cites MoESD: Unveil Speculative Decoding’s Potential for Accelerating Sparse MoE,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoESD: Unveil Speculative Decoding’s Potential for Accelerating Sparse MoE,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.472875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.472875Z digest=sha256:1026cf8db818d837d31c86ebb8222107f10f907f4900da8a509258b8d4bc2526

Observation dfe7d41f-1e8f-4dcc-9ad0-2d43a8c8b853 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.475775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.475775Z digest=sha256:92899063951a29f2f508dd4bbd8e9d85127cae443f5b538b9ac3621d5e084913

Observation fed56ccb-64d4-4da0-84f1-215bb62753bd · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding OLMoE: Open Mixture-of-Experts Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.478932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.478932Z digest=sha256:5452fa095fe1b0165609851d1c9390db8d5a4847b9b5017a5afee1bdc1451e90

Observation a17e14cd-43c7-43ad-a5bd-340afe275ac9 · outbound

This paper cites Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.482175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.482175Z digest=sha256:7d813b4ffb71982f8fd425051e7f15cfb12aa0e8574a2b8b3c7f7dae0721e117

Observation ea6b6aec-b0aa-4b41-9583-7c5cc053537a · outbound

This paper cites Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.578776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.485453Z digest=sha256:14532ca881b9b62ea501615c043df724cbb6a5cb73ee039ed01bc45e20b15803

Observation 0251c8c5-b9f9-48cc-ab79-2af095ca6c76 · outbound

This paper cites Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.568814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.488860Z digest=sha256:22e9c8510fdbf7cb2da7ed09ecaee8572d31d532028e7995ba27d7fa78914a73

Observation e67b1ddc-3bec-4b49-9202-303e48de200a · outbound

This paper cites 20.8 Space-Mate: A 303.5mW Real-Time Sparse Mixture-of-Experts- Based NeRF-SLAM Processor for Mobile Spatial Computing,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 20.8 Space-Mate: A 303.5mW Real-Time Sparse Mixture-of-Experts- Based NeRF-SLAM Processor for Mobile Spatial Computing,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.559672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.491756Z digest=sha256:9ad92a3d6d89019e8ef1c67c28035f3f11eb878739a9b8ac8994bfd44b5433c0

Observation 8c15e677-9f27-46b1-bb99-04049c5179a9 · outbound

This paper cites SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification,

Reference 38

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:06.818755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.495329Z digest=sha256:b1d4015f625c2a2fe06cc8b4658ed85dee110e19925cbe4da418ac6bdd7aa5e1

Observation 1de07ba8-0bac-4dc7-a0db-c8acf3181940 · outbound

This paper cites SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.498093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.498093Z digest=sha256:21155779bd9e6268485e451b6e1e09d50319d1e654a046d4e8446582b96bfb6f

Observation d7160d2e-1d4a-4abe-a1ef-4aeb4170b491 · outbound

This paper cites Fast best-of-n decoding via speculative rejection,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast best-of-n decoding via speculative rejection,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.550479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.501944Z digest=sha256:d19c4e618ba3a88598258833e856aec5b5d29c322795ed1a0d371b691cfa2630

Observation 2055f79f-efb3-4781-b2ee-6dea2c53cc7a · outbound

This paper cites Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.504790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.504790Z digest=sha256:1880b47593831cd319e8e15aef3f76e51dae4caec0082e99ba666c764c0452f2

Observation 5934ef46-3f3a-4b75-92bd-25a4b590bc57 · outbound

This paper cites Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.541433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.508383Z digest=sha256:5f52270985b2ed8230585c44b724a9ae63cedcc0c32b8dc25f096d1b60d17153

Observation 8385853b-6233-4470-88db-f6eccc8cde97 · outbound

This paper cites Moe-i2: Compressing mixture of experts models through inter-expert pruning and intra-expert low-rank decom- position,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Moe-i2: Compressing mixture of experts models through inter-expert pruning and intra-expert low-rank decom- position,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.532354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.511406Z digest=sha256:7f0bafe0366b9bc2beddea5893bc9f86255bdc3544b1e4a42b6577f9a3845422

Observation 5ed985b9-0324-446d-8bd4-5eb1f49936ce · outbound

This paper cites Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.515013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.515013Z digest=sha256:5b71af51a6aab129e66c7113b6ba321ba6bc7e118deb2e3012a9a28817b236e6

Observation c2592ca9-03d6-4912-86e9-b7e0c254c2f5 · outbound

This paper cites Self-speculative decoding for on-device moe acceleration,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Self-speculative decoding for on-device moe acceleration,

Reference 45

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:06.627756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.518084Z digest=sha256:35cb85e0e3091c0a3fd5a37195fdf7720e585cf2f663dd8efc7f082e2e619d67

Observation 5b962831-68d3-4223-a6e7-fc29abcb80b6 · outbound

This paper cites MoE-Spec: Ex- pert Budgeting for Efficient Speculative Decoding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoE-Spec: Ex- pert Budgeting for Efficient Speculative Decoding,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.521241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.521241Z digest=sha256:09c1810aade3228956e9fc38a28c55a83faedeffc57bd7206b97f3742c2d9f32

Observation 671a1651-b0a5-4294-a807-7dd54cf9ccbf · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.524086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.524086Z digest=sha256:4942ff3a707b5febf1cd6fef66b51359bfaa4ca2dc5c13b55105969e6365db16

Observation f549ddb7-a3b5-4147-a11c-f38f85d2d55e · outbound

This paper cites Emerging properties in self-supervised vision transformers,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Emerging properties in self-supervised vision transformers,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.527731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.527731Z digest=sha256:634be3a9f25b3b01998088971e6f43b346bd945539f8ce93dbd01ec0c18fba04

Observation a764ec8c-961b-4968-96a1-09507f4cd167 · outbound

This paper cites Dense passage retrieval for open-domain question answering,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Dense passage retrieval for open-domain question answering,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.530625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.530625Z digest=sha256:23500886ba4f6ca71933f45802c2f63b1871defbaf88bcc9ed80f0cf2ee0e7d9

Observation 00f90a41-8bbb-41cb-a3f2-5cb76184f4b1 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.534196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.534196Z digest=sha256:cd6030cb7420d362ae4672481f6dca9117c2a34d23153f968df30f4852cdc9df

Observation 825ed448-8683-47ed-81eb-7260bb47b5e6 · outbound

This paper cites HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models (Student Ab- stract),.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models (Student Ab- stract),

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.507759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.537368Z digest=sha256:579905a46ff6b2ace12d06f10ceb92ea328eb876b32044bfe1d9e1fdef5ac6ed

Observation 50900402-f295-4453-b4d4-e4d01cb9b06d · outbound

This paper cites Atp-llava: Adaptive token pruning for large vision language models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Atp-llava: Adaptive token pruning for large vision language models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.498714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.540750Z digest=sha256:9db587d0b47b5d73fee3d2ca243702694e8313969ead26253aa5bb4c1b55ef95

Observation 6bc4df81-91b6-4a07-8c93-e95bf1956905 · outbound

This paper cites all-minilm-l6-v2: Sentence transformers model,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding all-minilm-l6-v2: Sentence transformers model,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.489593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.543761Z digest=sha256:4cf012856d25c2d2def5065afe88d05d1f93d829df97bd4bc9719321d103be46

Observation 29504228-7919-4bce-8800-7b54eac0093d · outbound

This paper cites Pre-Gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Pre-Gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference,

Reference 54

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:06.180319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.547508Z digest=sha256:1b2873d5378776b747bf35c2962d17142ba730111de01e61ab252f66ee6af256

Observation c2805c7d-735b-49ee-baec-2298e994a166 · outbound

This paper cites Sigma: A sparse and irregular gemm ac- celerator with flexible interconnects for dnn training,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Sigma: A sparse and irregular gemm ac- celerator with flexible interconnects for dnn training,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.550292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.550292Z digest=sha256:839877a881534bb13f53f83ad917a518290a29a895cfe2f6f63f0a5c0eaa4f80

Observation 81e16d87-8484-431b-a7dc-94a1f2e2fd3d · outbound

This paper cites 2022, rev.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 2022, rev

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.475928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.554023Z digest=sha256:0096e17f7563e36a63913ccbf88f8bc1298273879b0743f0113604a2f42fb776

Observation d188d3c3-0f45-45c3-9001-e3110bebfdb0 · outbound

This paper cites SCNN: An accelerator for compressed-sparse convolutional neural networks,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SCNN: An accelerator for compressed-sparse convolutional neural networks,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.467223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.556825Z digest=sha256:150035d9bb56503cbaaf9386eb382660e0446ec8d7cd7126240440668acecfaa

Observation d61f043c-d1fe-4200-a1e9-596323a909b5 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.560519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.560519Z digest=sha256:75245a7c486f33278e95c511168206d52c58135f45b774feb719d2d402ec6226

Observation 49e0ac6f-c82f-421f-870a-ebb14c9e39f8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Training Verifiers to Solve Math Word Problems

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.563489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.563489Z digest=sha256:ffe21fd64a049151824c26e72f009a91b1053642091f865b107875fbc4d3b70a

Observation ff4a0b63-6cce-403f-be61-5e4fa02f69d8 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Measuring Massive Multitask Language Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.566935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.566935Z digest=sha256:b3d849cbae6fab2fcaec50d0c1f536cb8d8da010f48745a40c0a276714096399

Observation 78b6eddc-74c7-4cde-9c81-9e55b7cab3fd · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.569995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.569995Z digest=sha256:4fb22af18206f23719542e6a28c0fb0288f25c8193db5ac8a05612f2c3aa87a1

Observation 3a9f32e8-afbd-456d-870b-e8a4b73faf5d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.573207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.573207Z digest=sha256:b4beb366fe0ec6435083e03eb548defb37b1d184efce76242082af9328968134

Observation 69a296cd-70d0-443a-9e69-0552bbec13ea · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Winogrande: An adversarial winograd schema challenge at scale,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.576315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.576315Z digest=sha256:b833bc22f223ff115d584846892f9687508b46dc825be39e77fb5448ba3b6f18

Observation 88be2782-5a1d-43a9-b183-231a003128c7 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Piqa: Reasoning about physical commonsense in natural language,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.579324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.579324Z digest=sha256:34db613a4733d0863595eb6c1c90855b29e1dc79f632094e3be4e7a9c42bb387

Observation 5a9cb099-3d82-4296-8501-7d41cfc99c69 · outbound

This paper cites Pointer Sentinel Mixture Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Pointer Sentinel Mixture Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.582063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.582063Z digest=sha256:4af42a23b1327de82f4269e72f67cbc17fe1991833df2a4cbd90456b92f4a1fb

Observation 3c6fcc8c-25af-4813-b23b-866bc30e9d7f · outbound

This paper cites MLPerf Inference interactive benchmark,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MLPerf Inference interactive benchmark,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.443859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.585984Z digest=sha256:da80ed56b94f6a0e277522dd92cdf2773219c86b06bdb339711bb6eb506b9ba7

Observation dc73aff7-80fb-4e2a-8b99-4cc40d8d1e95 · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Spatten: Efficient sparse attention architecture with cascade token and head pruning,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.589074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.589074Z digest=sha256:f90bc5e74f80697e2242aecd0f69cbc7da1223c8d9c807208763c6954bff7810

Observation 0dbda71c-3190-4436-87ef-c3b05591e032 · outbound

This paper cites FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation Prediction,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation Prediction,

Reference 68

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:05.921856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.592441Z digest=sha256:0aedb84c3a2139258ef78c4ee825cd3f0c8c41ed0beeb79a215e1401a8270fa8

Observation 55ad2281-72cf-4765-8890-eabf11a64b56 · outbound

This paper cites Available: https://github.com/ibm-granite/granite-3.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Available: https://github.com/ibm-granite/granite-3

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.630644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T15:48:05.414109Z digest=sha256:292b7132eb136bc7fd2c970f82b0de8d07fca73ca877b9f7c8770e674b4b6e7b

Pith citing papers

No inbound Pith citation observations are available.