Pith. sign in

Paper Citation Record · LEDGER

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding

As of 19 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.05303.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05303 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:48:05.592441Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact5
  • verified fuzzy25
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed695317-8d7c-41bc-837b-f2d21b5eca00 · outbound

This paper cites MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.370776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.370776Z digest=sha256:c390fe736b72af5507505cda5bf9a97a38b321d14c7c2712a88a459174e49054

Observation f3f5d51a-8acc-4b44-ae49-dbfe0854c943 · outbound

This paper cites EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.701101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.374909Z digest=sha256:2d5e8c91e6023f30c649be085e96fdb67429e1b89f5e43b48867cc1e1d4aca8c

Observation ae263b66-9294-4466-8a4f-00936893b512 · outbound

This paper cites SMoLPU: 122.1µJ/Token Sparse MoE- Based Speculative Decoding Language Processing Unit with Adaptive- Offload NPU-CIM Core,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SMoLPU: 122.1µJ/Token Sparse MoE- Based Speculative Decoding Language Processing Unit with Adaptive- Offload NPU-CIM Core,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.693024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.378530Z digest=sha256:111a8881a0ac55e7617c7ea4368c370e17576b4864fe62f8abf77a4987e7380e

Observation 5b2ecdfb-c0dc-4ed0-a45a-9bcbf43469f4 · outbound

This paper cites GPT-4 Technical Report.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.381623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.381623Z digest=sha256:0cad3e86de41c6e091913e8d6a716c062b21c8afeb8985be61a70c29debd5a3f

Observation db22c0f0-b175-463e-9034-0cf8ee242061 · outbound

This paper cites The Llama 3 Herd of Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.385130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.385130Z digest=sha256:b5727c21c9f4356d8333fe5df2bbca3c3175cfdd7a0f1a5628a45858485cb2c1

Observation 051c5fe7-3177-483f-a8ef-098e51f1b1e2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.388512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.388512Z digest=sha256:012d76d11dd0084156decdc4ca26a32e8c7b9371f16de21dd2d7143e89cc2ca1

Observation 4e9b8a5e-85bd-4193-bf90-dce9372c11d4 · outbound

This paper cites Squeezed atten- tion: Accelerating long context length llm inference,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Squeezed atten- tion: Accelerating long context length llm inference,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.392025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.392025Z digest=sha256:a3c47344d925bb0b140ec9c785d6110e560025166e958f41ad0bbe88b447e0e7

Observation 809ce551-d722-49fb-b103-277daee70d60 · outbound

This paper cites ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.685342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.394901Z digest=sha256:24d348fbbcc927689548fe422fc42bbf0f1d5c3f001ebc519ae35c27e790bf19

Observation ac65b07a-5a20-485d-aca6-501ffe8b44d5 · outbound

This paper cites 23.7 BROCA: A 52.4-to-559.2mW Mobile Social Agent System-on-Chip with Adaptive Bit-Truncate Unit and Acoustic-Cluster Bit Grouping,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 23.7 BROCA: A 52.4-to-559.2mW Mobile Social Agent System-on-Chip with Adaptive Bit-Truncate Unit and Acoustic-Cluster Bit Grouping,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.676879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.398249Z digest=sha256:d16408df7324e4b5f677398edf05b5cf4ef0157053b5d1aa096f45ccb60913f8

Observation ba28f471-836c-4fd3-a67c-ef42ce85d1c0 · outbound

This paper cites Fast on-device LLM inference with npus,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast on-device LLM inference with npus,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.667360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.401099Z digest=sha256:d83983630354eb9f87b7697ea9eaba15d4fd677400b44508ce2bb160aac90a8a

Observation 511be75b-39ab-4849-b3e2-28be5a08d4d8 · outbound

This paper cites C-Transformer: An Energy-Efficient Homogeneous DNN- Transformer/SNN-Transformer Processor for Large Language Models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding C-Transformer: An Energy-Efficient Homogeneous DNN- Transformer/SNN-Transformer Processor for Large Language Models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.658221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.404815Z digest=sha256:08324ffaf7aa1fa1dbafe5c112c26da39d51607886c6bc9e9864c9296ac9ae5f

Observation c4fb3c66-d9e3-42e2-8ef6-7be519996abf · outbound

This paper cites MECLA: Memory-Compute-Efficient LLM Accelerator with Scaling Sub-matrix Partition,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MECLA: Memory-Compute-Efficient LLM Accelerator with Scaling Sub-matrix Partition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.648874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.407945Z digest=sha256:ad19708ea212b97dd8949d4af2874536124567c410d8bab72dd009b6b62f682c

Observation 1838530b-5a93-4da0-8371-ce4af794968f · outbound

This paper cites Granite 3.0 Language Models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Granite 3.0 Language Models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.639501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.411220Z digest=sha256:8259a55158c7a20770e1fff8adc55fcbc299cedf2dee89f260edf47153d2f124

Observation 4abd0f5e-0d08-47be-a10b-d0e55a78211a · outbound

This paper cites Qwen3 Technical Report.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Qwen3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.417780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.417780Z digest=sha256:e4431130c34904cef9bb612da50c8033d33e2cc37a962c907eef1db33a1c22be

Observation a5d2494a-4a8b-45d7-94f5-a145dd962520 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.420966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.420966Z digest=sha256:4cfcae2eac7b758bba0d3b2670c04d7b5f475ae2c1e6c0d05b5e3baf825f0aff

Observation b6b9e935-9268-43a5-bac9-02781e31ef48 · outbound

This paper cites GlaM: Efficient scaling of language models with mixture-of-experts,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding GlaM: Efficient scaling of language models with mixture-of-experts,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.621676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.424336Z digest=sha256:47e2684d1d97157c566ce38e02d8075e334e448a11334e1fd2f090598107c43e

Observation 664ae81a-5b4f-4a52-919c-51ae0d8ffa87 · outbound

This paper cites Mixtral of Experts.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Mixtral of Experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.427000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.427000Z digest=sha256:aac63d69416019e78ab8c8d5901f447d2d2c3c72d59c2149530bf213cfc4c679

Observation 7a846a52-6e18-4057-9cca-fb3633c892d5 · outbound

This paper cites LLaMA-MoE: Building mixture-of-experts from llama with continual pre-training,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding LLaMA-MoE: Building mixture-of-experts from llama with continual pre-training,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.612630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.430825Z digest=sha256:c196fc72be322fb5f03c02166cba1a3c4ad277af93c954055e4dd89cdc962311

Observation 88b87ab0-4e6a-4802-9d8f-310c81853211 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.433834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.433834Z digest=sha256:e1df591139ae268384c370de29777f3fe0b2c78d6b7aeac7f9fa448cc25507bb

Observation 1823bac8-8314-4eac-823e-f9ef4c1401c8 · outbound

This paper cites Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.437417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.437417Z digest=sha256:3e034124bd5d91de88bdc51ae216023dcecae0affb21a95f7dad1cbf86c69615

Observation 8c9a80ed-3034-4919-b947-ed19cdf9a461 · outbound

This paper cites Fast inference from transform- ers via speculative decoding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast inference from transform- ers via speculative decoding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.440658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.440658Z digest=sha256:3737cdcb8dd982f2ee43eda69e7724584e53ebe7b4d6e660118f4307ca66b3ae

Observation 961ebce8-44ad-4258-a45a-2eba2a942fc6 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sampling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.444304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.444304Z digest=sha256:6617dca7cb880b2e64d07a4d5a0ec80176af69fcbcbcd661cdeabf91b517ecf3

Observation e6f76a28-735c-4b5f-9efa-b3e29193f342 · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.447443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.447443Z digest=sha256:4ad9da4977a03abb969ab1172744ac4f41c0007bcdfb34988c83552a62ca6cbb

Observation 7f16ba3b-926f-441d-91b0-9aa9211f3dee · outbound

This paper cites Speculative decoding with big little decoder,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Speculative decoding with big little decoder,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.597995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.450810Z digest=sha256:8be67ff3e728ef0826bbcddc582a623e41c04ed1c2dc5197f2332f0d343ba917

Observation eed58b92-de0e-40ab-9dd9-621f04b8eca5 · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.453468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.453468Z digest=sha256:3695f8050ada9e5c3d4f836d79a013719e805427933959f5db22472d6b8f2458

Observation 85835bdf-8a21-4e13-bce0-e53038b268d5 · outbound

This paper cites ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.456995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.456995Z digest=sha256:3e23b2216c0f4a33dff0504671d42ce31989283a866bce7dfe4a1e1639a32ed7

Observation fb2305c1-23da-42da-8dc2-8b7077efd693 · outbound

This paper cites EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.588032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.460225Z digest=sha256:4e65edc1d88be5863fa71edce80310ec886e94d3816b08b4c090eecb354e6265

Observation cc9a7068-a057-40fd-b7fe-df79c43812e4 · outbound

This paper cites SpecMemo: Speculative Decoding is in Your Pocket.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecMemo: Speculative Decoding is in Your Pocket

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:48:07.063112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.463273Z digest=sha256:89c8e1da58b33bc7556c93dc19d2e6d53ad2bb40358f16c6acb70dc692f52d7e

Observation 82ee0902-2a62-4b5e-90af-0586845a430d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.466359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.466359Z digest=sha256:a9052fd362daa0c45b5338286371a0e15acd31371b4d4c801d84f29dad1b4866

Observation aa1bf1a1-79ba-4f9a-860e-05bd8e02feb6 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.469674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.469674Z digest=sha256:0576d048adb1c292002c0707e26bffb162e194706ab68c4031445934e340823d

Observation abf13411-7af7-4b04-82e5-104d3b16802a · outbound

This paper cites MoESD: Unveil Speculative Decoding’s Potential for Accelerating Sparse MoE,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoESD: Unveil Speculative Decoding’s Potential for Accelerating Sparse MoE,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.472875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.472875Z digest=sha256:e6f582f6226f0ffbcb11c14b93544cfbcfaa9fafbffbbecb182830daa90ef397

Observation dfe7d41f-1e8f-4dcc-9ad0-2d43a8c8b853 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.475775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.475775Z digest=sha256:87b30d1a8f5aa2d9667ed6dafac6986380b0bff870725165c3019a7ee179a951

Observation fed56ccb-64d4-4da0-84f1-215bb62753bd · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding OLMoE: Open Mixture-of-Experts Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.478932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.478932Z digest=sha256:b31074270830ae52081f8970d818a3ecdddd8c5c9dec32a11bd1e52cfc741c76

Observation a17e14cd-43c7-43ad-a5bd-340afe275ac9 · outbound

This paper cites Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.482175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.482175Z digest=sha256:a593710c03e53352905a67a04919b5e3c3f1e728da08dcc744cbf6effb78f7df

Observation ea6b6aec-b0aa-4b41-9583-7c5cc053537a · outbound

This paper cites Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.578776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.485453Z digest=sha256:b3dea9f04092f02846e860fdf6fc7b293e710e61e21ef483179c2f731aae9567

Observation 0251c8c5-b9f9-48cc-ab79-2af095ca6c76 · outbound

This paper cites Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.568814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.488860Z digest=sha256:8273a37c66ae7d0517c70315f26b3b8e7b67463c6a697ace3b288ebf2039950f

Observation e67b1ddc-3bec-4b49-9202-303e48de200a · outbound

This paper cites 20.8 Space-Mate: A 303.5mW Real-Time Sparse Mixture-of-Experts- Based NeRF-SLAM Processor for Mobile Spatial Computing,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 20.8 Space-Mate: A 303.5mW Real-Time Sparse Mixture-of-Experts- Based NeRF-SLAM Processor for Mobile Spatial Computing,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.559672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.491756Z digest=sha256:7bafef87a16e88bc042e6c041e3d0a2aee266efb7e978efa331a790ef027e922

Observation 8c15e677-9f27-46b1-bb99-04049c5179a9 · outbound

This paper cites SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification,

Reference 38

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:06.818755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.495329Z digest=sha256:3953047750a29aa9ab7ea8affd327905270a526a509439763d1d593ab1adff02

Observation 1de07ba8-0bac-4dc7-a0db-c8acf3181940 · outbound

This paper cites SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.498093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.498093Z digest=sha256:eeb64234471bc124c125b5e1ac4e385dc0b7f8c1a60c2c500d1cc1a7db9dbe9c

Observation d7160d2e-1d4a-4abe-a1ef-4aeb4170b491 · outbound

This paper cites Fast best-of-n decoding via speculative rejection,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast best-of-n decoding via speculative rejection,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.550479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.501944Z digest=sha256:9d8fe21968ccdb725cc71d6e0dc1e62afdca340974f1307bb956dc6f06ba649c

Observation 2055f79f-efb3-4781-b2ee-6dea2c53cc7a · outbound

This paper cites Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.504790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.504790Z digest=sha256:e790bfc437e8189348856bd7655b9b6d7b23515f7106fe8ba8d2c302e2a07bfe

Observation 5934ef46-3f3a-4b75-92bd-25a4b590bc57 · outbound

This paper cites Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.541433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.508383Z digest=sha256:442ca6a3812fad8bef8ba242f15579ba05b609049a5512ea9f59c4b54186eb0e

Observation 8385853b-6233-4470-88db-f6eccc8cde97 · outbound

This paper cites Moe-i2: Compressing mixture of experts models through inter-expert pruning and intra-expert low-rank decom- position,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Moe-i2: Compressing mixture of experts models through inter-expert pruning and intra-expert low-rank decom- position,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.532354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.511406Z digest=sha256:d4659f1753e37c4938bdc392bb94fc5f5c6a273dd23425018fcee61a215ec17f

Observation 5ed985b9-0324-446d-8bd4-5eb1f49936ce · outbound

This paper cites Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.515013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.515013Z digest=sha256:1a87ccf4ff9fc16015624c0f5a2d9dbddd66d0e0e23924db27a8c95394b29d1a

Observation c2592ca9-03d6-4912-86e9-b7e0c254c2f5 · outbound

This paper cites Self-speculative decoding for on-device moe acceleration,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Self-speculative decoding for on-device moe acceleration,

Reference 45

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:06.627756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.518084Z digest=sha256:639d9a1a719299c4f2261e69d22ca7571430cc8a232ae41eee119bbf805e0984

Observation 5b962831-68d3-4223-a6e7-fc29abcb80b6 · outbound

This paper cites MoE-Spec: Ex- pert Budgeting for Efficient Speculative Decoding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoE-Spec: Ex- pert Budgeting for Efficient Speculative Decoding,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.521241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.521241Z digest=sha256:0efbdc583804a1093c7921f1cf6603287ca9f07da770a80fb5c04b397ba0ac67

Observation 671a1651-b0a5-4294-a807-7dd54cf9ccbf · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.524086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.524086Z digest=sha256:9262a3e0056546da0304c8aa2a5fa0269c8dbb500d130f4fcc50d289d151657c

Observation f549ddb7-a3b5-4147-a11c-f38f85d2d55e · outbound

This paper cites Emerging properties in self-supervised vision transformers,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Emerging properties in self-supervised vision transformers,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.527731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.527731Z digest=sha256:af2b920b09d19d42901ee1506016ceba8543d744633ed7fabab964ca9042db38

Observation a764ec8c-961b-4968-96a1-09507f4cd167 · outbound

This paper cites Dense passage retrieval for open-domain question answering,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Dense passage retrieval for open-domain question answering,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.530625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.530625Z digest=sha256:c9916f284de652c0c17d911754dbb44b9ee7e9d644fef6e7c6e5799d05b10603

Observation 00f90a41-8bbb-41cb-a3f2-5cb76184f4b1 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.534196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.534196Z digest=sha256:85a3ccce9ecd763c50dd6da177f85d3e7bd0398737d8990e0f29a3ba366902f6

Observation 825ed448-8683-47ed-81eb-7260bb47b5e6 · outbound

This paper cites HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models (Student Ab- stract),.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models (Student Ab- stract),

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.507759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.537368Z digest=sha256:4ef7e6a7db29a9655d9bd4e96cfcd857ea130a1f7b7e947cfe936cc8e71a24ee

Observation 50900402-f295-4453-b4d4-e4d01cb9b06d · outbound

This paper cites Atp-llava: Adaptive token pruning for large vision language models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Atp-llava: Adaptive token pruning for large vision language models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.498714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.540750Z digest=sha256:f95a3748dec9a4311eda45cb5ff2e334728b1196fc94b482d89da8d8373a6b43

Observation 6bc4df81-91b6-4a07-8c93-e95bf1956905 · outbound

This paper cites all-minilm-l6-v2: Sentence transformers model,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding all-minilm-l6-v2: Sentence transformers model,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.489593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.543761Z digest=sha256:9322292c48556e51218f7f509a05185364309edb080c52b2d3e61149564a0d78

Observation 29504228-7919-4bce-8800-7b54eac0093d · outbound

This paper cites Pre-Gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Pre-Gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference,

Reference 54

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:06.180319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.547508Z digest=sha256:d9bbf086e1976c74d315da1abcb7a118bc16361753369074c616c23cdcd47f6b

Observation c2805c7d-735b-49ee-baec-2298e994a166 · outbound

This paper cites Sigma: A sparse and irregular gemm ac- celerator with flexible interconnects for dnn training,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Sigma: A sparse and irregular gemm ac- celerator with flexible interconnects for dnn training,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.550292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.550292Z digest=sha256:45f7f27220ec731ff4e7efe80eba51c592d45cadd293a3c47e17d6344be16dbe

Observation 81e16d87-8484-431b-a7dc-94a1f2e2fd3d · outbound

This paper cites 2022, rev.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 2022, rev

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.475928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.554023Z digest=sha256:4e1239a4faff4cef7702d15c261a4c099aa9f23c24d5d257f285eca59704828f

Observation d188d3c3-0f45-45c3-9001-e3110bebfdb0 · outbound

This paper cites SCNN: An accelerator for compressed-sparse convolutional neural networks,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SCNN: An accelerator for compressed-sparse convolutional neural networks,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.467223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.556825Z digest=sha256:a75c9e862ffb6d255fc4e9c3002f720f438db36cf3d02aff474dc028a844a5cc

Observation d61f043c-d1fe-4200-a1e9-596323a909b5 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.560519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.560519Z digest=sha256:c2d4d6ecc243aeca9a0ab7a743b0571a8ba4b3b8e0468702c9fa7a561fd16b39

Observation 49e0ac6f-c82f-421f-870a-ebb14c9e39f8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Training Verifiers to Solve Math Word Problems

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.563489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.563489Z digest=sha256:658c435dde47ec8a45107394d8c8a5694ed2f10b18e021d7ad17f32845b528db

Observation ff4a0b63-6cce-403f-be61-5e4fa02f69d8 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Measuring Massive Multitask Language Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.566935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.566935Z digest=sha256:5ae35263d5ba7bf82a096c2abf3f43a218d2b12b490fb92d90f81f021c8d330d

Observation 78b6eddc-74c7-4cde-9c81-9e55b7cab3fd · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.569995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.569995Z digest=sha256:bdeb839b99219ccb38c2d8a915bb6df363af4442891266ed167f2502b309fcb7

Observation 3a9f32e8-afbd-456d-870b-e8a4b73faf5d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.573207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.573207Z digest=sha256:41252c61bc5f27cb1571f933c0e17212da0f9261980bc0254e79bbd767ab092c

Observation 69a296cd-70d0-443a-9e69-0552bbec13ea · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Winogrande: An adversarial winograd schema challenge at scale,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.576315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.576315Z digest=sha256:8d891197537fa78c948f5e027a64f10b7d65267457ca4b07e6d53303aaf580f8

Observation 88be2782-5a1d-43a9-b183-231a003128c7 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Piqa: Reasoning about physical commonsense in natural language,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.579324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.579324Z digest=sha256:98848220cee781371ca8b44861e2ce22a1001552b338d25bb2e1cdd8afb5898d

Observation 5a9cb099-3d82-4296-8501-7d41cfc99c69 · outbound

This paper cites Pointer Sentinel Mixture Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Pointer Sentinel Mixture Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.582063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.582063Z digest=sha256:5619ea26bf2365989101340218215841635397829cbc9a4ffea59accc517af4a

Observation 3c6fcc8c-25af-4813-b23b-866bc30e9d7f · outbound

This paper cites MLPerf Inference interactive benchmark,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MLPerf Inference interactive benchmark,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.443859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.585984Z digest=sha256:606c71fb44fbd1ea417c8d0d9c53752b5532ff11d0badc4cb7d2a7ea4501d501

Observation dc73aff7-80fb-4e2a-8b99-4cc40d8d1e95 · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Spatten: Efficient sparse attention architecture with cascade token and head pruning,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.589074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.589074Z digest=sha256:0b506617d24198dfa6c1835a774974caf8c8808e05062041723bcca3113d6f5d

Observation 0dbda71c-3190-4436-87ef-c3b05591e032 · outbound

This paper cites FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation Prediction,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation Prediction,

Reference 68

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:05.921856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.592441Z digest=sha256:7e4f0d44caef7dbb92f7768eb33dd59a0c2d67a2d558baac40f3d3fbb66323dc

Observation 55ad2281-72cf-4765-8890-eabf11a64b56 · outbound

This paper cites Available: https://github.com/ibm-granite/granite-3.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Available: https://github.com/ibm-granite/granite-3

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.630644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-08T15:48:05.414109Z digest=sha256:cc34c321cd1748bea7cbafd0de21cde6283b7ac9da8cdb90d6675343a684aaec

Pith citing papers

No inbound Pith citation observations are available.