Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T15:48:05.592441Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.05303.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T15:48:05.592441Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
69 of 69 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ed695317-8d7c-41bc-837b-f2d21b5eca00 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3f5d51a-8acc-4b44-ae49-dbfe0854c943 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ae263b66-9294-4466-8a4f-00936893b512 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SMoLPU: 122.1µJ/Token Sparse MoE- Based Speculative Decoding Language Processing Unit with Adaptive- Offload NPU-CIM Core,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5b2ecdfb-c0dc-4ed0-a45a-9bcbf43469f4 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db22c0f0-b175-463e-9034-0cf8ee242061 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 051c5fe7-3177-483f-a8ef-098e51f1b1e2 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e9b8a5e-85bd-4193-bf90-dce9372c11d4 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Squeezed atten- tion: Accelerating long context length llm inference,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 809ce551-d722-49fb-b103-277daee70d60 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ac65b07a-5a20-485d-aca6-501ffe8b44d5 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 23.7 BROCA: A 52.4-to-559.2mW Mobile Social Agent System-on-Chip with Adaptive Bit-Truncate Unit and Acoustic-Cluster Bit Grouping,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ba28f471-836c-4fd3-a67c-ef42ce85d1c0 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast on-device LLM inference with npus,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 511be75b-39ab-4849-b3e2-28be5a08d4d8 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding C-Transformer: An Energy-Efficient Homogeneous DNN- Transformer/SNN-Transformer Processor for Large Language Models,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c4fb3c66-d9e3-42e2-8ef6-7be519996abf · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MECLA: Memory-Compute-Efficient LLM Accelerator with Scaling Sub-matrix Partition,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1838530b-5a93-4da0-8371-ce4af794968f · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Granite 3.0 Language Models,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4abd0f5e-0d08-47be-a10b-d0e55a78211a · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Qwen3 Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5d2494a-4a8b-45d7-94f5-a145dd962520 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b9e935-9268-43a5-bac9-02781e31ef48 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding GlaM: Efficient scaling of language models with mixture-of-experts,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 664ae81a-5b4f-4a52-919c-51ae0d8ffa87 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Mixtral of Experts
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a846a52-6e18-4057-9cca-fb3633c892d5 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding LLaMA-MoE: Building mixture-of-experts from llama with continual pre-training,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 88b87ab0-4e6a-4802-9d8f-310c81853211 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1823bac8-8314-4eac-823e-f9ef4c1401c8 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c9a80ed-3034-4919-b947-ed19cdf9a461 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast inference from transform- ers via speculative decoding,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 961ebce8-44ad-4258-a45a-2eba2a942fc6 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sampling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6f76a28-735c-4b5f-9efa-b3e29193f342 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f16ba3b-926f-441d-91b0-9aa9211f3dee · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Speculative decoding with big little decoder,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eed58b92-de0e-40ab-9dd9-621f04b8eca5 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85835bdf-8a21-4e13-bce0-e53038b268d5 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb2305c1-23da-42da-8dc2-8b7077efd693 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cc9a7068-a057-40fd-b7fe-df79c43812e4 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecMemo: Speculative Decoding is in Your Pocket
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 82ee0902-2a62-4b5e-90af-0586845a430d · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa1bf1a1-79ba-4f9a-860e-05bd8e02feb6 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abf13411-7af7-4b04-82e5-104d3b16802a · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoESD: Unveil Speculative Decoding’s Potential for Accelerating Sparse MoE,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe7d41f-1e8f-4dcc-9ad0-2d43a8c8b853 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fed56ccb-64d4-4da0-84f1-215bb62753bd · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding OLMoE: Open Mixture-of-Experts Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a17e14cd-43c7-43ad-a5bd-340afe275ac9 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea6b6aec-b0aa-4b41-9583-7c5cc053537a · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0251c8c5-b9f9-48cc-ab79-2af095ca6c76 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e67b1ddc-3bec-4b49-9202-303e48de200a · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 20.8 Space-Mate: A 303.5mW Real-Time Sparse Mixture-of-Experts- Based NeRF-SLAM Processor for Mobile Spatial Computing,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8c15e677-9f27-46b1-bb99-04049c5179a9 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1de07ba8-0bac-4dc7-a0db-c8acf3181940 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7160d2e-1d4a-4abe-a1ef-4aeb4170b491 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast best-of-n decoding via speculative rejection,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2055f79f-efb3-4781-b2ee-6dea2c53cc7a · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5934ef46-3f3a-4b75-92bd-25a4b590bc57 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8385853b-6233-4470-88db-f6eccc8cde97 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Moe-i2: Compressing mixture of experts models through inter-expert pruning and intra-expert low-rank decom- position,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5ed985b9-0324-446d-8bd4-5eb1f49936ce · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2592ca9-03d6-4912-86e9-b7e0c254c2f5 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Self-speculative decoding for on-device moe acceleration,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5b962831-68d3-4223-a6e7-fc29abcb80b6 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoE-Spec: Ex- pert Budgeting for Efficient Speculative Decoding,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 671a1651-b0a5-4294-a807-7dd54cf9ccbf · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Bert: Pre-training of deep bidirectional transformers for language understanding,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f549ddb7-a3b5-4147-a11c-f38f85d2d55e · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Emerging properties in self-supervised vision transformers,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a764ec8c-961b-4968-96a1-09507f4cd167 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Dense passage retrieval for open-domain question answering,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00f90a41-8bbb-41cb-a3f2-5cb76184f4b1 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 825ed448-8683-47ed-81eb-7260bb47b5e6 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models (Student Ab- stract),
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 50900402-f295-4453-b4d4-e4d01cb9b06d · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Atp-llava: Adaptive token pruning for large vision language models,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6bc4df81-91b6-4a07-8c93-e95bf1956905 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding all-minilm-l6-v2: Sentence transformers model,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 29504228-7919-4bce-8800-7b54eac0093d · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Pre-Gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c2805c7d-735b-49ee-baec-2298e994a166 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Sigma: A sparse and irregular gemm ac- celerator with flexible interconnects for dnn training,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e16d87-8484-431b-a7dc-94a1f2e2fd3d · outbound
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d188d3c3-0f45-45c3-9001-e3110bebfdb0 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SCNN: An accelerator for compressed-sparse convolutional neural networks,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d61f043c-d1fe-4200-a1e9-596323a909b5 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Judging llm-as-a-judge with mt-bench and chatbot arena,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e0ac6f-c82f-421f-870a-ebb14c9e39f8 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Training Verifiers to Solve Math Word Problems
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff4a0b63-6cce-403f-be61-5e4fa02f69d8 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Measuring Massive Multitask Language Understanding
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78b6eddc-74c7-4cde-9c81-9e55b7cab3fd · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a9f32e8-afbd-456d-870b-e8a4b73faf5d · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69a296cd-70d0-443a-9e69-0552bbec13ea · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Winogrande: An adversarial winograd schema challenge at scale,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88be2782-5a1d-43a9-b183-231a003128c7 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Piqa: Reasoning about physical commonsense in natural language,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9cb099-3d82-4296-8501-7d41cfc99c69 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Pointer Sentinel Mixture Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c6fcc8c-25af-4813-b23b-866bc30e9d7f · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MLPerf Inference interactive benchmark,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dc73aff7-80fb-4e2a-8b99-4cc40d8d1e95 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Spatten: Efficient sparse attention architecture with cascade token and head pruning,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dbda71c-3190-4436-87ef-c3b05591e032 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation Prediction,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 55ad2281-72cf-4765-8890-eabf11a64b56 · outbound
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Available: https://github.com/ibm-granite/granite-3
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.