Pith. sign in

Paper Citation Record · LEDGER

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs

As of 19 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 3 inbound Pith citation observations for arXiv:2504.16266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16266 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:12:12.027472Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-09T19:06:18.309062Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:51:27.298033Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1fcbbf87-326d-41a0-a155-010f9b692f0f · outbound

This paper cites Language mod- els are few-shot learners,.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs Language mod- els are few-shot learners,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:12:11.908072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:12:11.908072Z digest=sha256:7361937fffe6d3de5924ba939b0306a6562b472458e589afcc31faffaeb9548a

Observation 4ebdf683-005e-40c3-af8a-cde5f9a07130 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:12:11.913879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:12:11.913879Z digest=sha256:07aa63d34964882d7637ee4fbebd88c11870ff8a0371f63a6c73a496ce7832d7

Observation 69ef7ba5-e2b5-490a-b9f9-04e8a0bfad9f · outbound

This paper cites A two-stage efficient 3- d cnn framework for eeg based emotion recognition,.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs A two-stage efficient 3- d cnn framework for eeg based emotion recognition,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:12:12.380073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:12:11.923677Z digest=sha256:366d821b0c59682923df9d309ebfe8225784669b63899960d6289fb94f33fd6e

Observation a2a2012a-e994-4d55-9375-f73dc497e20e · outbound

This paper cites Bnn an ideal architecture for acceleration with resistive in memory computation,.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs Bnn an ideal architecture for acceleration with resistive in memory computation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:12:12.365582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:12:11.929770Z digest=sha256:123d66cbe5eceff254aec13a7ca9b1b375b5ddeba0449ae60a4dff9b778fc7f6

Observation 8896b396-eb9d-40f4-ab90-eee97880ee23 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:12:11.934753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:12:11.934753Z digest=sha256:fa52a3f3514ae59a891a5505ff5cc5c2932075dde0be65256c5f6bba9cee2acc

Observation 375ec7aa-e3d7-44ec-aa3d-6754f01c8611 · outbound

This paper cites The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:12:11.940413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:12:11.940413Z digest=sha256:7207a355c65c3312b0bf21a33416227ce229899d6a3b6e248f676effd07555c4

Observation 19a237eb-9dcc-4556-9c49-d8189e269fc3 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:12:11.946205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:12:11.946205Z digest=sha256:4acde19be9232f229a5f7cd8e763946f86a827dc114d1e03288153f2ab29269a

Observation 91c1d3e8-5fa9-4a56-be39-1dcd4d8c27de · outbound

This paper cites Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:12:11.951884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:12:11.951884Z digest=sha256:d3cc96b4180286f71fc4969b94df600bd09c0ab7185c7d2ab9c5fe161765ac2a

Observation 6c96c758-2c2a-4018-abe1-1d902c3be773 · outbound

This paper cites Understanding the potential of fpga-based spatial accel- eration for large language model inference,.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs Understanding the potential of fpga-based spatial accel- eration for large language model inference,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:12:12.349465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:12:11.958490Z digest=sha256:88140f96c5e13aeb5d36beb91add396adfb4d048e2836f04a74c30263aa88bd4

Observation 2a75ffbf-02db-4237-a577-2c3f7ce488c0 · outbound

This paper cites FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive Distillation.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:12:11.965727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:12:11.965727Z digest=sha256:e5efe9aa6741707ed569d1b98b0500dd341a021ba7ddc3a2fca7d2416e7d56e5

Observation cd4a10aa-f0c7-42e6-9ac9-42a86e4a5f9d · outbound

This paper cites Onebit: Towards extremely low-bit large language models,.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs Onebit: Towards extremely low-bit large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:12:12.331178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:12:11.973937Z digest=sha256:8b6cbfe3d9f7ed28808a6d017be3068664ad79536cf5f5b48f0f8f232c9238d2

Observation f1a51205-de58-4d2c-903d-30a5735ac138 · outbound

This paper cites Bitdistiller: Unleashing the potential of sub-4-bit llms via self-distillation,.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs Bitdistiller: Unleashing the potential of sub-4-bit llms via self-distillation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:12:12.316137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:12:11.980320Z digest=sha256:75f308db4d93cdd788d56cc5a2e345f6c844a3fb8acf26e593fc9b049e591cc1

Observation 7372115b-f171-4c0d-8498-72f0b1546050 · outbound

This paper cites QuIP: 2-Bit Quantization of Large Language Models With Guarantees.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs QuIP: 2-Bit Quantization of Large Language Models With Guarantees

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:12:11.986624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:12:11.986624Z digest=sha256:bb9b50c005ba754be3b777c2b46d6b3c513bd5cf24cd8c2d5ca457de41c0515f

Observation 8e73ae95-50dc-49d5-a921-098bde8fe1d5 · outbound

This paper cites T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:12:11.993156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:12:11.993156Z digest=sha256:2d0e75a45379d8a88c196a475a1120187116729d9e977ce025544b2b5147190c

Observation 21b859ee-1452-4092-b33f-dd7da05fb8af · outbound

This paper cites LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:12:11.998609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:12:11.998609Z digest=sha256:89233ea7e6746809c2ff8277837cb22a358ab8b48ab495f928117752d3c003bb

Observation 207f4096-854e-461c-9d39-3a55e4146da4 · outbound

This paper cites Edge-moe: Memory- efficient multi-task vision transformer architecture with task-level spar- sity via mixture-of-experts,.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs Edge-moe: Memory- efficient multi-task vision transformer architecture with task-level spar- sity via mixture-of-experts,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:12:12.301205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:12:12.003606Z digest=sha256:8c1b159f9c170245c25993132007c79ded293f94de7f321a3738a80b1849b7c9

Observation 52dd96e2-493e-4779-80e4-ea56fe257e59 · outbound

This paper cites Designing Efficient LLM Accelerators for Edge Devices.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs Designing Efficient LLM Accelerators for Edge Devices

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:12:12.008702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:12:12.008702Z digest=sha256:b402960df524726d1d8a795b311d4a2725e114f872e6b535533fa27edf2ff590

Observation 75c503f5-bf81-47c9-8e08-9e2ffa8ce75a · outbound

This paper cites Table- lookup mac: Scalable processing of quantised neural networks in fpga soft logic,.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs Table- lookup mac: Scalable processing of quantised neural networks in fpga soft logic,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:12:12.284325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:12:12.014685Z digest=sha256:e534935872b8bf7ad3d068520dbccd1e474dc42ff97a16366232b59ef5a9a047

Observation 7f52e666-9468-441c-a13b-dcf0676ed204 · outbound

This paper cites Uram storage bind,.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs Uram storage bind,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:12:12.267221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T11:12:12.021281Z digest=sha256:e30a44e08651c876c179aa751d5f5012b2115fe9257f17a670425422170a6561

Observation 55ab18f8-1cef-4b15-b077-5dc3fe978f8f · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:12:12.027472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:12:12.027472Z digest=sha256:a51333a320350e0f3e97d6b51c89dca6fc2e9462125409574035a1c3009bd05c

Pith citing papers

Observation 25a31676-9d8a-4a98-9fc2-7959df4fee3f · inbound

Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference cites this paper.

Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:41:16.416573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T14:52:39.193613Z digest=sha256:1282d90041fe0df91d558152437bceb813169ed3236edf129fd6a5de2f9e3304

Observation 4147db7f-5dd0-4e91-8774-6ffb86da20e3 · inbound

VitaLLM: A Versatile, Ultra-Compact Ternary LLM Accelerator with Dependency-Aware Scheduling cites this paper.

VitaLLM: A Versatile, Ultra-Compact Ternary LLM Accelerator with Dependency-Aware Scheduling TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:51:27.301632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T09:05:12.937377Z digest=sha256:f37c9db8b00cdd049457d78feb4befdbd150712f60a46ee714433a9f823559f1

Observation a88e72cc-2c61-412c-afb2-98c499eb21ff · inbound

VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices cites this paper.

VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:43.970023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T19:06:18.309062Z digest=sha256:c998387307dda4784cc27b073fa55af2da3ded64262cbc34068921b787ba35f1