Pith. sign in

Paper Citation Record · LEDGER

Inference economics of language models

As of 22 August 2026, this Paper Citation Record lists 2 of 2 outbound references and 13 inbound Pith citation observations for arXiv:2506.04645.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04645 v1

Coverage vector

measured 2 of 2 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:43:48.474937Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:21:29.458359Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:57:51.165925Z

Reference resolution

2 of 2 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a6db3ea-157f-4845-9f1e-2529f6abccbe · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Inference economics of language models PaLM: Scaling Language Modeling with Pathways

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.385356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.385356Z digest=sha256:b56e5ad12895389375da64711288caca08775034e3632a66e10ec736591ee6b4

Observation 35eb7c2d-c09b-4705-ad24-12896af4fd21 · outbound

This paper cites Fast Inference from Transformers via Speculative Decoding.

Inference economics of language models Fast Inference from Transformers via Speculative Decoding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.474937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.474937Z digest=sha256:14ba1035d372e3d6b6913c4fe2f715225016f8938f87ceb83b400c8da44bf164

Pith citing papers

Observation 2e1dac95-2868-46bb-aa47-e2a6dd506406 · inbound

Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches cites this paper.

Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches Inference economics of language models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:41:25.963890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T13:36:55.938673Z digest=sha256:e1de186fe347de131aaa976b389ac7ef3ef47136e3bceab24fabe136def3d116

Observation a72bb78f-99c1-4ada-aa8e-fcb66231a207 · inbound

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN cites this paper.

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN Inference economics of language models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:43:32.112169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T00:41:18.795730Z digest=sha256:c4fa7285426548ffccedd0eed6051171f044ff6c0d8a8f4127e4d3ae0ba60ae3

Observation f2329126-04e4-4dd8-9b45-7dde365fa7df · inbound

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN cites this paper.

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN Inference economics of language models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.200797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T16:54:58.406395Z digest=sha256:3ed9a39c128ddd86564815d8cc9ad3e675cf3e4cfd4430afc687b78ff853e257

Observation d9dea4af-5767-436c-aa02-d6ee6dc4ae9a · inbound

Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods cites this paper.

Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods Inference economics of language models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:38:14.423669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T20:37:49.558450Z digest=sha256:7e7829e8f3182966fa75cf19bb3e5f9d55da33c304beacc91174e8a631b8796d

Observation e8c8a6e4-0338-4c8e-8f8e-2b827045b8b8 · inbound

Tokalator: A Context Engineering Toolkit for Artificial Intelligence Coding Assistants cites this paper.

Tokalator: A Context Engineering Toolkit for Artificial Intelligence Coding Assistants Inference economics of language models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:36:03.158798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:34:16.075235Z digest=sha256:6746c085a9a38bcd314ecf82610ae5b7fcf5904d98f0ccfde0fa394f5298631b

Observation 97493b5c-aa08-4af4-81b6-0f5dcdc1aa20 · inbound

The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development cites this paper.

The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development Inference economics of language models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:12.070833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T18:29:18.485336Z digest=sha256:5e5bec7556f9c66b19ca3b988abb55e8a2de2aca49ad3bcb4799f53827f593d8

Observation dd6cd589-49bd-437a-93f3-2fc4e4b46384 · inbound

Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference cites this paper.

Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference Inference economics of language models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:06:24.637089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T12:36:23.320194Z digest=sha256:0553956f37235fe85254ed57e8c9860a708e28ab00b4de1ebb6c120499d01228

Observation 1adaf6bc-9594-4451-8151-2bf8b9e2e8e5 · inbound

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation cites this paper.

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation Inference economics of language models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:48:11.782979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T08:45:42.781160Z digest=sha256:6a75f89f4009ffb5f17c1437324eb9687995083fa3dab071c1109c9dcde35848

Observation 122dca54-a96d-4745-ba2a-23dba8e89d7b · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Inference economics of language models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.141713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:694c1ce87299f38e9c5b2e78f70cdf936d978a179987c8a64ca76e4b5a353a1e

Observation d8fd2f96-9403-41d8-a630-009a86a578d4 · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Inference economics of language models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.196297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:146fe79ff14268893f19d52a5ead43c7aa9aac9cfe09fbb026a62595fc0070c4

Observation 987461ec-5ba4-481e-bfc2-a79665687c7f · inbound

Efficient Clustering with Provable Guardrails for LLM Inference at Scale cites this paper.

Efficient Clustering with Provable Guardrails for LLM Inference at Scale Inference economics of language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:03:18.785055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:03:18.785055Z digest=sha256:036b1a41f1a18b2239cd7f3588fe1b84339bd8f1b612b0eb2c1186ae859c589b

Observation c8dc26db-bf8f-4bb6-b57e-7126535a1c92 · inbound

Scale Weight Decay and Train Better cites this paper.

Scale Weight Decay and Train Better Inference economics of language models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.155145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.155145Z digest=sha256:56e2f8be0e3efad7b857cc3b4b5996db58ef60f79b1adbfaa02e7087bb6262f7

Observation 9931379b-adf8-4bd2-87e6-95d5d15fe743 · inbound

NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems cites this paper.

NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems Inference economics of language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:21:29.458359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:21:29.458359Z digest=sha256:05d84dc6b269c7ee9a6c239a022885db8653c835ad1dc60b2473a208da463e21