Pith. sign in

Paper Citation Record · LEDGER

Inference economics of language models

As of 19 August 2026, this Paper Citation Record lists 2 of 2 outbound references and 13 inbound Pith citation observations for arXiv:2506.04645.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04645 v1

Coverage vector

measured 2 of 2 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:43:48.474937Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:21:29.458359Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:57:51.165925Z

Reference resolution

2 of 2 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a6db3ea-157f-4845-9f1e-2529f6abccbe · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Inference economics of language models PaLM: Scaling Language Modeling with Pathways

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.385356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.385356Z digest=sha256:4ad529c6b0e49f162cb7a48c5785ff1bc0a192ee1efcac8275a20485cba2bdf1

Observation 35eb7c2d-c09b-4705-ad24-12896af4fd21 · outbound

This paper cites Fast Inference from Transformers via Speculative Decoding.

Inference economics of language models Fast Inference from Transformers via Speculative Decoding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.474937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.474937Z digest=sha256:a12ed507329d6c1ee8f2693fb7ca325a10d496e9507b423535b4d4343cbf1b94

Pith citing papers

Observation 2e1dac95-2868-46bb-aa47-e2a6dd506406 · inbound

Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches cites this paper.

Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches Inference economics of language models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:41:25.963890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T13:36:55.938673Z digest=sha256:45b362e7d68daeddbec27e54fc0f4b617764f4a0e2994a1a894ed1915074ab4a

Observation a72bb78f-99c1-4ada-aa8e-fcb66231a207 · inbound

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN cites this paper.

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN Inference economics of language models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:43:32.112169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T00:41:18.795730Z digest=sha256:84a2ddda77a85c15b3fd33e452b2269779d058d265ab2e4ef02b2e81c62ecfaf

Observation f2329126-04e4-4dd8-9b45-7dde365fa7df · inbound

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN cites this paper.

A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN Inference economics of language models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.200797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T16:54:58.406395Z digest=sha256:ae5c40756407cc46263a020cae8ba5e1c8abae2810f4a9e6a2e0a86075210155

Observation d9dea4af-5767-436c-aa02-d6ee6dc4ae9a · inbound

Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods cites this paper.

Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods Inference economics of language models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:38:14.423669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T20:37:49.558450Z digest=sha256:1ec28b789fb6d7402b384d060834d71ca268ffd682857f7bb42e2ca49f9fbc14

Observation e8c8a6e4-0338-4c8e-8f8e-2b827045b8b8 · inbound

Tokalator: A Context Engineering Toolkit for Artificial Intelligence Coding Assistants cites this paper.

Tokalator: A Context Engineering Toolkit for Artificial Intelligence Coding Assistants Inference economics of language models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:36:03.158798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:34:16.075235Z digest=sha256:958d222eb527cb0d3d6af9c228d9383383647957b7e0bc567809f34f0879b911

Observation 97493b5c-aa08-4af4-81b6-0f5dcdc1aa20 · inbound

The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development cites this paper.

The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development Inference economics of language models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:12.070833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T18:29:18.485336Z digest=sha256:6aaad555599cb0d7a6d1683106b5925477c5bf2f47c4d1ee852bee5324762ae6

Observation dd6cd589-49bd-437a-93f3-2fc4e4b46384 · inbound

Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference cites this paper.

Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference Inference economics of language models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:06:24.637089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T12:36:23.320194Z digest=sha256:f1dc462cfc0e9467c882622fd7ea99c5e3d587b0453297894d227d959b921e88

Observation 1adaf6bc-9594-4451-8151-2bf8b9e2e8e5 · inbound

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation cites this paper.

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation Inference economics of language models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:48:11.782979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T08:45:42.781160Z digest=sha256:bb4999fba42a4411b594c778ade1f0df1d62ddb6732b53b6ee79bb0c20712b7b

Observation 122dca54-a96d-4745-ba2a-23dba8e89d7b · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Inference economics of language models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.141713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:13d7278cbb210dc307f08e0b5276aa367f810e029f8e7378c8606788d9dfa17d

Observation d8fd2f96-9403-41d8-a630-009a86a578d4 · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Inference economics of language models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.196297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:1926579b4878b8b4b5e90eaa08cffcfe7396c75201580beb6133f9700c3ac9c3

Observation 987461ec-5ba4-481e-bfc2-a79665687c7f · inbound

Efficient Clustering with Provable Guardrails for LLM Inference at Scale cites this paper.

Efficient Clustering with Provable Guardrails for LLM Inference at Scale Inference economics of language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T12:03:18.785055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:03:18.785055Z digest=sha256:cb39589291a80bf9dc7f1f06d05d6b8b107e90840038fa314e47011e227dcf8f

Observation c8dc26db-bf8f-4bb6-b57e-7126535a1c92 · inbound

Scale Weight Decay and Train Better cites this paper.

Scale Weight Decay and Train Better Inference economics of language models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.155145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.155145Z digest=sha256:ea33320fe0666b13c56558197e97d1184ea062d0149234e2df841e88089c32df

Observation 9931379b-adf8-4bd2-87e6-95d5d15fe743 · inbound

NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems cites this paper.

NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems Inference economics of language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:21:29.458359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:21:29.458359Z digest=sha256:3b7898f5079f4dec7319522fed46f5da160a8fd0067753977441610174f85347