Pith. sign in

Paper Citation Record · LEDGER

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy

As of 14 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.07427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07427 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:51:16.342069Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e86c4175-46a7-4286-b62e-f555d045ec46 · outbound

This paper cites Can US infrastructure keep up with AI economy?.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Can US infrastructure keep up with AI economy?

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.682845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T04:51:16.250855Z digest=sha256:31202f9676f1c8f7e5df489a120da0b22c7c709c82e0ced5aef29e794f979e02

Observation 7d0ca9c2-e7cd-4c01-917e-2319287b237d · outbound

This paper cites The future of datacenters,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy The future of datacenters,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.669592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T04:51:16.255689Z digest=sha256:aacfab08086dc1864f9b02c8382548e9ddee6b52bb5fbf1ffd6c77148de19cd4

Observation fc289954-9359-4e9d-9304-67464b88dc4a · outbound

This paper cites How hungry is AI? Benchmarking energy, water and carbon footprint of LLM inference,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy How hungry is AI? Benchmarking energy, water and carbon footprint of LLM inference,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.260016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.260016Z digest=sha256:1f150036b2d7e70211d63eaafe590ac0684d648eba7506683a3905e7171cccc9

Observation dd8b3348-8555-4d0f-bf7d-735982898ac6 · outbound

This paper cites Artificial intelligence and energy overcon- sumption: Data center electricity demand, cooling burdens, and regional sustainability constraints,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Artificial intelligence and energy overcon- sumption: Data center electricity demand, cooling burdens, and regional sustainability constraints,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.655713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T04:51:16.264462Z digest=sha256:6c60d5a9c9ee51995565d1b1b5d5a345bf7bc65f68f3a585d04699ac8ab5bab8

Observation bbc5bbef-a87a-45da-b37f-69c669e01afd · outbound

This paper cites Comparative analysis on developed optimization techniques for reducing energy consumption in AI training and inference,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Comparative analysis on developed optimization techniques for reducing energy consumption in AI training and inference,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.640829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T04:51:16.268990Z digest=sha256:186c5b9c4acaf97b038a4253a7154401a50f238e7dce190f74f0977f8da617dc

Observation bbb05f71-675b-423b-a571-ba990905c72a · outbound

This paper cites From prompts to power: Measuring the energy footprints of LLM inference,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy From prompts to power: Measuring the energy footprints of LLM inference,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.273744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.273744Z digest=sha256:7e5e69e5cb44388a0c948e4e1caa1370adb064cb45bfe8be00b59aee6c62df8c

Observation 1aeb037a-8a26-4336-9779-2a5426fc744a · outbound

This paper cites Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.278459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.278459Z digest=sha256:0b1a754edb44d7d0d7ae64b0a7f6cb1da771e047a5e6d184347e367c9a3831b1

Observation e30abeef-7af0-4cce-90e5-89d990b2da18 · outbound

This paper cites How different tokenization algo- rithm impact LLMs and transformer models for binary code analysis,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy How different tokenization algo- rithm impact LLMs and transformer models for binary code analysis,

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-08-10T04:51:17.217901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T04:51:16.283034Z digest=sha256:05e80d4ec0e062bd1dbee8825a5f27e1be52223078fb954400177e0fc1c06252

Observation b4287e23-e609-4ad9-9b7d-6c048b1df8f8 · outbound

This paper cites Harnessing vision-language models for time series anomaly detection,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Harnessing vision-language models for time series anomaly detection,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.286983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.286983Z digest=sha256:cb46d54d8be5c1146a507112f92dbd4d6d6f06c189db8c1e2725b6232d01e5e2

Observation 0d2f52fe-fbe6-4f04-a526-cada29527f67 · outbound

This paper cites TokenPowerBench: Benchmarking the power consumption of LLM inference,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy TokenPowerBench: Benchmarking the power consumption of LLM inference,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.291049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.291049Z digest=sha256:e6debaafd92e5ec71e520b17c17b4573fb75bf6e790eb8c35816837d7e530ed1

Observation cf8957cc-a73d-4287-94c5-1b4f3a91685f · outbound

This paper cites The ML.ENERGY benchmark: Toward automated inference energy measurement and optimization,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy The ML.ENERGY benchmark: Toward automated inference energy measurement and optimization,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.294955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.294955Z digest=sha256:5f8b5669b57ec5a7d862c78e1b439cb5645c624c41af4b19913df608d82a1489

Observation fc361cb8-1430-4d56-8148-8ea9cdff5c19 · outbound

This paper cites An energy-efficient vision language model inference with importance-aware token pruning,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy An energy-efficient vision language model inference with importance-aware token pruning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.626613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T04:51:16.298990Z digest=sha256:ef9316d5b3cd8707025115c0bd952bf95009b3a6976ff02b4ef948ec1f97884e

Observation 92c823a5-1470-42e2-8e75-c50bdc24a828 · outbound

This paper cites Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.303359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.303359Z digest=sha256:b233a058e8bcbd53ce347424342f87fee24300091d02a1043202d482af76ce06

Observation f3381d6d-fe4e-4a4e-9902-ea80c3f64b56 · outbound

This paper cites A Picture is Worth A Thousand Numbers: Enabling LLMs Reason about Time Series via Visualization.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy A Picture is Worth A Thousand Numbers: Enabling LLMs Reason about Time Series via Visualization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.308177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.308177Z digest=sha256:f5aa576d4e875a888382f9aca128af5705ac1ab001ed08ed09355832043787ad

Observation 23f19057-c6ed-4369-8bdb-6a68b50353a2 · outbound

This paper cites Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:51:16.426717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T04:51:16.312840Z digest=sha256:07e640c77e84f346e2f53ed9817c0d9cced083b1d649fa642b22cf071c4bf994

Observation f42fc48d-43db-4062-b758-6986cc746f6d · outbound

This paper cites Zeus: Understanding and optimizing GPU energy consumption of DNN training,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Zeus: Understanding and optimizing GPU energy consumption of DNN training,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.612703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T04:51:16.317365Z digest=sha256:f7552be4e4bd6fba5a03bc62211526c134abcdfd6e7813920d12b36ed5b895a1

Observation d5f6ba40-85f2-42cf-acd7-1412f693b57d · outbound

This paper cites The Llama 3 Herd of Models.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy The Llama 3 Herd of Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.323125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.323125Z digest=sha256:ae9a753ef274d6654683d95b4474e176ded05b20c93a44ed631c7237ca3f0132

Observation be40f393-fb5b-4a06-a150-58204cb18b2e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.328390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.328390Z digest=sha256:dc938bc147375f076ccbf427316fa1a4825c48d40dd618f7e54977b105737958

Observation 8c5fe114-94a8-444f-93aa-c622da1dfad6 · outbound

This paper cites Pixtral 12B.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Pixtral 12B

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:51:16.333215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:51:16.333215Z digest=sha256:021edda3763b6a315f84a2e39bc5f34c24131f6463f08a648d55625f8ccc1663

Observation a8f2d798-5472-4240-b21b-e5888612a969 · outbound

This paper cites Causal inter- vention sequence analysis for fault tracking in radio access networks,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Causal inter- vention sequence analysis for fault tracking in radio access networks,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.598284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T04:51:16.337747Z digest=sha256:44a00d5a290410caecb0380fc240da13d4f368c48a506a9c0ab8775effe66686

Observation bda76dd2-3446-4364-821f-8e760aa0dadd · outbound

This paper cites Cost efficient GPU cluster management for training and inference of deep learning,.

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Cost efficient GPU cluster management for training and inference of deep learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:51:17.583501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T04:51:16.342069Z digest=sha256:ab33288be9afb80fe04c9cdb8187bbec171006549203e760ef684551d79c36bf

Pith citing papers

No inbound Pith citation observations are available.