Pith. sign in

Paper Citation Record · LEDGER

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models

As of 22 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2509.05925.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.05925 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:53:40.274702Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:53:37.109195Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T04:53:40.793301Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4da74bc6-7662-495c-8f57-b1581587e833 · outbound

This paper cites Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:53:40.956069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:37.109195Z digest=sha256:7ec1f65dadb89c42073ad6ecc8831e5e79c3ea92e2401edf74e5e5bb3cac6722

Observation 3e4b8b92-64d7-45bf-ae39-e9e3b8a96ae0 · outbound

This paper cites In this paper, we exploit the CLIP model [10], which aligns images’ visual features with corresponding textual descriptions using contrastive learning.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models In this paper, we exploit the CLIP model [10], which aligns images’ visual features with corresponding textual descriptions using contrastive learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:45.779523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:37.212450Z digest=sha256:6ea0feb6288fc9696a5f1d809419afed3cfca7ccf21a02c88bceffb760d87691

Observation b930a5d7-96a3-4c70-88c8-709ac71c5bb9 · outbound

This paper cites Overall architecture The proposedPQVAE-sharedscheme for feature compression is illustrated in Fig.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Overall architecture The proposedPQVAE-sharedscheme for feature compression is illustrated in Fig

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:45.460737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:37.365495Z digest=sha256:b260147dfd3ca703b98f663f215f0f7562401b76fdbff9f7971fcab302b36bf4

Observation 98ba6847-84da-4c22-837f-686021cfa1aa · outbound

This paper cites ViT-L/14@336px.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models ViT-L/14@336px

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:45.157943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:37.545821Z digest=sha256:539df159fcdbe3cff454c7f3094933ad1fcfa5db467ead4ce0e01ddb9e4eaf15

Observation 49c524aa-4c30-4682-96c8-ff1e1e88ae93 · outbound

This paper cites Experiments demonstrate that original CLIP features can be compressed more than 30-fold while maintaining satisfactory semantic preservation.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Experiments demonstrate that original CLIP features can be compressed more than 30-fold while maintaining satisfactory semantic preservation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:44.881018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:37.732831Z digest=sha256:ea69dc50447eec60d3122b3e7219fc2bc4b27e8ede0b4526207d786a5c583930

Observation 4c18d651-de56-42e2-948b-856674ad208a · outbound

This paper cites Variational image compres- sion with a scale hyperprior,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Variational image compres- sion with a scale hyperprior,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:44.538181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:37.846515Z digest=sha256:668a297147c32b157fda16756094bd57ac9c9e44da7c4d8c2a63a760604a353b

Observation c30253b3-e511-40cb-ae82-a0938a3096a3 · outbound

This paper cites Learned image compression with discretized gaussian mixture likelihoods and attention modules,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Learned image compression with discretized gaussian mixture likelihoods and attention modules,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:44.229520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:37.990074Z digest=sha256:0409394630138fe95553b0cbaa83ef90dae39c1eabe3a202bae18d9ad583834d

Observation 0faa6861-b94f-4fd5-8f15-ff29a2154ebb · outbound

This paper cites LotteryCodec: Searching the implicit representation in a random network for low-complexity image compression,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models LotteryCodec: Searching the implicit representation in a random network for low-complexity image compression,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:43.937660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:38.128128Z digest=sha256:8f2cf4b96bf81572d26d82abee0a22baae5da938e8c8c5a6052bc1831e319579

Observation 79f7c581-45c2-4254-93f9-a973395c26ef · outbound

This paper cites DiffCP: Ultra-low bit collaborative perception via diffusion model,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models DiffCP: Ultra-low bit collaborative perception via diffusion model,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:43.585584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:38.325003Z digest=sha256:3474e8b16168865eb01e4c2e8e8d8652b4628070ee22f35c5102d5bfdc922798

Observation 3f440b27-71b5-44d2-baa4-fd5054b0754a · outbound

This paper cites Edge computing with artifi- cial intelligence: A machine learning perspective,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Edge computing with artifi- cial intelligence: A machine learning perspective,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:43.228545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:38.471273Z digest=sha256:d961904e06f788b26cfcfac24b123e520a5831da096efaaed9fc8a9e79990295

Observation 8344dd89-30d4-4e30-9f33-5b5450bedc06 · outbound

This paper cites Zero-Shot Semantic Communication with Multimodal Foundation Models.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Zero-Shot Semantic Communication with Multimodal Foundation Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:53:40.618283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:38.668546Z digest=sha256:5be1ac822bd8f30effd4db2476780ea98709b373b3c5c8231dea0b20e4c6e764

Observation 5f18c227-aa9e-44d9-8dd5-6fea3e89f022 · outbound

This paper cites Beyond transmitting bits: Context, semantics, and task-oriented communica- tions,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Beyond transmitting bits: Context, semantics, and task-oriented communica- tions,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:42.994790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:38.799433Z digest=sha256:d64ab1318e0d6d1bb28d93b29d16782d8005e65e24f8ae8491e40d7640ed6b1c

Observation 726ecdf7-0423-452b-bb21-9a7230df89af · outbound

This paper cites DeepSIC: Deep seman- tic image compression,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models DeepSIC: Deep seman- tic image compression,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:42.704977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:38.928515Z digest=sha256:98b336fc68c9d3712c554a538906eadb11dc7730eb81159be42d669fc47f2e2d

Observation d44ca2ab-d800-49d6-bc68-a7c436f92386 · outbound

This paper cites Semantic-aware video compression for automotive cameras,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Semantic-aware video compression for automotive cameras,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:42.441709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:39.116202Z digest=sha256:60c8b2298cf86407b5da507918b4e06494f35cc8af1269a150531a042a7ba6b2

Observation bd65ecac-78b6-4594-bdb6-76471a14d328 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Learning transferable visual models from natural lan- guage supervision,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:39.260529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:39.260529Z digest=sha256:4b389385c6c834ec2dd61fedfcc6cf639fc2d06b3d8c47461e9bc26465320e46

Observation 08a3f453-71aa-47f5-8cdb-67d81a778822 · outbound

This paper cites Visual instruction tuning,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Visual instruction tuning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:42.182134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:39.427350Z digest=sha256:b2d7c86bd3cf0945ef4cd01974e25e2d8133049d1f1394c0ed451a5d19239874

Observation 45688dfc-2174-42b4-b4c8-c112b8b33677 · outbound

This paper cites Neural dis- crete representation learning,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Neural dis- crete representation learning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:41.912848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:39.607838Z digest=sha256:9792c8c26db80763776c0602c9e125993baec8848a127ed5b1420c5a678892f4

Observation 56b51f31-9b2b-414a-8f29-7f39a4720ab9 · outbound

This paper cites Learning product code- books using vector-quantized autoencoders for image re- trieval,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Learning product code- books using vector-quantized autoencoders for image re- trieval,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:41.541895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:39.802706Z digest=sha256:344eb6a2eb5f386e910490885929e33416fc9822cf38ff4be7255d69daeb2d08

Observation 767f95a8-06b9-401a-92d3-e93b484ed41f · outbound

This paper cites Image coding for machines with edge information learning using segment anything,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Image coding for machines with edge information learning using segment anything,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:53:41.188624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:39.940748Z digest=sha256:bd61d9d59c7c4c3659713684c3fb1147c255010362402964128d899cf67bdf28

Observation cdb27136-dcf8-43a4-98ed-2f7bc45059c5 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models ClipCap: CLIP Prefix for Image Captioning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:40.077127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:40.077127Z digest=sha256:0f7d2bb506b4c1b389d747457ce017510636f3885c1c3b898b585ed31866bbd2

Observation 070a47ef-e630-4fa9-8a2c-d9f9301890bd · outbound

This paper cites Segment anything,.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Segment anything,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:40.274702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:40.274702Z digest=sha256:8a930ecd1a3bef0a640004438896a65c024533e31373cc76595c01d66c57b6ea

Pith citing papers

Observation 4da74bc6-7662-495c-8f57-b1581587e833 · inbound

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models cites this paper.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:53:40.956069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T04:53:37.109195Z digest=sha256:7ec1f65dadb89c42073ad6ecc8831e5e79c3ea92e2401edf74e5e5bb3cac6722