Pith. sign in

Paper Citation Record · LEDGER

VeS: Teaching Pixels to Listen Without Supervision

As of 12 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2507.22008.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22008 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:11:49.190965Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d79d9b3-fd12-48c7-8530-7ee7bc8526b9 · outbound

This paper cites write newline.

VeS: Teaching Pixels to Listen Without Supervision write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:11:49.153861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:11:49.153861Z digest=sha256:57f465429fc3d5e6634a8c7df710b9efcca070681dc1378755542657d3ca2e13

Observation 9ffb9103-45f8-4a9b-8619-8164c84f4b08 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

VeS: Teaching Pixels to Listen Without Supervision wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.305108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.158237Z digest=sha256:d10da0ab31af2bcdb8c1ab20cea353da45a18447c963a2788f9ccd0d0d4f1a33

Observation b7396ad5-a7e3-40c5-9ab8-09a77b1fbc98 · outbound

This paper cites DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden-unit BERT.

VeS: Teaching Pixels to Listen Without Supervision DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden-unit BERT

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:11:49.161523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:11:49.161523Z digest=sha256:7d968f88926d0505b9750727dcc5f5116d872df3ecb5764cdacb2e39cc42b2c3

Observation 7eff41cd-a1ea-450e-998f-cc800ad4779a · outbound

This paper cites chirp" from the.

VeS: Teaching Pixels to Listen Without Supervision chirp" from the

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.295297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.164942Z digest=sha256:4f1e1a640567ed640e38e32400dfe909c80963f385135ac4fc08885193d2d0f8

Observation d9a7fc29-d51d-4283-962d-a24acd71215c · outbound

This paper cites Project vaani (huggingface dataset).

VeS: Teaching Pixels to Listen Without Supervision Project vaani (huggingface dataset)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.285721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.168255Z digest=sha256:a88a525d60aa7846beb0cc7c064499d927874f85d0f112da0235f9001ebe7d2d

Observation 4df5d955-ad3a-4b3d-9c9a-5e0b45218d97 · outbound

This paper cites Project vaani.

VeS: Teaching Pixels to Listen Without Supervision Project vaani

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.274559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.171434Z digest=sha256:066641e8b7e1b2d2a914fb856677ef49221a5c01ac90c9e3db9bd7d9defb3742

Observation bbdd71e6-9a7c-4701-a062-80cc9f3c674b · outbound

This paper cites Vo, Patrick Labatut, and Piotr Bojanowski.

VeS: Teaching Pixels to Listen Without Supervision Vo, Patrick Labatut, and Piotr Bojanowski

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.265612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.174798Z digest=sha256:e0d68095e482aea822671a0f4fea9f9ddb9a7cf53ae1894842cdac66e6590780

Observation 68e10076-8d75-4e86-8417-2c27eba76b15 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VeS: Teaching Pixels to Listen Without Supervision DINOv2: Learning Robust Visual Features without Supervision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:11:49.178168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:11:49.178168Z digest=sha256:1775296581e7f99f3b58a32406be51a569161b2687753ff8263ad56b9f165dd3

Observation f35682fe-08d5-49b1-a949-def07971cdd1 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VeS: Teaching Pixels to Listen Without Supervision Learning transferable visual models from natural language supervision

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.256044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.181378Z digest=sha256:e207f4451aeae8bf8f052eb6d55100afce5f265a8017907c666b476b1f111bc6

Observation b4d4be16-a38c-4ff1-9b15-04a265516705 · outbound

This paper cites Filip: Fine-grained interactive language-image pre-training.

VeS: Teaching Pixels to Listen Without Supervision Filip: Fine-grained interactive language-image pre-training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.246621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.187601Z digest=sha256:8698738a0b2529a93dc21203ca14c57f696f761ef3698b3ab73e87aae600e038

Observation 5eb1de39-ff1c-4f5b-84fc-47a35be25fc3 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

VeS: Teaching Pixels to Listen Without Supervision Lit: Zero-shot transfer with locked-image text tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.236473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.190965Z digest=sha256:0b3e7b7c37d857eba903be49155f611a3490653c0cb28b5e716664470c1832ca

Pith citing papers

No inbound Pith citation observations are available.