Pith. sign in

Paper Citation Record · LEDGER

VeS: Teaching Pixels to Listen Without Supervision

As of 12 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2507.22008.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22008 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:11:49.190965Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d79d9b3-fd12-48c7-8530-7ee7bc8526b9 · outbound

This paper cites write newline.

VeS: Teaching Pixels to Listen Without Supervision write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:11:49.153861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:11:49.153861Z digest=sha256:57f465429fc3d5e6634a8c7df710b9efcca070681dc1378755542657d3ca2e13

Observation 9ffb9103-45f8-4a9b-8619-8164c84f4b08 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

VeS: Teaching Pixels to Listen Without Supervision wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.305108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.158237Z digest=sha256:6931046286a362181a53643de02d8631af31153934372b8696396070331cc58b

Observation b7396ad5-a7e3-40c5-9ab8-09a77b1fbc98 · outbound

This paper cites DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden-unit BERT.

VeS: Teaching Pixels to Listen Without Supervision DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden-unit BERT

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:11:49.161523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:11:49.161523Z digest=sha256:7d968f88926d0505b9750727dcc5f5116d872df3ecb5764cdacb2e39cc42b2c3

Observation 7eff41cd-a1ea-450e-998f-cc800ad4779a · outbound

This paper cites chirp" from the.

VeS: Teaching Pixels to Listen Without Supervision chirp" from the

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.295297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.164942Z digest=sha256:73ab5cd54cfd649032f79cb4ec1061c2fa3a1ace8a0ce8c43794763306943fc9

Observation d9a7fc29-d51d-4283-962d-a24acd71215c · outbound

This paper cites Project vaani (huggingface dataset).

VeS: Teaching Pixels to Listen Without Supervision Project vaani (huggingface dataset)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.285721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.168255Z digest=sha256:38866804c69ca7f50dfb417d4248cb6cf6e9015e5fb7baac34ee1b6c085d476f

Observation 4df5d955-ad3a-4b3d-9c9a-5e0b45218d97 · outbound

This paper cites Project vaani.

VeS: Teaching Pixels to Listen Without Supervision Project vaani

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.274559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.171434Z digest=sha256:dfc0a9e8a8b4cbbea983a19ffcbe5acc98abfe0468e10e9567823172c1ee77c2

Observation bbdd71e6-9a7c-4701-a062-80cc9f3c674b · outbound

This paper cites Vo, Patrick Labatut, and Piotr Bojanowski.

VeS: Teaching Pixels to Listen Without Supervision Vo, Patrick Labatut, and Piotr Bojanowski

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.265612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.174798Z digest=sha256:32ba699f510dc7774371f8d412f98fd342dedda62022962271f4b6f7e1fd41b3

Observation 68e10076-8d75-4e86-8417-2c27eba76b15 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VeS: Teaching Pixels to Listen Without Supervision DINOv2: Learning Robust Visual Features without Supervision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:11:49.178168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:11:49.178168Z digest=sha256:1775296581e7f99f3b58a32406be51a569161b2687753ff8263ad56b9f165dd3

Observation f35682fe-08d5-49b1-a949-def07971cdd1 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VeS: Teaching Pixels to Listen Without Supervision Learning transferable visual models from natural language supervision

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.256044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.181378Z digest=sha256:73f23198e50de9ad040957847b99ce2c4f2a3ba60f9ef2a9ed126b48fc0eac53

Observation b4d4be16-a38c-4ff1-9b15-04a265516705 · outbound

This paper cites Filip: Fine-grained interactive language-image pre-training.

VeS: Teaching Pixels to Listen Without Supervision Filip: Fine-grained interactive language-image pre-training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.246621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.187601Z digest=sha256:59aea2be7fec00ded8806fd60b377084bc5c958f2f28f762636dad3fbed32fdf

Observation 5eb1de39-ff1c-4f5b-84fc-47a35be25fc3 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

VeS: Teaching Pixels to Listen Without Supervision Lit: Zero-shot transfer with locked-image text tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:11:49.236473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-06T12:11:49.190965Z digest=sha256:5cd45159113cc8463b6bb54e824b2e1eee6c61a27f8bbf6511fb507ef5630ccb

Pith citing papers

No inbound Pith citation observations are available.