Pith. sign in

Paper Citation Record · LEDGER

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation

As of 10 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2608.04902.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04902 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:02:31.138343Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45471de0-2efc-470e-b0c4-ccd50d499613 · outbound

This paper cites The Million Song Dataset.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation The Million Song Dataset

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.499333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.067776Z digest=sha256:deeb629c9c8f4c9505efe8c9eb3d70054ee0204c518e8080237c89c8e6eddf60

Observation 02dce2a7-0f15-4d4f-9322-65f13d136b38 · outbound

This paper cites InICCV, 4195–4205.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation InICCV, 4195–4205

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.429678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.094932Z digest=sha256:96eb627c07794f20ef2d3c468693ba16863ba5dac40005007aa28eb625bc41d4

Observation dfdb8c95-54b5-42cc-bb6f-9ceff5005d74 · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:02:31.099756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:02:31.099756Z digest=sha256:695038d9620c8c7cffa695a8a47c828d08a661df2e0739015d93781e2b2dc615

Observation 2b52b8b3-99e7-4d01-add4-fafc8a8e1fea · outbound

This paper cites InICASSP, 1–5.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation InICASSP, 1–5

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.414717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.104491Z digest=sha256:b572cceb426ac8bb2e74efbaf9c28c3a52e1497d06527da423fb7df0a5ad997f

Observation 3945e800-7001-402d-8428-18eb7951b55e · outbound

This paper cites an unresolved cited work.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:02:31.386492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.112670Z digest=sha256:55b6c63d1af2bc30b7b6ab8cd40ef218f054d09123fc88d31991ac998d8bcaf1

Observation 6c0ff6be-6ef1-4bc6-9df1-5f6eb67ecdd7 · outbound

This paper cites InICML, 62596–62626.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation InICML, 62596–62626

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.371486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.117017Z digest=sha256:e576abace4b55849f2180065e66171f5eb7e35f43c81b634d78ddaeffe55f5f6

Observation 66eed211-a3b8-45f9-88d3-94138eb509cb · outbound

This paper cites Video-to-Audio Generation with Hidden Alignment.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation Video-to-Audio Generation with Hidden Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:02:31.121546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:02:31.121546Z digest=sha256:2a89ce164d6f34bcd886a5c46adc48ae879838d2e951cc35f2c4f947236a4f15

Observation 6a24dab8-44da-4224-8599-e384618994c1 · outbound

This paper cites baby crying.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation baby crying

Reference 17

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T14:02:31.310949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.138343Z digest=sha256:2f7d92af97b3711f118f5bacfa01749266f25911716478b294c5a156d52c90f4

Observation c311c4f5-a012-47b8-ba7a-5376eeac1565 · outbound

This paper cites Snake ac- tivation functions are applied throughout the network, and no final tanh activation is used in the decoder.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation Snake ac- tivation functions are applied throughout the network, and no final tanh activation is used in the decoder

Reference 640

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.326121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.134405Z digest=sha256:e7e41bbb42d29ff2cd6b074160a4fb61e29a1f33a3ea8214200bc48adf5a1999

Observation dd7e5510-caa3-461c-95c7-a3634daa1938 · outbound

This paper cites K.; Ishii, M.; Hayakawa, A.; Shibuya, T.; Schwing, A.; andMitsufuji,Y.2025.MMAudio:TamingMultimodalJointTrain- ing for High-Quality Video-to-Audio Synthesis.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation K.; Ishii, M.; Hayakawa, A.; Shibuya, T.; Schwing, A.; andMitsufuji,Y.2025.MMAudio:TamingMultimodalJointTrain- ing for High-Quality Video-to-Audio Synthesis

Reference 725

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.485720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.072369Z digest=sha256:7cb28285e8215bc4a0b7d0d73fa98926a8957f6b6ef497abab3f07a5e13f8635

Observation 6f2360e0-6179-4b73-99c8-0f77d5b01876 · outbound

This paper cites Dataset Task Hours (h) Source AudioCaps T2A 109 (Kim et al.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation Dataset Task Hours (h) Source AudioCaps T2A 109 (Kim et al

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.341681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.130418Z digest=sha256:94a71495269b247902edfed6d9271731c0873c20e6cdfaf34bc5cf3ddca96c61

Observation e5e2d683-b662-418a-8a5c-cffad615da28 · outbound

This paper cites In ICASSP, 776–780.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation In ICASSP, 776–780

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.458005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.086327Z digest=sha256:bcbda49db1b3df1cfd5295f885cb555f24648a90c8bdeebc0666c3153a6013ad

Observation 2f9cbe26-a44d-4c5a-8ffd-a1aca9ac9a19 · outbound

This paper cites InECCV, 570–586.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation InECCV, 570–586

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.356595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.126304Z digest=sha256:6005ac490dddfa98b9d5dc4c1c416c64e6ab7f5174c504cd764b2e49b5e8a2be

Observation 301fe667-d513-4690-8b29-1bcc1ac6b95f · outbound

This paper cites InICML, 21450–21474.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation InICML, 21450–21474

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.443727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.090647Z digest=sha256:5a8e63c39d5bf8f2f7202ebd2993f11f47b7bb10b704ebac90a72ed01b1988d7

Observation 5de61743-68ed-4801-ae66-0c8a82a90731 · outbound

This paper cites InICML, 46804–46822.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation InICML, 46804–46822

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:02:31.400600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.108705Z digest=sha256:4d68e00c2537f817e414e6e2f2a07a78c2ce3ee6e27a523beb303ad6ac1da4e4

Observation 7e551071-0bbc-4330-9fb5-20bc92722f03 · outbound

This paper cites an unresolved cited work.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:02:31.471592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.081777Z digest=sha256:bb28c38b40e809df41cf075c98db260e436f16269431c3d4e9799c9acd7fd257

Observation 3cdb3f24-129f-4066-a256-8e4cced7f829 · outbound

This paper cites arXiv:2606.15956.

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation arXiv:2606.15956

Reference 2026

Resolution
verified exact
raw_fallback, observed 2026-08-06T14:02:31.296477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:02:31.076800Z digest=sha256:05d805bb0f8dd47b2a51946db7e315f00d9077dddf11f119da89679174c91ec2

Pith citing papers

No inbound Pith citation observations are available.