Pith. sign in

Paper Citation Record · LEDGER

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters

As of 17 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2508.11074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11074 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:43:11.904898Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e78136a4-2488-442c-8a20-46682029c1ca · outbound

This paper cites write newline.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.806244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.806244Z digest=sha256:2f9b8e45fc942f10d01fe2facb13ade286b97267c055b0dbc5337286dd2ad7ce

Observation 4ae74a23-daf0-4b98-9b18-944e3de46d5f · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Vggsound: A large-scale audio-visual dataset

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:43:12.211198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:43:11.812578Z digest=sha256:91fa4c88bcf6ebac3dd68bdd26136e6b84a9d90a087706de7f962639d1d49420

Observation d9381e5b-49e4-451c-b97a-762ff65b9d3b · outbound

This paper cites Video-Guided Foley Sound Generation with Multimodal Controls.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Video-Guided Foley Sound Generation with Multimodal Controls

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.816985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.816985Z digest=sha256:0518cd7893a39b8c8884fcbd289cfde58bb6e34cdabb32864b9ad74ecf7a3ee2

Observation 349471a0-cb46-4725-b636-b734422574e2 · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.821697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.821697Z digest=sha256:20277f714c18f27d8817e039f66c2a14bbb8d467a9908c62d92b0894c6c62149

Observation dcc35b06-a553-4e38-8fcd-f0645061720e · outbound

This paper cites LoVA: Long-form Video-to-Audio Generation.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters LoVA: Long-form Video-to-Audio Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.829975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.829975Z digest=sha256:ec0e934410d46b299235dde481f57cd4a0351e0abe510f3056959466db3d6b83

Observation 58fcf3ef-a609-4ed0-8a0c-d49c3eb76351 · outbound

This paper cites $^R$FLAV: Rolling Flow matching for infinite Audio Video generation.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters $^R$FLAV: Rolling Flow matching for infinite Audio Video generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.834522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.834522Z digest=sha256:64e33754bc4f9a8a387944ea33564dc78c0926b89c5a920ded525f89bf13a0aa

Observation ebf64fa7-471c-4894-9ad7-ee90ce4071a9 · outbound

This paper cites SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.839045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.839045Z digest=sha256:e15f57f57bd49ca949b3a3b08e9b501b409a264fc996b402ab10b589fe24358e

Observation ecd00d11-a5b1-43fa-9782-87994a77c396 · outbound

This paper cites Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:43:12.197816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:43:11.844187Z digest=sha256:8d8e04fd77043fe542972845a636d478fcf58f3a470ed6050f2793a70a0bd650

Observation 1a92ee21-30e9-4c54-b588-93e1fa8c8a75 · outbound

This paper cites Taming visually guided sound generation.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Taming visually guided sound generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:43:12.184383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:43:11.848386Z digest=sha256:5d729caef2e9118c5676ec351f1dcb114a5a20755e1828c7c32fc86b8d9f941c

Observation da7961a6-da39-420d-ac9c-e7920d23a809 · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Synchformer: Efficient synchronization from sparse cues

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:43:12.170363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:43:11.852676Z digest=sha256:de5cff82c6803f957934b104dda391aa1d6a411d08e90fb9676dc3ce381d2076

Observation d7254799-242c-4f6d-b964-b468697508a8 · outbound

This paper cites Sophia Koepke, Olivia Wiles, Yael Moses, and Andrew Zisserman.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Sophia Koepke, Olivia Wiles, Yael Moses, and Andrew Zisserman

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:43:12.154980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:43:11.856889Z digest=sha256:685a79e3a7c2fb197dda126e62ccfc08f8fd698b2f9f4ac9fdfb1afa3d02d449

Observation e919d59f-57ec-4852-8490-837cf7b98d77 · outbound

This paper cites Crandall, and Christopher Raphael.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Crandall, and Christopher Raphael

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:43:12.140815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:43:11.861211Z digest=sha256:1f395b57fe0d401298506669913a456fdec5297cd8db513bb1207bc9088bb24b

Observation 899315b1-22f0-4d7b-9a63-3ea2b9173a07 · outbound

This paper cites Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.865655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.865655Z digest=sha256:1534ef59c4acc8bb0a1a2f2e5d7c6337a064b32162b7da9392e2ce1f52c98398

Observation 2107d52f-67dd-4cab-b998-43323ffafb7f · outbound

This paper cites Factorized contrastive learning: Going beyond multi-view redundancy.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Factorized contrastive learning: Going beyond multi-view redundancy

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:43:12.127389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:43:11.870621Z digest=sha256:fe59566bfbaeaba7496fb002b4905d3b9575e1aa297f283829ea91333ec9478a

Observation 1b7e0018-c120-4f04-9244-e1b7f0b3deec · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models, 2023.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:43:12.113277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:43:11.874497Z digest=sha256:94b927d94219678950d776561d547abccf44e89cf5f81e723de4a72a80effdaf

Observation fabbcdef-04ae-47c0-bd7d-e0808ee01439 · outbound

This paper cites Foleygen: Visually-guided audio generation.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Foleygen: Visually-guided audio generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:43:12.099030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:43:11.878506Z digest=sha256:cc5cd302aa865f6a3ac8ea2ad6c31e49412cb5dbe6d72ce9eb9619561e3e658a

Observation e950f710-a6ee-49b5-8512-54705c2bcbe4 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Learning transferable visual models from natural language supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.882574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.882574Z digest=sha256:b68e8ce4b66173d1292074706ec76e06bed571687419abd89dd8af261ce1980d

Observation 702af5aa-9c8d-4f6e-8fa0-14ed86edf5ad · outbound

This paper cites Temporally Aligned Audio for Video with Autoregression.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Temporally Aligned Audio for Video with Autoregression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.887020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.887020Z digest=sha256:9e639bc377632ff21a39c6708ba278f7d4886dfce1134144a72543330829e8ff

Observation b542e464-cd3a-4c0c-9ab9-03e4af67c341 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models, 2023.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.891567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.891567Z digest=sha256:6e4e9e282963f2b8db5bc596177003bd2490150209e7ed1705fd6f138f776ff5

Observation 751f4d28-aa1d-4f69-8aa6-76ea3027d447 · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.895807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.895807Z digest=sha256:31dd68cb77e65ee2532d4864fd485b1f0093ea2dc7c94986c1d219e6be45c8e4

Observation 961eab08-0c24-4cff-80d1-e4afeacfda5f · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:43:11.900566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:43:11.900566Z digest=sha256:b6ad51c3a9ae2252c68c51e0681d3d665931d44cf38e449fa216d921889057ad

Observation 3d4964a5-4651-4968-abf3-8f88c76afcf0 · outbound

This paper cites Long-video audio synthesis with multi-agent collaboration, 2025.

LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters Long-video audio synthesis with multi-agent collaboration, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:43:12.066434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T17:43:11.904898Z digest=sha256:0e6434e89ef7c4aefa4f8b64473f475b49c0af14e09f2fb301be52a4c9b4bfda

Pith citing papers

No inbound Pith citation observations are available.