Pith. sign in

Paper Citation Record · LEDGER

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper

As of 10 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2509.04957.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04957 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:50:17.733574Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f842bb75-a603-4117-9da5-2b066087a606 · outbound

This paper cites Attend-Fusion: Efficient Audio-Visual Fusion for Video Classification.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Attend-Fusion: Efficient Audio-Visual Fusion for Video Classification

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:50:18.135702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.519723Z digest=sha256:14c95052455458dab110d7301eb83bac4b8c770cedefe94d1c79db513870f563

Observation d92612f5-0f85-4dc3-8993-f96e18ac2a2f · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.606540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.524856Z digest=sha256:c444d13e3f95804620c886bd2f88b5083f7ad95833fd0794982ef7db3da936f8

Observation 64b80587-6110-4140-a7e1-9e808d6a1dfe · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.591054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.529177Z digest=sha256:292a2aec9d3b8b545f70de64cee6d8ebba7d582c3d4106a2165f41887575c2d6

Observation b083d87c-44e8-4c1f-b798-83715b8af64d · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.533640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.533640Z digest=sha256:015ad637482d29c57da5633c9161450debfd290d319cde2d333122d966fcac79

Observation 0a4974b8-f164-4803-9450-e0320ed4036e · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.576690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.538704Z digest=sha256:de688fd744e3b658c0ea64c6e789492de4692681a452360a89155b4bc3f3a808

Observation 8f7afa3b-c735-43ec-b80b-1d9279f0be44 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Scaling Instruction-Finetuned Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.543238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.543238Z digest=sha256:b2410ad299dd2913bd1c46332ddb1163ff5f02063a44ee2284395d17cfc16fca

Observation 920e750c-87cb-44a9-90a8-487cc2e9d6d9 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.548517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.548517Z digest=sha256:de4c50f49a71ea4b042ee0d55fea96f7ef7cc2c7b1316e5c614f001e184900ef

Observation 2477458d-f157-461f-a77c-3dec81e27925 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.562436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.553217Z digest=sha256:cee3ae02a4e2b5ff75ff278edeba27111ce62fc282e0d8e875b5efff35ef47dd

Observation 8b65355f-770c-42a0-9601-7638765eca5f · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.547989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.557627Z digest=sha256:bd3e8bac3fc5118d4fe4d89224c0ec768618a7edfdfcaafcf399a7a74aaecfdd

Observation 01ee2aa1-7ee7-4fd0-95d3-481ff5b6f6b9 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.533267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.562046Z digest=sha256:72e01b6619431b934e7449f1ece3e38ad80c8127416b79f32eb2f50e795ce4e5

Observation ccad6bb4-a840-49e3-8c12-ce157375b9ed · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.566407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.566407Z digest=sha256:c4f423260e3a50b730e165225bf221e084c679bb47e01974f007027075c141de

Observation f2c81e68-5173-487d-a076-ae32126e42b7 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.570895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.570895Z digest=sha256:9f6b4a6e2314c47c4c0f067b66da17a844e578805f874f5e903439a3dbc85814

Observation 2daa0361-24c5-4086-8f2b-9dbb3aa95cba · outbound

This paper cites Classifier-Free Diffusion Guidance.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Classifier-Free Diffusion Guidance

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.575495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.575495Z digest=sha256:d657023653fad0b635e0c3dfac34f499555908eca9a9ae3bb1df264a1d700953

Observation adad0761-b7c1-48a1-868f-c62c50dd4d3e · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.499605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.580136Z digest=sha256:432045ed07fe69cf87437669f51e4d79054bd6b2414bbba93caa8e65946333d6

Observation 99e9fa52-c8de-4799-bf42-bdf00c1033ba · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.484811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.584412Z digest=sha256:5d1d8e7b648c719cf838341502d2cd39e4ab28f844fad1550e06a99650c13106

Observation 18893818-3c1a-4f82-b771-ea52e6e531cf · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.469860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.588724Z digest=sha256:59783e378711fd9e2c64983d35c52174839f9d052a20db4080098b99b1a13123

Observation b7ccc4d3-bb00-44a2-a790-a168969d09b1 · outbound

This paper cites Kingma and Jimmy Ba.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Kingma and Jimmy Ba

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:50:18.456063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.593020Z digest=sha256:52ba0afe2a5ea7191856bb61743e3e34fbfcacbabd5e808ebb031965b5241f3d

Observation 17df4267-1f6d-436d-96a8-c81db972266b · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.441960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.597490Z digest=sha256:2cb8f709ac12cadbad8ebc472903093f4d530d8ccf3d24a39b73e500ae5b298b

Observation 1108955e-afe4-412a-af83-d5f6f9f2b367 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.426643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.601835Z digest=sha256:93be2edd349417f68c536721c680522c3296a77ec5108c5b1b3d654d22fbb785

Observation f37833bb-2ede-4a41-a0f5-d21a26a422e8 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.606300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.606300Z digest=sha256:bb922f77aefbfce7ab435dfd74474e5c1ebb0de4cb428e91bb07661d6ffd64f0

Observation 1bf7075f-6193-42a6-bd20-d612f08a4319 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.411781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.610788Z digest=sha256:7f3e4e61faac158ee13e2e5ba4ae96e10501f8b36689f5e8aaf94e25fe073dea

Observation 975cc498-caf8-4376-b9ce-d025e0bd1d04 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper VideoChat: Chat-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.614967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.614967Z digest=sha256:98cdb7b7f1cfb8a0a6de01cee67ac63fb1bb0c1ccbe0ec693dea95593503084c

Observation 66111a53-d5f5-48bc-9bb2-83037c25e61c · outbound

This paper cites Flow Matching for Generative Modeling.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Flow Matching for Generative Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.619399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.619399Z digest=sha256:ce2e79d434f7c60764adb7ec8c8e6f08ee2dea2b20e25ce59f4d146031904623

Observation 0d32f6cb-e52f-4189-9917-074c376419be · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.397292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.623593Z digest=sha256:9dd22dbb7dc76c3e5c56bd99641442bd83cdb9f48117c1714214d4ee6f8e987f

Observation d382af3b-d5a5-497c-b424-bca310888960 · outbound

This paper cites Plumbley.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Plumbley

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T05:50:17.945212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.627972Z digest=sha256:c9d864b0cfcb4e8dcd02c54fc3728830edb9503b999fddb5e6365ec13f8af686

Observation 67980384-74a7-4b10-a69f-40e82ae3e346 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.383059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.632277Z digest=sha256:ca44a6af4e8942538bedbe15e30052ea6730b9a83575a91994ffeb0027ce0c4a

Observation 2004328e-cec0-4701-ad35-823601a41df6 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.368955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.636526Z digest=sha256:5d52f9d18288475284f26e4888b1653f5c0164282b62d3b4a055bac0e15a5e5a

Observation 52cdf0cd-88c2-47ff-8f27-7cf897e9e091 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.354042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.640702Z digest=sha256:19d52c1958a6c71d9e8e518f751a675b9210e6e330360eb355acc5d75aceef94

Observation 31785116-109a-4ac6-a46d-f4f50b0ebf38 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.339848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.645138Z digest=sha256:b5ae365eb8e98927eee3bc149446a4c7ae0d96e393d5b56ec15e77654a16d415

Observation e9fa5e75-9146-4be3-b3ce-2a4d7c1b470c · outbound

This paper cites FoleyGen: Visually-Guided Audio Generation.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper FoleyGen: Visually-Guided Audio Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.649506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.649506Z digest=sha256:2ccb251eaca831473bda0c7799f07decac2bba0cd53e5c7087e12135a347f347

Observation 319ee91f-0d9e-4a29-a079-c133b1f3be5e · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.325583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.654064Z digest=sha256:18d12d41cbb256d3412801d5b36606307c32d4c323eeccea75515fecee51e6cb

Observation 4b766f44-b95b-4d07-87df-f4552eaff303 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Movie Gen: A Cast of Media Foundation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.658718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.658718Z digest=sha256:5d7bfadcafbafab0af64f35ce80ffe34e29f7c8116435e4b1b86ba7bdc1d3e52

Observation 17ed6f0b-e814-4182-b78b-2bb016ad8d92 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.311338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.664093Z digest=sha256:ff9484984ee657b404c0792018cb058b31e70c29eb2cc6639128bd2eb98aa853

Observation 24abc524-30d2-4a93-b5fd-f162697875f6 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.297367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.668943Z digest=sha256:aeea5f7e742525d17bc6f4053204f91d5e7d7ffeacfcdacda14365b23e6d6c9e

Observation b7494c12-2a84-418d-92fb-3603cd230f03 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.283544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.673895Z digest=sha256:b8ae33bf1c15b21ea744e66ac618857b164ee2f639a2b088167c0a02694e610b

Observation 18a912f6-1e23-49fe-8f26-677563723251 · outbound

This paper cites STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.678651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.678651Z digest=sha256:450c173737a86779684b8d177197e9ac35e3d0a7332ce9449d5a0c26b4c0a1ac

Observation 7da1f194-15b0-4b70-b537-2d88fe5436de · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.268582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.684187Z digest=sha256:d1dddb84014167178dd048d6863a9eb12242fb396cc1aa333afbf273449adbff

Observation 104718a5-30aa-4099-98b6-1dcbbc970e99 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.254648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.688905Z digest=sha256:429210c927db132ee59f9c77d6e5d891f8913526da935a1a819b099a7430717d

Observation 99e60a90-e50d-4f49-8c2b-09ca4e1d6b7e · outbound

This paper cites Stilwell.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Stilwell

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:50:18.240302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.693279Z digest=sha256:7c1f76c8e27c9b0781f1f60d1de2c629c1745bb8a39a1964453d6348d7531826

Observation 105322e9-47c3-4fe6-a387-86a7f29ce25e · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:50:18.226263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.697683Z digest=sha256:b706159f0b135fe79dd4e2da8a13712ab2dcf2585ce5932b930ddd9c53501a4d

Observation 0fdf94d3-d15b-4f2c-90d6-0d5ded89b42b · outbound

This paper cites Temporally Aligned Audio for Video with Autoregression.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Temporally Aligned Audio for Video with Autoregression

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:50:17.811973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.701833Z digest=sha256:9db7cf2086f5ea01b63ae05737002b7d0d9bff51d77c80d5e791d113be11e37c

Observation a01090d7-0673-45c0-9e59-9b88d94a7785 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.210766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.706257Z digest=sha256:ac621cb02c5f594d882532545e6899e1a01851331377f0fef629994f36c7c99e

Observation 91c118c5-e7c3-42f4-b2af-1d7d9026b95d · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.195399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.710751Z digest=sha256:2d8d0116f4e8cdc403bf1fbe413314f5781c8ce628f3765666a3519c5cd3898b

Observation 9719fa83-b505-4803-b709-146fe5f4a38a · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.180556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.715111Z digest=sha256:7a9a7121e38d86c6ccb5adaf7135e8308466fa5ce34b4ed6371d0a58b9d3fd01

Observation f64d7a23-799d-4b28-95d8-91d5f431c613 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.165361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.720146Z digest=sha256:25434c1b4d3c4a0df1bc06a800930beccc39e8a1f8cbcfb83305d29aa44f859b

Observation 27c7109f-22b1-4a40-9825-173789221db5 · outbound

This paper cites an unresolved cited work.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:50:18.150775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T05:50:17.724549Z digest=sha256:c96750cf2d7c552e2c6d0a36746049cb8acc87088c7b2b2718cce128badedf0c

Observation 74f1ff39-134b-41a1-b0fa-820359ac3e42 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.728841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.728841Z digest=sha256:256a5a1103ecaf17a0983bb2f130213fd7000a5dfef1d24b056ef464befc5e0a

Observation a6f882af-01b9-441a-9d80-225faee1280e · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.733574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.733574Z digest=sha256:3e43ae2efb6071fd2c1fd10c6a7edaa5e4fb342b8191d4ff3aa8cf3767abd67b

Pith citing papers

No inbound Pith citation observations are available.