Pith. sign in

Paper Citation Record · LEDGER

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

As of 11 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2606.21661.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.21661 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T14:30:48.687642Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:44:39.879025Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T13:04:15.827084Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact28
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3109daf9-48c2-4938-922c-20b6ea6c267d · outbound

This paper cites Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Di- dac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Di- dac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:36.927383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:bd6d4ed3c7864afc09d969f0324ce3f676760b6833a821670c8305618c27a5e6

Observation c5dddc9b-193f-412c-a5aa-e4af4bbd8733 · outbound

This paper cites Qwen3-VL Technical Report.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.918036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:9abc0438a0e6f2dc17b6621c51524cbbe9f2980008e5c27c301bd96c4e3812f3

Observation 3df3c858-676b-4256-9542-88db439d00ff · outbound

This paper cites WhisperX: Time-Accurate Speech Transcription of Long-Form Audio.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating WhisperX: Time-Accurate Speech Transcription of Long-Form Audio

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.915075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:6d108f1976bc5d06dd64c16da1533eb8be24d69f58f3219909fe9140c3192b3c

Observation 46b26646-5b5d-4453-b3ce-17cbbf741618 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.965088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:564f5b22fd123ca482ab1fe230de379b8faa98ab1e5354d11b0f44d81e08e007

Observation 076fcbcf-dfaf-4599-8249-d034dada056a · outbound

This paper cites Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:d9e9b5772538f5ef49f2576e551e34321551bd99928c18487f8135093eabde0e

Observation 8f9091b4-7e61-4e5c-bcb0-3a88a881841f · outbound

This paper cites Seedance 2.0.https : / / seed.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Seedance 2.0.https : / / seed

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:1057d1789c2a5aa9fbc644991fe8b924407710aa2706b8b126fc5a98475c65bd

Observation cad5b838-8358-4c86-a801-193602d54785 · outbound

This paper cites Id-lora: Identity-driven audio-video personalization with in-context lora.arXiv preprint arXiv:2603.10256, 2026.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Id-lora: Identity-driven audio-video personalization with in-context lora.arXiv preprint arXiv:2603.10256, 2026

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.850711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:27402f86b55e468f780b81f31fe69b6f873d358e37f540b26869e87c7d41a7a0

Observation 601b8e3f-a1d3-41a2-b808-2432f618b09f · outbound

This paper cites Hybrid Spectrogram and Waveform Source Separation.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Hybrid Spectrogram and Waveform Source Separation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.882109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:3f8bd82fe23cb5170b6c3c1671c61bbfe9862ccb033859ff249bda1082387ded

Observation da475eba-75cc-412b-96bf-e1e8260596ad · outbound

This paper cites Clap learning audio concepts from nat- ural language supervision.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Clap learning audio concepts from nat- ural language supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:fcf3f83d3f6f987e723f0101824b079788c28572c39b9aad038c955e6331c38c

Observation b7b08f5c-a7ca-4800-867e-9a8c9d9732f3 · outbound

This paper cites Beat this! Accurate beat tracking without DBN postprocessing.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Beat this! Accurate beat tracking without DBN postprocessing

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.924623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:1dfc3a7c9932c358dbcbe531fc9c3ebe514023fbf04d54f6a5187faa43377c99

Observation 2d0830ed-7749-42f1-a842-88582beec87a · outbound

This paper cites Dreamid-omni: Unified framework for controllable human-centric audio-video generation.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Dreamid-omni: Unified framework for controllable human-centric audio-video generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.887697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:324cdf6514d0e15829bc1201129f5d6016520ffb6a0ddcc622676074acf0e8f6

Observation 073bfb7a-b621-4518-9f95-be6ef80f2a32 · outbound

This paper cites Long Context Tuning for Video Generation.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Long Context Tuning for Video Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:36.920808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:997c2c8b2c2e82c830212197406e5cfceaf5b99608fdd130e37160e227f93ba4

Observation 519a5368-c539-4a88-b084-bb11b06e56e1 · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.957109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:90dbd009f6e76785024022005539665cb0d46313dfc342aae114f48f524529e6

Observation 0edea6f5-b840-4b02-9fea-aa0c6224da8b · outbound

This paper cites Cut2Next: Generating Next Shot via In-Context Tuning.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Cut2Next: Generating Next Shot via In-Context Tuning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.921720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:7e54e36c1e1f0aef3cdeb8cb771b4b6d77b359ddf10362b2bc88374eec7972ca

Observation f235913b-f681-4233-aeb4-435459498be6 · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:c6186fd4e9065284ba8b8e7de305c08cb0c06487f9ef9ed4973f8e9b0685b23c

Observation 1d0e5dbf-e0c9-4441-8640-6aa2681feb0c · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Imagen Video: High Definition Video Generation with Diffusion Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.950324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:8adf872b684d3c485b17cd9ce5b14ffb615171db6e3cd15090de7eb5e024d193

Observation ffa644c8-9983-49df-9b0e-599a5e7b91f0 · outbound

This paper cites Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646, 2022.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646, 2022

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:063a53628755542c1c0a994eba5409b3ce6a4abe365053a7b0ad7d748b8f5d43

Observation f3b0e743-ccab-466a-bb1c-0f7c7b86e2e0 · outbound

This paper cites In-Context LoRA for Diffusion Transformers.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating In-Context LoRA for Diffusion Transformers

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.953901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:e7c7c000e5256bbdb7b3223d3a2d5653c36cf2c2c10a9b8d4e963bb68aed5ce2

Observation 9bba47a9-0141-4765-8cec-a0ddd537f8d0 · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Vbench: Comprehensive bench- mark suite for video generative models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:03c944bf83762d446c88e67a8f5e44468704d53c8dcafff76bea8012b4e84058

Observation ecc88e60-d16a-436e-83a3-0b1cde99f00a · outbound

This paper cites Shotadapter: Text-to- multi-shot video generation with diffusion models.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Shotadapter: Text-to- multi-shot video generation with diffusion models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:0c911bb7016bd8f688c34778c06a746c40f7d6cc29272c22ab077bff1ad560c6

Observation f1901d49-8938-45b9-9da0-d2963c6fc661 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.927894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:42f871ce1cd6b178befda31bcaa76fe9228ac56abb0708344f4055370c11a075

Observation 1922648c-fbbd-45c5-a4ce-ce692ac4225d · outbound

This paper cites Kling ai.https://klingai.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Kling ai.https://klingai

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:379c85c9d85b12ce5be81bb87e3462ec1d8ed0e5c63da34f33ef505b36c17b37

Observation 62b7d4ee-a8e3-459f-ac4d-161813ccd1d5 · outbound

This paper cites DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.930647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:66a990be3431ea186f7878c463b66cbeab276c8c259d4e0be45498f21041a405

Observation 3a1dcf94-3ccd-425d-84db-1764b114ba7f · outbound

This paper cites Univa: Universal video agent towards open-source next-generation video generalist.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Univa: Universal video agent towards open-source next-generation video generalist

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.953703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:fac7506ef3452d5b26aa55b059193adb4f4bb8976ab10930742c4d6ecdfff913

Observation 5619a780-7c4b-41ed-aca1-2dbe2cb69118 · outbound

This paper cites Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, 32: 2871–2883, 2024.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, 32: 2871–2883, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:3ab25d6a64cc6604d02d9854763a3d32447ddb9d1b9fd3824b6d9be3b03221f8

Observation abfa0db6-9399-46ca-bc0c-b07c21a587b6 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.885002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:59fba6248b334578a4f5b03ad715107698ea08d19735fec793ee99c7cc6decee

Observation b4cf5651-b829-44c6-a510-06f0cf93c009 · outbound

This paper cites Filmweaver: Weaving consistent multi- shot videos with cache-guided autoregressive diffusion.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Filmweaver: Weaving consistent multi- shot videos with cache-guided autoregressive diffusion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:0e2f4ef1727157aef795699956a22110dd2bebee827ef18026c41d2521a7a964

Observation 8a85c64c-1c23-49bd-afae-485c195d41b7 · outbound

This paper cites Shotstream: Streaming multi-shot video generation for interactive storytelling.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Shotstream: Streaming multi-shot video generation for interactive storytelling

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.907596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:b7acb72a5379ebeebb352e6f4d9c771a11b2cd6d303294b24c5adee5f8987a68

Observation 4a92260a-e066-45e0-b1c7-d3c296fec6a0 · outbound

This paper cites Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, et al.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, et al

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:36.873321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:4114bf52993523f4edf72dfc89963171de602a2f29ffe897670a2d53528f0a5c

Observation 799e665c-290b-41ad-8a11-fc51f623ea55 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating DINOv2: Learning Robust Visual Features without Supervision

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.946480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:9ac0f25c0a53f68638da974dacc62ec8e4a6c5ddd6633d9a71e6918a07f45861

Observation 293e1ccf-edd6-4df8-a705-65c965b628c3 · outbound

This paper cites Scalable diffusion models with transformers.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Scalable diffusion models with transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:3310f49fd6c5522b16af7efce0dd34b9ed499e2fa2f5eea071ca486a4d114314

Observation 41ce377b-b91a-45e1-938d-3934b104614c · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Movie Gen: A Cast of Media Foundation Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.847624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:44bcc0faf13a6f12c8ca90205114b62401e81b469dac0241c00d3dac259219c3

Observation e45e663f-d35c-457a-8183-1f8cc722b268 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Zero: Memory optimizations toward training trillion parameter models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:4236f9712e672848606d1a935e1a6a763dba6099cae07908d2399dadf6cf10c7

Observation f6cab276-1fc6-470d-ac0e-fdd1cd32f168 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating High-resolution image synthesis with latent diffusion models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:31ecc5ca1a863f6d501ce3c02737bfd401fdd2a1fd1f6ee5c8f643aa9b6b4224

Observation f89166b0-29ff-48f5-b47f-9d1dc3a435ae · outbound

This paper cites Denoising Diffusion Implicit Models.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Denoising Diffusion Implicit Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.950093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:819736a2922bb83e40dd3eca05e783e360628a5357acf521d1e2bf6669b7ebf6

Observation fa4d55f7-1953-44c1-968c-9f2ffa5e6d96 · outbound

This paper cites TransNet V2: An effective deep network architecture for fast shot transition detection.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating TransNet V2: An effective deep network architecture for fast shot transition detection

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:36.933814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:c752165e21c3d7e411c73c63f18f8114e0805984e51bbac678f081d0ca90c829

Observation f3c8f0b1-2e2b-4306-b694-9b3c6c554522 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063,.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:48dc9d66a788c4aaab751117df22894486c4331328a10bc17295adc34848d3db

Observation 6484f5c3-c581-4307-9432-d3d966cf96d6 · outbound

This paper cites Mova: Towards scalable and synchronized video-audio generation.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Mova: Towards scalable and synchronized video-audio generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.941749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:d1402f4d3848df65efd57e7f60c3343ed29224a22aa5578e8a0391e767a20db8

Observation a4540cb8-2811-4ee5-bcd5-e665da92aa63 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.904301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:2582f012428e91673f76aa34123939f69620ebeb11d755c6d5b06ef0572aa1e1

Observation e8207912-e7ed-41e2-b2f0-24f606bae433 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:c26bfb6262193a4e070a5719e056a0fcd81dc0f91fc67bd6e47c1adad5869ac2

Observation 3d3b7d49-1593-4422-9f5f-449cc5f02cd5 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Wan: Open and Advanced Large-Scale Video Generative Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.930406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:735c22a50127f159e778c6816ad0851a7649994d19c446d456612e8f7ffec7e8

Observation b5f50de4-1858-42f0-9a4f-4b097ba9ad88 · outbound

This paper cites Cinemaster: A 3d-aware and controllable frame- work for cinematic text-to-video generation.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Cinemaster: A 3d-aware and controllable frame- work for cinematic text-to-video generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:73f736787f24dfca56817c49be939f39b61816c30d77ef4e78889cb05bfe91b5

Observation aa45f9f4-602a-489a-9e7f-bc7fa84edd58 · outbound

This paper cites Xierui Wang, Siming Fu, Qihan Huang, Wanggui He, and Hao Jiang.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Xierui Wang, Siming Fu, Qihan Huang, Wanggui He, and Hao Jiang

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:36.961388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:1901c6fe1b5f4cc6ba616848b5bd634488adaa531116c65409a131533141beda

Observation 6317b0f0-5ac7-43dd-b38a-38d9f7337531 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:ffb07b07a73e3159baf6377a4112532c4b94a998b42b14a34c3b9f82d75bb2d6

Observation ea3d9d4c-2ddf-489c-b6d4-a4ea9aeac680 · outbound

This paper cites Qwen-image technical report,.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Qwen-image technical report,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:0e5d01388152ca131f58bc6eaf4749e45e0e4f983e0a641a1e96e0a018220b27

Observation 62639505-b66b-4ea5-b757-a88253d6e3e8 · outbound

This paper cites Cinetrans: Learning to generate videos with cinematic transitions via masked diffusion models.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Cinetrans: Learning to generate videos with cinematic transitions via masked diffusion models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.864400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:2ff90f48eebb2ab2e80ad22082d6ff89eac147e5dedc67a7d6463f22dca3fe96

Observation 4d3da473-bdd5-40e8-b45a-188af54f1830 · outbound

This paper cites Captain Cinema: Towards Short Movie Generation.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Captain Cinema: Towards Short Movie Generation

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:36.876998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:76ced4c5c064d89d64400d04eb1e6aae7b3801e3b59d4943eb600ade102e422e

Observation a0e08016-4948-4e52-b565-a38efeb9da29 · outbound

This paper cites Motioncanvas: Cinematic shot design with controllable image-to-video generation.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Motioncanvas: Cinematic shot design with controllable image-to-video generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:db7b5635c5a0deedf35464257d70b943e7fbc6cc37144f1c87555bdeb7974f33

Observation d4d5dacd-a63d-4981-99d5-63adbba7a6db · outbound

This paper cites Cogvideox: Text-to- video diffusion models with an expert transformer.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Cogvideox: Text-to- video diffusion models with an expert transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:352c6a5dfc833d9159de0548b3e45ef7a90c62dbf7e7f2424ddb966faf529bd1

Observation 1cc83777-54b0-46f2-9c4d-1e39b737dced · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.941115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:a8db375aea025c813e797236b71015ef7c9b4ee3345f97f18b96418c14756faf

Observation eeb79f86-c8f1-4c8b-b782-16eeff218219 · outbound

This paper cites Storymem: Multi-shot long video storytelling with memory.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Storymem: Multi-shot long video storytelling with memory

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.937641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:7773b058f9820d717e0644a463bd2934520d3c5dc49c037e7b7f7c2ebfbd957b

Observation 969368aa-4cff-4f34-8b28-4c7c298ac10b · outbound

This paper cites Frame context packing and drift prevention in next-frame-prediction video diffusion models.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Frame context packing and drift prevention in next-frame-prediction video diffusion models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:b32355db6ebf2f60ac17cb0fa8173c97beb8a880bb69ed9539f8dd0bdbd20cca

Observation 95a24e8b-dc25-4377-a410-4fa1a0f544f5 · outbound

This paper cites Stage: Storyboard- anchored generation for cinematic multi-shot narrative.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Stage: Storyboard- anchored generation for cinematic multi-shot narrative

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.911140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:d31bc20c3d296af3a619124cf41ded02ef862bd8fdf8bf488cbb14b1a9a23bcb

Observation 995f29a0-0641-4568-bd24-c488280c7223 · outbound

This paper cites Fo- leycrafter: Bring silent videos to life with lifelike and syn- chronized sounds.International Journal of Computer Vision, 134(1):46, 2026.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Fo- leycrafter: Bring silent videos to life with lifelike and syn- chronized sounds.International Journal of Computer Vision, 134(1):46, 2026

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:b3cf8686dd84f4c8fb20c49417854c1952be11bf90371bb202ae582db3644a62

Observation 98404115-dc36-47d4-b24d-086680a09b85 · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Open-Sora: Democratizing Efficient Video Production for All

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:29:36.904071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:6cdf4ac473b73ebe0f3df9a83b5eb647013d0b6b09ea9e5393357b537675a563

Observation 2a32fd5b-a60b-4a49-a335-c67ff390caa9 · outbound

This paper cites arXiv preprint arXiv:2601.03655 (2026) 30.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating arXiv preprint arXiv:2601.03655 (2026) 30

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:36.946921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:78b0fcb0b65dfcea9d1e576113ec815c9636399d1293f245afae203c973f6172

Observation 81b16b97-5f97-4a6b-aea4-22d1d54811c6 · outbound

This paper cites Nipemaembemawili,haraka.\.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Nipemaembemawili,haraka.\

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-26T14:30:48.687642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:cdfa582346f077f6f0add5ee0b2f53d27dc0b0a8bca10dbe68c88eef5b43db67

Observation ad478994-600f-4eae-afb0-7efe41f60cc5 · outbound

This paper cites an unresolved cited work.

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Unresolved cited work

Reference 58

Resolution
malformed identifier
arxiv_id, observed 2026-07-04T06:29:36.917896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:30:48.687642Z digest=sha256:4a1259cc7b0163d01e58cec44344ba41601cdd3c67a1c214d49480d41c36a377

Pith citing papers

Observation 8e99a2f7-594a-4de4-9d06-1534d92a080d · inbound

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing cites this paper.

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T13:04:15.832846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T13:04:13.631681Z digest=sha256:b19fd5365ef58df58b19a5ea97305338223b7db9218efd00be3c1c8f1ac69bb0

Observation f79588df-4aae-4004-b8a4-4d776be64733 · inbound

Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification cites this paper.

Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T23:44:39.879025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:44:39.879025Z digest=sha256:2074fe62d5d40ccae89f3f4afd6cefdbba0865f660bca0e1e3aa7ef3158721a5