Pith. sign in

Paper Citation Record · LEDGER

Do Joint Audio-Video Generation Models Understand Physics?

As of 19 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2605.07061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.07061 v2

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T23:39:22.070629Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact27
  • verified fuzzy23
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a8a2620-d35a-4f03-b086-7b814fd71ea4 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Do Joint Audio-Video Generation Models Understand Physics? Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.260783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:fadd2c28976e958b62188f8f38467ec8e0bb0b3d76c93b248c4d982f230298e9

Observation c0795d41-239c-494c-919d-5e8ec55ae934 · outbound

This paper cites VideoPhy: Evaluating Physical Commonsense for Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? VideoPhy: Evaluating Physical Commonsense for Video Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.231096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:e23077e2fcef7b7835534783d971129eb61d1fb09e80c39b2146ab7c0f60e826

Observation 68ca9a9f-28a4-45f6-97c7-000e535ba148 · outbound

This paper cites VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.203083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:c7dcbcf2229f1306a6c53bfe133b950eed23db8f08c2d35e63e3f9d47bfc5d67

Observation 70f089c7-d8d9-45fc-871e-9cc4bfdee1b2 · outbound

This paper cites Video generation models as world simulators.

Do Joint Audio-Video Generation Models Understand Physics? Video generation models as world simulators

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.268019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:c7fe4656ccd931b9b0c0f6d51df45e16475fa25bdb57acab3e715df422c5fbc9

Observation 76e56aab-8d82-414a-b7e8-908829fa1ca6 · outbound

This paper cites Genie: Generative interactive environments.

Do Joint Audio-Video Generation Models Understand Physics? Genie: Generative interactive environments

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.269905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:5866e71704a004978d22c2e6dbcd87944508ea86c74f50d248c0559cc0e84e17

Observation a69173e7-7a2c-468a-9dcb-74d5437933cb · outbound

This paper cites T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.236205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:a2bd4582cc17ce9f18fa431dc97288d565264ed60771c001474e0bbe032352e0

Observation b2617e39-eb0d-4709-b816-ddbac5719cac · outbound

This paper cites Soundspaces 2.0: A simulation platform for visual-acoustic learning.Advances in Neural Information Processing Systems, 35:8896– 8911.

Do Joint Audio-Video Generation Models Understand Physics? Soundspaces 2.0: A simulation platform for visual-acoustic learning.Advances in Neural Information Processing Systems, 35:8896– 8911

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.271754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:b999d697a2c9a32beb8345dd0348d1163966e16805d64c3d3c8052d4baa6f104

Observation 162511d4-657a-4d71-8e94-51887078ef76 · outbound

This paper cites SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing.

Do Joint Audio-Video Generation Models Understand Physics? SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.223574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:cbe32a80363da1a73a194a5ff74049d6a054d189dc95b816f90bcbb42ae4dd60

Observation 0eb73f90-8817-4f51-b983-38ff60e71dbb · outbound

This paper cites Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model.

Do Joint Audio-Video Generation Models Understand Physics? Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.241509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:86c044fac5e09569ba95ceecbd0d7a4e569af1d42f48bb7f8b567908b241241b

Observation 75847fe0-e924-4df2-ac3d-91fb842a887a · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Do Joint Audio-Video Generation Models Understand Physics? Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.215209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:bcf959fa1838de5308c4f261e8f925b90c7fe980a84c37b7dd618e102e2bca4c

Observation 385760c1-0697-4437-8535-80c3b003ef5c · outbound

This paper cites Introducing Veo 3.1 and advanced capa- bilities in Flow.

Do Joint Audio-Video Generation Models Understand Physics? Introducing Veo 3.1 and advanced capa- bilities in Flow

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.266378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:b1edba952308db9011291642d9faf21444d4b183b08473aa01608844c066af31

Observation 19520228-7172-4c56-929d-2f27853d27b7 · outbound

This paper cites Technical details inherited from the Veo 3 Tech Re- port,https://storage.googleapis.com/deepmind-media/veo/Veo-3-Tech-Report.

Do Joint Audio-Video Generation Models Understand Physics? Technical details inherited from the Veo 3 Tech Re- port,https://storage.googleapis.com/deepmind-media/veo/Veo-3-Tech-Report

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.273514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:5684350df6a5c89126780d1f09c680795e0e5cf47d162f5ccd97bd63919aa2cd

Observation dddbdb33-a9e4-4c17-bce5-d2c8fef8bbde · outbound

This paper cites Look, listen, and act: Towards audio-visual embodied navigation.

Do Joint Audio-Video Generation Models Understand Physics? Look, listen, and act: Towards audio-visual embodied navigation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.259561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:737e9fcab80efd0d257a017e1e31f4fdaefe1b21fbf57f5573bffbf93658f564

Observation 1ca743af-1dfe-440b-bb86-c2ecf398f61f · outbound

This paper cites "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models.

Do Joint Audio-Video Generation Models Understand Physics? "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.238608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:f45f406bd45a6f79c161078b315b8187f225fa774be0767ee66fd4857bfb1ff4

Observation 9cbc7b9f-9d62-433f-81d8-56ab5f6b919b · outbound

This paper cites T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.258303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:728ddf975b87e19860e45dfdab1d64c44ebd508618a905115f3ac586dcfb8bce

Observation 5975024c-9ed4-454a-b844-4ca3c6bf49ab · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

Do Joint Audio-Video Generation Models Understand Physics? LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.209925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:eea55b692265b0d0602ba9f1b301d70b3adb008caecdb95230fda2bd2704a10d

Observation 03d9c9c7-4aa9-46ad-8eea-6c77da3d07b2 · outbound

This paper cites Videoscore: Building automatic metrics to simulate fine-grained human feedback for video generation.

Do Joint Audio-Video Generation Models Understand Physics? Videoscore: Building automatic metrics to simulate fine-grained human feedback for video generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.257251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:b15b115fa36de660e54e660483b568050c3cc3686c1d05e54b90a2a7828ab754

Observation d4e9e043-088a-44f1-af6c-7cb6fa3e6914 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning.

Do Joint Audio-Video Generation Models Understand Physics? Clipscore: A reference-free evaluation metric for image captioning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.261363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:72442508faded48cbc6b5f9f96283f3021833053bee64109e1319a47dbfd2451

Observation 3d3576bf-83f4-4c9a-9439-5d73cfdcf179 · outbound

This paper cites VABench: A Comprehensive Benchmark for Audio-Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? VABench: A Comprehensive Benchmark for Audio-Video Generation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.228667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:aa19b738f57beb3a639a7a90a87999e6660c9a10b62a6524ec833263beaab2d2

Observation ad0976e0-f54c-4ff2-9944-6ad393062896 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Do Joint Audio-Video Generation Models Understand Physics? Vbench: Comprehensive benchmark suite for video generative models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.264640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:4680552e27fd785f87ac181e076c15e1cdd42bf8ab10679fd91ea699265cc288

Observation f3b0d21a-3025-45eb-a22c-c9dee7fb4e05 · outbound

This paper cites A reference-free metric for evaluating music enhancement algorithms.

Do Joint Audio-Video Generation Models Understand Physics? A reference-free metric for evaluating music enhancement algorithms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.251522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:c7c0801a4c4a8227f1bd0df2815de131c6d5c88afe58407c8a889b68560af4f4

Observation 7085bb9c-47fe-4f79-8bca-cc182a94ab33 · outbound

This paper cites The measurement of observer agreement for categorical data.biometrics, pages 159–174.

Do Joint Audio-Video Generation Models Understand Physics? The measurement of observer agreement for categorical data.biometrics, pages 159–174

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.242940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:fddc94c85b35533bf293d2c6ec76fb5fc5478531d0d51dc9817d61b40e9819db

Observation ad03681c-d52a-49b6-a5ef-e09c1e0c8229 · outbound

This paper cites Video generation models: A survey of post-training and alignment.

Do Joint Audio-Video Generation Models Understand Physics? Video generation models: A survey of post-training and alignment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.240842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:4bdd10a74e3c185502a1984b502eda01c36ef277fae4e714a4ce7755c802e750

Observation 0317efd0-ebfc-4283-9024-76ef8e8627be · outbound

This paper cites JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization.

Do Joint Audio-Video Generation Models Understand Physics? JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.195881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:2c1a610b613a7f0d65d7c3be1070eceb79958013d916b249afd8be818857626c

Observation 0770c935-9b86-4a56-92d2-4906dbd1ab50 · outbound

This paper cites Ilya Loshchilov and Frank Hutter.

Do Joint Audio-Video Generation Models Understand Physics? Ilya Loshchilov and Frank Hutter

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:45:08.249848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:14d929017ec84d8dce168a25fc81a585feca46692e59df43cc26c519cedc9eed

Observation 748f8c98-a274-41a2-8867-6dff29af8db8 · outbound

This paper cites Caven: An embodied conversational agent for efficient audio-visual navigation in noisy environments.

Do Joint Audio-Video Generation Models Understand Physics? Caven: An embodied conversational agent for efficient audio-visual navigation in noisy environments

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.247169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:72aba2ca5c461140a7039ffffb535b480294e35c4863037cf579fe3887438e6f

Observation 9eb70423-d81d-4dac-a462-08c57376539d · outbound

This paper cites Tell what you hear from what you see-video to audio generation through text.Advances in Neural Information Processing Systems, 37:101337– 101366, 2024.

Do Joint Audio-Video Generation Models Understand Physics? Tell what you hear from what you see-video to audio generation through text.Advances in Neural Information Processing Systems, 37:101337– 101366, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.249135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:e0915226ff4d3f80279511c89c9625232a2cfb50a1c6c879f633ff0b1d33db20

Observation 32cb4ad9-c698-4217-9117-a4f494c3acf7 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.269827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:e0f19cdb626bbc23dcd61519b557810d91dd9da1f491f9cc10836fe95d6a1557

Observation 5bd31d78-326f-4c67-bd83-a5f8bc71f341 · outbound

This paper cites Tavgbench: Benchmarking text to audible-video generation.

Do Joint Audio-Video Generation Models Understand Physics? Tavgbench: Benchmarking text to audible-video generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.239063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:e39fbaeba58efc70ddc168b74021f95f364fc6d915ac3f3dd6644b9919ea42bf

Observation 5a7ff603-bb63-429c-94d6-992a18facb2a · outbound

This paper cites Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.217598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:eedfead058c8ab543d5f9d90e2184be35822006fd8d2663eebb50e1c9c500fca

Observation 3dfafc7a-5b28-4f36-b88c-4f1d3cc6e3f8 · outbound

This paper cites Do gener- ative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958.

Do Joint Audio-Video Generation Models Understand Physics? Do gener- ative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.237267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:f33d7982384057008f81f0b662c96878c2ba816731f8141ce64fd66503d1571f

Observation 566884d4-dc78-4da6-b2f2-fe67cda4520c · outbound

This paper cites Sora 2, 2025.

Do Joint Audio-Video Generation Models Understand Physics? Sora 2, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.233544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:057c81c421bba57e881e26fbf8726e3b8296f60e7475939bb798a949dd987821

Observation f703365e-a69d-494b-bbf3-5a12f52a8deb · outbound

This paper cites OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text.

Do Joint Audio-Video Generation Models Understand Physics? OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.233900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:8b806923754534ec840166564b3f55249ec3a25101954290453309a6d44429f0

Observation d376540a-2246-45e3-be74-25853a0dd631 · outbound

This paper cites Seedance 2.0: Advancing Video Generation for World Complexity.

Do Joint Audio-Video Generation Models Understand Physics? Seedance 2.0: Advancing Video Generation for World Complexity

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.275677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:9274179f2d6d4b1b3a2d3c74ffd89de4942797dacfc32e1c68a69a3c80b853d0

Observation 084b07ca-7a41-4ddc-9b31-a0b1e4913807 · outbound

This paper cites Savgbench: Benchmarking spatially aligned audio-video generation.

Do Joint Audio-Video Generation Models Understand Physics? Savgbench: Benchmarking spatially aligned audio-video generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.231900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:31655486816c2da3ed61a18b38b99602d27b7180d07ccbc86dd2988482a376c8

Observation 6f2e1459-40fc-48d7-b3f0-8ab806ece0f3 · outbound

This paper cites OpenAI GPT-5 System Card.

Do Joint Audio-Video Generation Models Understand Physics? OpenAI GPT-5 System Card

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.207476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:ec4221894b5dcc4e286493192dd211cc929c9c32cecf9e67c9a33eee78aea321

Observation 4c469386-d564-4fbb-88f1-64ad2f9fc04e · outbound

This paper cites From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation.

Do Joint Audio-Video Generation Models Understand Physics? From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.234269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:47f2e1ca88d4b61cdc2a793ae21d9378782bbae7296ce3a26d96822a5a95f3d8

Observation 8b2f2e9b-780d-47b1-9ca5-5bb0fc51fc2d · outbound

This paper cites T2v- compbench: A comprehensive benchmark for compositional text-to-video generation.

Do Joint Audio-Video Generation Models Understand Physics? T2v- compbench: A comprehensive benchmark for compositional text-to-video generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.235367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:3e1a7110f3833b122c1da273183ca5ebd2d213fc780f06725e70c44379077348

Observation f75f7dd6-302c-441e-990d-695c2314aa8a · outbound

This paper cites Sonicbench: Dissecting the physical perception bottleneck in large audio language models.

Do Joint Audio-Video Generation Models Understand Physics? Sonicbench: Dissecting the physical perception bottleneck in large audio language models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.208655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:66444712dd8f5e3c0680e837815973ffa5a7f760420ecde4c8454eb2d7768dec

Observation b99f848b-3f92-44c0-b1c3-d1f9628e34e9 · outbound

This paper cites Kling-Omni Technical Report.

Do Joint Audio-Video Generation Models Understand Physics? Kling-Omni Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.219778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:db00e3db810541654b616d1390b36ccecc16e612806a0aa187695284f79eb2a5

Observation 50d22eea-2a04-447b-aa7b-793f65bc697f · outbound

This paper cites Qwen3.5-Omni Technical Report.

Do Joint Audio-Video Generation Models Understand Physics? Qwen3.5-Omni Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.255291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:18a446433f2cd593b6d11820b78ce34a6c1fddae974d2f8fda92e103db43aa2d

Observation cd596f24-076e-4356-a092-3b74d53b140e · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Do Joint Audio-Video Generation Models Understand Physics? Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.226099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:74ca993b6a66c3b4babde7bc9e09387fc96df9d888b9fb94953e73c473b1c1d0

Observation 158c2ec1-1e00-4bfa-9c73-b4ea28a3f037 · outbound

This paper cites UniVerse-1: Unified Audio-Video Generation via Stitching of Experts.

Do Joint Audio-Video Generation Models Understand Physics? UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.252808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:7156a4c4e5574cc02819098a4ba822fa89fdaa60742de610fa39e9ecbd592f19

Observation 6b087763-ae3e-4f2a-bce9-60e6a324b030 · outbound

This paper cites PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.237129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:85e60c8f033c1c37da08a01dff30887c9db9e23b8c4099eb53db90f941960ae4

Observation 5ebebb9b-8765-4ce0-aac4-710fb43f94cd · outbound

This paper cites A survey on video diffusion models.ACM Computing Surveys, 57(2):1–42.

Do Joint Audio-Video Generation Models Understand Physics? A survey on video diffusion models.ACM Computing Surveys, 57(2):1–42

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.253226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:0bb73c53ec74b44c2f1e9af3e222b4dfab43358edac45899f63da23144e98bf9

Observation f7c47df5-db7c-48c4-af0f-dd15dd4d8eca · outbound

This paper cites A Systematic Post-Train Framework for Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? A Systematic Post-Train Framework for Video Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.213661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:cb35c311c62dd47a223a4d3ecb41903436badd78915e14669573f8c4286eb217

Observation 3221fd1e-30a0-4801-a237-d9b5617ed036 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Do Joint Audio-Video Generation Models Understand Physics? ReAct: Synergizing Reasoning and Acting in Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.266662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:7f1b1a4d8c575060966be3e62edc0e32ba715a439529e1c64546bbd37c09dfe2

Observation eebe27e9-d3a5-46bd-80e6-a1cc79960dd5 · outbound

This paper cites Diverse and aligned audio-to-video generation via text-to-video model adaptation.

Do Joint Audio-Video Generation Models Understand Physics? Diverse and aligned audio-to-video generation via text-to-video model adaptation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.255030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:f2ca02ec0e3ae9444182fead75ecc53a09f0e5ca9809e3e20e89bbd2dc891ec1

Observation da48796b-6c81-455b-8b34-3c8f4ba82f53 · outbound

This paper cites Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing.

Do Joint Audio-Video Generation Models Understand Physics? Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing

Reference 49

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T23:45:08.273174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:bacd66b7dbc8891a80bf4028e0f4c6620665ab63c7ace3f5e5a80dcfaff06e7d

Observation fb640097-828b-45fa-a630-04466d5e7927 · outbound

This paper cites an unresolved cited work.

Do Joint Audio-Video Generation Models Understand Physics? Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-07-07T10:33:40.263014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:69e6dafa90215a088558f127c6fe7211825a057e5dcc2837dfb236a1eb352b87

Observation b2ccf433-5f3d-40cc-a41d-fa46ed0d7a46 · outbound

This paper cites {video.event}.

Do Joint Audio-Video Generation Models Understand Physics? {video.event}

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.275271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:867f46abdc390265e424cddc32f6a5396486d5cfa19673a41d7457a4b22a96f4

Observation 8b5a116c-98b3-4886-827c-d5f100057b03 · outbound

This paper cites would normally be audible if real-world physics held; answer Yes if they are appropriately represented as such (typically silent here).

Do Joint Audio-Video Generation Models Understand Physics? would normally be audible if real-world physics held; answer Yes if they are appropriately represented as such (typically silent here)

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.230063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:988645cd0b2fda665199299ff865bdb7f2aba2995a44689ccb0eeede18bce94d

Observation ea3732ce-2cc2-4940-ac13-347e774178af · outbound

This paper cites the clip is expected to be silent during the depicted event; answer Yes if it is appropriately silent throughout with no audible leak-through.

Do Joint Audio-Video Generation Models Understand Physics? the clip is expected to be silent during the depicted event; answer Yes if it is appropriately silent throughout with no audible leak-through

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.263931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:63b949cc9f57687b6298462e2ef4c717db23cf05cac8f4c978192a097b34cdb1

Pith citing papers

No inbound Pith citation observations are available.