Pith. sign in

Paper Citation Record · LEDGER

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

As of 15 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 3 inbound Pith citation observations for arXiv:2411.14423.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14423 v4

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:14:41.894668Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:29:49.747976Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T09:13:16.515188Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy29
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 284eec96-29a3-4564-97d0-09f49c7b2271 · outbound

This paper cites GPT-4 Technical Report.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.509394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.509394Z digest=sha256:9f2b02ffa7d9da3dfe0054773c16642fba7d7c28b0ce2dd736947015177c1995

Observation 3cdb9235-3623-4016-9342-36deb782dd2d · outbound

This paper cites Mip-nerf 360: Unbounded anti-aliased neural radiance fields.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Mip-nerf 360: Unbounded anti-aliased neural radiance fields

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.852144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.520651Z digest=sha256:07e8ada5a975289c4500bee777113ec0d8848bfb0e03a516ec03926d67198eef

Observation 6d963a4e-9d0c-4a52-89c0-76b0c9c2e4c2 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.528135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.528135Z digest=sha256:ebc076929c1f28f5682409163d086d4ba754aa0a9904eb87bd71909dfb7f97ed

Observation 90502584-a51d-4986-beea-57e1dffb9fe7 · outbound

This paper cites Gaussian- informed continuum for physical property identification and simulation.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Gaussian- informed continuum for physical property identification and simulation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.835386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.535737Z digest=sha256:6f327b881a3bfbca09518ec8d515f32386df728c9222122fe8533f9368e23ced

Observation bd754748-a657-4b35-be36-82028b332356 · outbound

This paper cites DynaSurfGS: Dynamic Surface Reconstruction with Planar-based Gaussian Splatting.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation DynaSurfGS: Dynamic Surface Reconstruction with Planar-based Gaussian Splatting

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.545761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.545761Z digest=sha256:2b35c244868520d719790bc128a58191cfbc979ffce07eaca9814dc7cac81049

Observation 3e817269-3e5f-4dbd-a972-21e6060d083e · outbound

This paper cites PGSR: Planar-based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation PGSR: Planar-based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.558606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.558606Z digest=sha256:0349e31cd763342ca4e5fde35a4a9021ce5de450a4c0a0e2ca4e6b6e5c09e1ac

Observation c81d11f1-372d-4af0-a612-2d193b56ff85 · outbound

This paper cites GigaGS: Scaling up Planar-Based 3D Gaussians for Large Scene Surface Reconstruction.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation GigaGS: Scaling up Planar-Based 3D Gaussians for Large Scene Surface Reconstruction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.565103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.565103Z digest=sha256:a4595e44858eaa53619d2df5b95a65d5fd0518c87a44a367d42e55379a103b64

Observation 46e94b50-7acc-4a0a-a21b-bb4b59fb3764 · outbound

This paper cites Dream- scene4d: Dynamic multi-object scene generation from monocular videos.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Dream- scene4d: Dynamic multi-object scene generation from monocular videos

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.819588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.574587Z digest=sha256:06ef2757eebc066ac555f4d56467aa8e5e53315d4cd6aa5e6ee611ae091e85e4

Observation bd18c3b9-b7a7-476c-b7a9-51a6bb365ea4 · outbound

This paper cites StreetSurfGS: Scalable Urban Street Surface Reconstruction with Planar-based Gaussian Splatting.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation StreetSurfGS: Scalable Urban Street Surface Reconstruction with Planar-based Gaussian Splatting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.580702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.580702Z digest=sha256:91d2a70e3e59fbd4dcaa5cbe3eb816bbbf50d54ca91087eaf4f9fb645a4e9c9c

Observation 23563879-5120-4684-82a2-7d77104beb57 · outbound

This paper cites Flownet: Learn- ing optical flow with convolutional networks.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Flownet: Learn- ing optical flow with convolutional networks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.803595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.587793Z digest=sha256:0227a025436565794372eb03fb50729acfd781437c75ff9ac17f8a266cc1f694

Observation 9eac2bec-74f8-49bf-af4c-5faa4797d0a9 · outbound

This paper cites A point set generation network for 3d object reconstruction from a single image.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation A point set generation network for 3d object reconstruction from a single image

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.788075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.594406Z digest=sha256:50120870b29771bec30ce420a6e39c099f9c12e52b1c5008ef3218185557f864

Observation 2c130928-5538-4f3d-b8bb-76a277fb30b7 · outbound

This paper cites SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.602856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.602856Z digest=sha256:47b95d9d07cf479882e5fa83c13a68acd8e879632641f88f5c1d3085be671ad0

Observation 8de34ad4-7709-4b74-a416-63ecdd5e2c6a · outbound

This paper cites https://github.com/nvidia/warp.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation https://github.com/nvidia/warp

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.771902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.610023Z digest=sha256:49e7c9da6b6ea4c62cbeabbb4885a0e2c71080726c16a0360cab27e684e6cf2a

Observation 9e57340c-ce43-498b-ba82-7ec48ded3c62 · outbound

This paper cites A moving least squares material point method with displacement disconti- nuity and two-way rigid body coupling.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation A moving least squares material point method with displacement disconti- nuity and two-way rigid body coupling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.756265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.617080Z digest=sha256:f91bab36246b6e9090b99c56335111b684a9f62ebb87a8b91f4b3e2169414f09

Observation 6630e152-833f-49f6-803f-9b5a3cc7bd83 · outbound

This paper cites NeRF-Det++: Incorporating Semantic Cues and Perspective-aware Depth Supervision for Indoor Multi-View 3D Detection.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation NeRF-Det++: Incorporating Semantic Cues and Perspective-aware Depth Supervision for Indoor Multi-View 3D Detection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.623952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.623952Z digest=sha256:dd6fac2c334625cd500ef184cea458c3590c1ba72e09e1ccd636cf2b16c947f2

Observation 06456792-6da6-4599-87a5-b2a0fc5cad6d · outbound

This paper cites DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion Priors.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion Priors

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.629698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.629698Z digest=sha256:cc5e5b09b8384786e9f1f653fa9508fd801564005b22b00783e14559dcbe65e9

Observation 50e35f0d-5d41-424f-9deb-d3ce5da8d8a7 · outbound

This paper cites The affine particle-in-cell method.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation The affine particle-in-cell method

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.742933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.635421Z digest=sha256:cc65316c69aa03c5cb462d23d2f1e40065a0786aede586087fed98773c594885

Observation 1575cfb1-afda-4ed2-88ba-753f4768094e · outbound

This paper cites Grid4D: 4D decomposed hash encoding for high-fidelity dynamic scene rendering.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Grid4D: 4D decomposed hash encoding for high-fidelity dynamic scene rendering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.727252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.639971Z digest=sha256:3e2b1643017ac8f1da264f5657170913923418118c07bf0abd2598a1d4a0e382

Observation 5998c403-2e73-4922-91df-2c0469bfa33d · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation 3d gaussian splatting for real-time radiance field rendering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.705106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.648753Z digest=sha256:0fd63de2fece3aa3dabf334f814a146a98fa14dfd3b996f372a61ea0558571fb

Observation 8e11271b-1c63-4c6d-a053-2d034516022c · outbound

This paper cites Kling ai.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Kling ai

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.679987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.655158Z digest=sha256:c7de17a8aa1aab6c53fb702298a6d96e11d308264c88143e689dfe53ce79c888

Observation 8ae17a44-7ba1-4f05-ab12-cb5d67befba8 · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.667536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.667536Z digest=sha256:34270ce68cb34381aa7b874edbb5c96be179a1e942cb52360f06de52ff127263

Observation 21789489-6f0d-4114-909a-f587fc61d98c · outbound

This paper cites Pac-nerf: Physics augmented continuum neural ra- diance fields for geometry-agnostic system identification.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Pac-nerf: Physics augmented continuum neural ra- diance fields for geometry-agnostic system identification

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.664040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.672091Z digest=sha256:4fd0a57a5174e4f89779e62a8a9f032db37643d376ae87384b5df8f228fcbe41

Observation b3480af6-57dd-4701-83f0-0d7600735429 · outbound

This paper cites Physics3D: Learning Physical Properties of 3D Gaussians via Video Diffusion.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Physics3D: Learning Physical Properties of 3D Gaussians via Video Diffusion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.679356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.679356Z digest=sha256:221f4313f52a1f3b24cc3725873ed621407420e52b6755baaf3bb14022c127fe

Observation e4844b72-11f5-4168-96a7-432197965373 · outbound

This paper cites Physgen: Rigid-body physics-grounded image- to-video generation.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Physgen: Rigid-body physics-grounded image- to-video generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.646190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.689222Z digest=sha256:acd6851350dcead58b2ded38735bd2639358be7d93b97b3ef0ff92d9f46ec895

Observation b8f580ad-5fb1-4e08-9e12-57cfce72bc3b · outbound

This paper cites Coxgraph: multi-robot col- laborative, globally consistent, online dense reconstruction system.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Coxgraph: multi-robot col- laborative, globally consistent, online dense reconstruction system

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.629455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.694578Z digest=sha256:df49dbe024db22d99d64fcc8572450b82036d8ee5186cd0c27cf84d11456a4f6

Observation ba64e558-9158-446b-9eb1-2ad850079e0c · outbound

This paper cites Raydf: Neural ray-surface distance fields with multi-view consistency.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Raydf: Neural ray-surface distance fields with multi-view consistency

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.606193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.704692Z digest=sha256:0bd540a1bddb2de8ec825dfeb0186dc610962d28eae6ba4de68537a836e138e6

Observation 2329190a-bd3b-404c-b663-cdef75e97da7 · outbound

This paper cites Ddf-ism: Internal structure modeling of human head using probabilistic directed distance field.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Ddf-ism: Internal structure modeling of human head using probabilistic directed distance field

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.589738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.713674Z digest=sha256:0535cea39930a1371bd80ece945722ac89ae16d0f221b116808fbd47cc6f3347

Observation 67b62869-4155-4707-a1ed-9208b9102b99 · outbound

This paper cites Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.731372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.731372Z digest=sha256:de8cdadc37f8033ea26106e8fd213129e85d3e7fc5d0179a4fcf4f3c59b39c68

Observation aafb8e1b-309b-4265-b5f9-7c35dcf3eb63 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view syn- thesis.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Nerf: Representing scenes as neural radiance fields for view syn- thesis

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.572160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.737469Z digest=sha256:66bd06270ffaf798ec341d18099ff080ec8cd284a528d16e21a95da091da9e85

Observation 43b806d0-7003-420b-b07b-f367c7b20c4a · outbound

This paper cites iDF-SLAM: End-to-End RGB-D SLAM with Neural Implicit Mapping and Deep Feature Tracking.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation iDF-SLAM: End-to-End RGB-D SLAM with Neural Implicit Mapping and Deep Feature Tracking

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:14:42.065070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.741684Z digest=sha256:705719c0629f1ef8e8dbd892ff93ad9ac84ff38a305f8b9d97faaad34059da0c

Observation ae2ff6de-390e-440a-93b2-2578e8645bf3 · outbound

This paper cites Instant neural graphics primitives with a multires- olution hash encoding.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Instant neural graphics primitives with a multires- olution hash encoding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.555703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.751842Z digest=sha256:7a013eab97ec2510f26e5f9f3e3ba15690704746bf84fb30f3501b3b50e756e8

Observation 888c1342-7f0e-463d-90df-0545db073d8f · outbound

This paper cites Mofa-video: Controllable image animation via generative motion field adaptions in frozen image-to-video diffusion model.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Mofa-video: Controllable image animation via generative motion field adaptions in frozen image-to-video diffusion model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.533858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.759680Z digest=sha256:9784b41a830646b25e15d04720f5e281114097d398fef6d91099e71686870aad

Observation bca84269-db33-4612-89fb-dbdc3e2a233e · outbound

This paper cites Dreamfusion: Text-to-3d using 2d diffusion.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Dreamfusion: Text-to-3d using 2d diffusion

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.766438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.766438Z digest=sha256:41ac74fe58d299a6cbd377635fe4aa4b8350239510b43871348ad2e93e0d83f3

Observation 05dd3248-ed3f-42b8-8512-bdbbec8a2baa · outbound

This paper cites Robocook: Long-horizon elasto-plastic object ma- nipulation with diverse tools.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Robocook: Long-horizon elasto-plastic object ma- nipulation with diverse tools

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.505452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.775173Z digest=sha256:4659137f04ed5c4e160dca16749c18aebc79784cdb7a39703b3ca8e13a65a752

Observation 99dfa758-d315-4e03-a9eb-7a6bcf5cf105 · outbound

This paper cites Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.786810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.786810Z digest=sha256:6aa273b4c527afe36c83c219d03c53b71e3a0038d8cbe146ab983b231ff6d737

Observation 2ef3fca0-8363-46ee-a6f3-b135a7a155ee · outbound

This paper cites Nerfstudio: A modu- lar framework for neural radiance field development.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Nerfstudio: A modu- lar framework for neural radiance field development

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.489879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.794903Z digest=sha256:4aa53b23ebbee68f07fb021cbd08636bfbfabf8c69b5f7531b855a5a62bfa71a

Observation 2ac6adb8-2821-491c-b855-11a9bc818178 · outbound

This paper cites HiSplat: Hierarchical 3D Gaussian Splatting for Generalizable Sparse-View Reconstruction.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation HiSplat: Hierarchical 3D Gaussian Splatting for Generalizable Sparse-View Reconstruction

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.809870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.809870Z digest=sha256:0ab60117b7a589b9efb231b7959299d27ed6eeb8721485afe3a9b9fa404df4ee

Observation dd16d9d7-877b-45b7-8994-5ada93216cbf · outbound

This paper cites ND-SDF: Learning Normal Deflection Fields for High-Fidelity Indoor Reconstruction.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation ND-SDF: Learning Normal Deflection Fields for High-Fidelity Indoor Reconstruction

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:14:42.001425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.818288Z digest=sha256:1223310d61851b9cc74a3970046226faeb6874fb53035c7e0fa83c7cbec6bd66

Observation 5ccad3dc-23dd-4b62-8432-485b37a2be6c · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Raft: Recurrent all-pairs field transforms for optical flow

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.470513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.823088Z digest=sha256:2e006fd0a8a7bf5674e692b59cfadb1c926ed98280476038984f839190a76aae

Observation 051a1cb6-89c0-49ac-9d73-728a4c921bac · outbound

This paper cites NeuRodin: A Two-stage Framework for High-Fidelity Neural Surface Reconstruction.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation NeuRodin: A Two-stage Framework for High-Fidelity Neural Surface Reconstruction

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.827259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.827259Z digest=sha256:355b50ab1e7dd14c1a6343ad9e3c70046574a6d129b4bc041b7629088ff3c800

Observation ea585545-cc44-4dda-b22f-591e470f4f76 · outbound

This paper cites Motionctrl: A unified and flexible motion controller for video generation.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Motionctrl: A unified and flexible motion controller for video generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.453559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.833260Z digest=sha256:da66f2c93e4b388b499bf939a61732c0d101e08ace68c8fc8c9ee395fbb8425c

Observation 1f6fa08a-ebda-4083-8ee9-3f700389b021 · outbound

This paper cites 4d gaussian splatting for real-time dynamic scene rendering.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation 4d gaussian splatting for real-time dynamic scene rendering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.437130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.840815Z digest=sha256:c858fe5466300e756236f4b59783cffbb4aaaad569c7abff81adfd5fe4df3e1d

Observation 4685184c-648e-4e90-b5d1-94558c8ad8e5 · outbound

This paper cites Physgaussian: Physics- integrated 3d gaussians for generative dynamics.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Physgaussian: Physics- integrated 3d gaussians for generative dynamics

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.847243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.847243Z digest=sha256:bfe2215ef8f7af4d95c8e812b942984350fb8b3d183c5d8f78b1ac170c971189

Observation 94526604-e33c-4d11-9c30-30c9f0979ccd · outbound

This paper cites Learning 3d dynamic scene representations for robot manip- ulation.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Learning 3d dynamic scene representations for robot manip- ulation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.407214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.853538Z digest=sha256:64596645e9a172ed015fbe4ce817ec9c3933c98a2212f82514236b1833797fc8

Observation ebe8f732-6cd7-47b9-ada7-1b78a5e2df74 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:14:41.865350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:14:41.865350Z digest=sha256:715769802e9b78cf3a8138061c64deee3dd1d0f5deade78c1fb0816832d327fd

Observation c99a63f0-d463-4125-9417-9f90f6b06894 · outbound

This paper cites IntrinsicNeRF: Learning Intrinsic Neural Radiance Fields for Editable Novel View Synthesis.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation IntrinsicNeRF: Learning Intrinsic Neural Radiance Fields for Editable Novel View Synthesis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.386817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.871634Z digest=sha256:3bb56609a56034cccb029501ae5bc4b9781f029da1a52482ae541b4c1f36bb79

Observation 72c3eb2b-45f8-4082-960d-7d9edee06963 · outbound

This paper cites Diffpano: Scalable and con- sistent text to panorama generation with spherical epipolar- aware diffusion.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Diffpano: Scalable and con- sistent text to panorama generation with spherical epipolar- aware diffusion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.365964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.877361Z digest=sha256:d1e7782a8b14f8c5b1218e774dd8985c8903c81ddeff43f19d39b877b1b8da2e

Observation 1097c071-ff02-4a26-ba56-8ea1c96281ee · outbound

This paper cites Feng, Changxi Zheng, Noah Snavely, Jiajun Wu, and William T.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation Feng, Changxi Zheng, Noah Snavely, Jiajun Wu, and William T

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.349192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.888893Z digest=sha256:a5a2dca0c56587051e3e33a009cce1e5ffceea516c37620c79aafc67e988ae74

Observation bae2e573-ffc5-47fc-8ece-c375d7daeefd · outbound

This paper cites The ice cream is slowly melting.

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation The ice cream is slowly melting

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:14:42.331186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:14:41.894668Z digest=sha256:6f2c0348ee8e21d19bb402d48e07edb946d7072ddacebe1919cf8aefe447ef96

Pith citing papers

Observation b69498fd-b26a-48e4-89b7-ef9c5134c419 · inbound

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding cites this paper.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.747976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.747976Z digest=sha256:b3d321f964564e0ea0ddfbd425e2a5459f895419df9bd18fec0ee548529a69df

Observation e3b704f0-aab2-4e93-86ab-32d1e3515bac · inbound

Zero-Shot 3D Visual Grounding from Vision-Language Models cites this paper.

Zero-Shot 3D Visual Grounding from Vision-Language Models PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:13.660944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:13.660944Z digest=sha256:0c6b2515d58994cd781a9fac0a8b63e6c5937a9b0f914272c25a3d9a3718b21f

Observation cc9fead2-e70e-46c7-a758-bfb90a40f465 · inbound

Physically Viable World Models: A Case for Query-Conditioned Embodied AI cites this paper.

Physically Viable World Models: A Case for Query-Conditioned Embodied AI PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:13:16.517222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T06:55:57.801162Z digest=sha256:33a2bba71a3c2725ca2fe8a65db3bf6f28b21b3c77ae58694df0ee4bcc6d6081