Pith. sign in

Paper Citation Record · LEDGER

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

As of 14 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 29 inbound Pith citation observations for arXiv:2505.23656.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23656 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:10.408876Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:30:52.394202Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:59:58.033217Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9a3f170-251e-44e0-964c-d8d060f8bb88 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:04.895413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:04.895413Z digest=sha256:d5ab86cb63416c21dfce4654611c8fd958371edd0580cf53066836691f9d7a96

Observation 8f1cb6ee-c7c2-4426-968d-ce5c80fcb68d · outbound

This paper cites VideoPhy: Evaluating Physical Commonsense for Video Generation.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models VideoPhy: Evaluating Physical Commonsense for Video Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:04.998138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:04.998138Z digest=sha256:2dbbce7830c069b8c177a3ed1be4f6345f26bddb9edf2fa0a3f331433d0fabfe

Observation 101b1ba2-fbfd-4999-b886-9c911a6af147 · outbound

This paper cites VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:05.096694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:05.096694Z digest=sha256:3078413435ba83400f81134c809f4eda69b72ef4414eee2937b792876e227310

Observation c0c39350-bc64-41c8-b6c0-30c32322d3e5 · outbound

This paper cites Bardes, Q.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Bardes, Q

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:16.676847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:05.246763Z digest=sha256:374117437deb94a9564072a80cc74a482bbb9a5407485f0ecd5c567ae305740b

Observation 4fad483b-9025-4a1c-8a8b-512cf76231b8 · outbound

This paper cites Physion: Evaluating Physical Prediction from Vision in Humans and Machines.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Physion: Evaluating Physical Prediction from Vision in Humans and Machines

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:05.353999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:05.353999Z digest=sha256:4660a499cc77137be8e1e61c172c1b864c3815f180bf34300f3aaaa580e24030

Observation cfcd3251-fdca-4688-8d8c-323dae142f30 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:05.474837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:05.474837Z digest=sha256:ef42b99fb4908fbe706747c453e1baf90bccd3f8af4219e204a8dc67e3054285

Observation 95245c2e-1a82-4ec9-ba5b-feb107250317 · outbound

This paper cites Caron, H.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Caron, H

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:05.614838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:05.614838Z digest=sha256:d8b2af01b92b463b282b12538e5fc8a5c52afdf0b50559dfb0d0a517b63c5def

Observation 875f1f3a-445c-4b35-b4c3-2412d5fd3381 · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:16.555436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:05.724748Z digest=sha256:51fa35120fc080a8dda7e79099f6a1565274a2d3172d6f6489bc7b5e4aa5b0a4

Observation c13630db-b951-4a44-a43f-3fd077b7858f · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:05.775220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:05.775220Z digest=sha256:e3429b5df2e4572bf69276f1a29b4a18e4ee71badfdc8cc62b1fefbcdb797ad1

Observation a7e632a7-185c-4143-a4dd-1c1652535c71 · outbound

This paper cites LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:05.855165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:05.855165Z digest=sha256:f86955dd8e12104110ec6b512c46053052bf048a8d4cfca1ece0b81b7de318af

Observation 393f91cf-09d7-4a5f-a18d-def4c31d7d79 · outbound

This paper cites Ehtesham, S.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Ehtesham, S

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:16.434952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:05.950664Z digest=sha256:92a3d8dc0f981394924ac59c7beb2cec929b4d3fea3916765fbdaa078a2fbc92

Observation 78cca29f-c722-48ac-a2bd-b0bf79f07ee0 · outbound

This paper cites Guiding Instruction-based Image Editing via Multimodal Large Language Models.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Guiding Instruction-based Image Editing via Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:06.001988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:06.001988Z digest=sha256:891d8d86d5fa868bf32d979b88e244cd9b38f9b017dcf742ffcb6fc73d1d2077

Observation 373c5a53-c8a9-402c-8ddc-2229e6e845e2 · outbound

This paper cites Intuitive physics understanding emerges from self-supervised pretraining on natural videos.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Intuitive physics understanding emerges from self-supervised pretraining on natural videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:06.055624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:06.055624Z digest=sha256:ad0458862e1f99f2185c715d73578f79a5fd0bd15e1a348e1d7a1abfe4262237

Observation 24ac5404-754b-4718-b8f3-717d54631f7d · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:06.140712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:06.140712Z digest=sha256:33e2599a1c8b0cee194b45f95630081517db7ab09b51f95bbf649721b5b4038d

Observation dde62608-8640-402e-be29-409149f4f4db · outbound

This paper cites Gerstenberg, M.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Gerstenberg, M

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:16.316639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:06.189807Z digest=sha256:34d64c6e77f34724645ff2baf2b7c1b631ac8caa7abe3054345c3ed9f4d04740

Observation 2408653f-a634-4ba3-a0c8-3dd06c78b033 · outbound

This paper cites Girdhar, A.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Girdhar, A

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:16.136359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:06.243049Z digest=sha256:0293e0271660bef2778260247e7fc28e3c6cf61d601dec505387fd0175ded72c

Observation 170d7212-91f2-4082-a84e-731cf4e6ce46 · outbound

This paper cites Grill, F.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Grill, F

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:06.276465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:06.276465Z digest=sha256:be70aa5f984707fd5b07dfd9657fd98023fcaa7a1e5e6034b7194fce8cd557eb

Observation 2eb06cf3-9388-474a-b890-4e6acd723a54 · outbound

This paper cites Groth, F.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Groth, F

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:15.897866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:06.314974Z digest=sha256:ab8a16c79329b15aa9ac59b81bffce4d4209e09657bd1307598c8fddc545a4b0

Observation 77cf8c38-19b9-49fa-bed2-5208afe588b6 · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:06.413996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:06.413996Z digest=sha256:d4a6bd3112d5ba48ba7a007b10707f0d384c918d5851bba58d21c284bf5d27a7

Observation 87274169-6784-4a76-884f-14c14aa99192 · outbound

This paper cites Huang, L.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Huang, L

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:15.661344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:06.504810Z digest=sha256:fc9822cdb50bebad5a728e1e71f32a06307b28e88da49fadb992afcd67637583

Observation 7e5334cb-3470-4afb-9254-9194b21af83b · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:15.372401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:06.655641Z digest=sha256:18c2c05d521861e5585b97bc9d2c1589759669d88eb8aca48493ea15cb0b6440

Observation e7de5213-8975-4c1d-985e-52b34c7b83ef · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:06.844842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:06.844842Z digest=sha256:177b44120272fe64bf2a2914307fc0cb6ce71164990bad3ee5904306e462c982

Observation 9fb36c7d-3ee9-4110-b2b6-ec1534011832 · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:06.995103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:06.995103Z digest=sha256:395d8b1fc2407ed0e75df33dfad4f78e9e07db328131d27aa793442a86cc19f6

Observation 32dea26b-8f0f-4624-80ec-2c2080399a59 · outbound

This paper cites ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:07.104940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:07.104940Z digest=sha256:f1ffe0d22d33213ff464de5cd3a5fa71033184dd824907dd46b020f1b8f91de2

Observation 3695e025-8b7a-402e-aa16-b81f58b0a790 · outbound

This paper cites Phy124: Fast Physics-Driven 4D Content Generation from a Single Image.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Phy124: Fast Physics-Driven 4D Content Generation from a Single Image

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:07.182308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:07.182308Z digest=sha256:ab892e9feb790a22bfd3c659599f85ed58cf945a5c1286fd0ab58d39a169b3cd

Observation dbfaa082-3d59-4d6f-acbf-0c5b8d928944 · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:07.329987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:07.329987Z digest=sha256:5aa1b6fc4b077cb701fc988e2b1ab113d8f14ec0a4f698e481cb557d153639a9

Observation 798eaab2-2d98-4c1e-903b-f5de3dd73e69 · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:15.115720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:07.394749Z digest=sha256:582c13917286a1bcdb2dc4fdc297edf3894923a2f7ca4f4b619b564ea9ff477d

Observation 87e02a5c-6ceb-48b0-8d58-e5c7f31e3626 · outbound

This paper cites Dream machine | ai video generator, 2024.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Dream machine | ai video generator, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:14.735260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:07.511215Z digest=sha256:094c48d7a6c7797382ccff784d5f2c069f69be38272a29b818c04ef80bb34760

Observation 36e24afc-ee16-458f-a9d7-94fbe5aba3da · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:07.615748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:07.615748Z digest=sha256:35f4f6f96ef3568e12da1dd19fcbcb8079d023b0fddb81c8335ba455a215b926

Observation df44bbb4-99aa-4065-a2da-2b9ee4dd14e7 · outbound

This paper cites Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:07.684840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:07.684840Z digest=sha256:392420c8576e9326622b909b508d6aa4c3d095073096b088e327722b1d8109ac

Observation fc687a02-04af-467e-a761-6376695e02c3 · outbound

This paper cites Montanaro, L.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Montanaro, L

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:14.467851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:07.752063Z digest=sha256:51ec574acf6aa9343d16bf46b2d4fed720e31fce9f29f951f2588befc9c7bd41

Observation 745040ec-2c8a-47df-8cfd-85588e937a36 · outbound

This paper cites Do generative video models understand physical principles?.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Do generative video models understand physical principles?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:07.830603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:07.830603Z digest=sha256:0767d74fc5ef154635dba7ccee85afa4070b1a897c9530c5bee34a0d9e98027e

Observation f87476b0-1aff-4c88-acb3-979c1a25d3a2 · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:07.913270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:07.913270Z digest=sha256:e60bbe8752bee7d20c63800147c6f15255738ebc7ad09385b638193400e32b2a

Observation ea57cb19-6d77-49e8-a864-2e7ac5846ed3 · outbound

This paper cites Openai: Gpt-4 for vision (chatgpt with image input), 2023.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Openai: Gpt-4 for vision (chatgpt with image input), 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:14.187770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:07.991574Z digest=sha256:cdd55bc502fbada2348b3976e423f1b950dd27ede17e9b41d2be335cc5820048

Observation 335718d4-43e0-4312-941c-e44897c8cb69 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models DINOv2: Learning Robust Visual Features without Supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:08.120763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:08.120763Z digest=sha256:0721c8dd37dd1a7fc886a7c07c4f10789d55c60a7fefbddd115f9756d6eb9153

Observation 50b93e7f-7ed5-48a0-8753-e08e2f998aef · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:13.910222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:08.264492Z digest=sha256:95e1dc2e2ca98f7885723594a380bc2fe095e7c9e40d5bc58557ac3e3bc1eca3

Observation f890dc14-d8ae-4956-a1dc-9f8cf8d285f1 · outbound

This paper cites LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion Prediction.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion Prediction

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:10.987749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:08.309117Z digest=sha256:c88c1c83f7aab30fc631a2576a40f78ca3b686b9c4282d12eca650d50c182107

Observation 8ba3a1c7-9831-4cdf-9fe3-4ab12a2fe39e · outbound

This paper cites IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:08.364785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:08.364785Z digest=sha256:4905cb4aa5ee88f37801d265953c6a2ebabe2bb63cd397471c863d04f1bca10a

Observation 55008d81-e02b-4ab4-8dd6-508ec808c995 · outbound

This paper cites Rombach, A.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Rombach, A

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:08.418335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:08.418335Z digest=sha256:f27cecfd370f5935ac5be82c154a60c2ea4acef34cfb41c0d43699b9b71be1ec

Observation 1caace0d-7ab0-4eca-9473-013c0dc5e957 · outbound

This paper cites Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:08.523044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:08.523044Z digest=sha256:129a1be07cbe2c868644c1e5000561fd37a0fd72371de4daa87ef471f8ca56f1

Observation 72841b68-76d3-43d5-a4dc-c9c93bba3a8f · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:08.648474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:08.648474Z digest=sha256:38ad92f5afe130c9dc858ef38edbdea395c24ebe642ca704e8ff2bc560e48992

Observation 6ab82908-610e-4351-9475-86e5d249b1d7 · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:13.645172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:08.725070Z digest=sha256:d43378fe90716845aabecb6a11ba8030dd3d1ff742ca51572c871d465b2df96b

Observation f9f6701e-37f3-4ef8-b955-3383b3e11379 · outbound

This paper cites Venkatesh, H.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Venkatesh, H

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:13.369592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:08.808723Z digest=sha256:abbc34a86853b28b676d0010eef4a4376be0ef288141559950508253b34b1fc3

Observation 2a625851-e3ee-4c90-be5d-15a49e52a717 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:08.895149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:08.895149Z digest=sha256:1ef9d36ceab609e3d0f0a840ccf3a66f45577f8673752ec72d4e51c300ff20ab

Observation 539c7b23-b222-489f-bc12-f7fdf0861580 · outbound

This paper cites WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:08.981619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:08.981619Z digest=sha256:f83bafaddd1010528a7bdd8cb05e2d4a15ed8ec3a2144ba2dfd2df2aa0e48e54

Observation cb720472-c50e-4a01-8cd3-fcce6ce555fc · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:13.107433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:09.060709Z digest=sha256:c6a61c3aa11c306db268c85ee5bd94ff35a891833a0e95ce5207f7c5596186f8

Observation 9af82351-901f-4a68-a4ae-51167f64062d · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.146797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.146797Z digest=sha256:d53406d2a0705db38dd120ceed81707d1d9f8ad16b87acf0c35f0544708fbafd

Observation 714741ad-6430-4188-a51f-069af5946aeb · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.237204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.237204Z digest=sha256:be4ecc9afc5e4348caeb35dff781821c59961514b9d805cd61db3a075e7d9f78

Observation 81964d07-186b-4f23-84cd-57c64f33666c · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.297377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.297377Z digest=sha256:6bdb912e15eca86930d04db90c5988589085bf6c56c83b6cc89c3a20aed62da8

Observation 44884a6f-6482-4b62-a5bf-e3353843c6c2 · outbound

This paper cites Automated Movie Generation via Multi-Agent CoT Planning.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Automated Movie Generation via Multi-Agent CoT Planning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.400264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.400264Z digest=sha256:bfc9b25d073236d8db40ef569f290a72af3485c7cb093c61111213c8f821eb33

Observation 2247713e-87dd-44bb-bdbc-e5eb78cbaf8c · outbound

This paper cites PhysAnimator: Physics-Guided Generative Cartoon Animation.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models PhysAnimator: Physics-Guided Generative Cartoon Animation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.529463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.529463Z digest=sha256:d986926fe72f9a1d3a4f481e695d39cb2a05eace5b39d4676990c179e56ea04a

Observation b18ec368-6288-47ba-a2d7-23e28e6ef8c4 · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:12.864827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:09.639749Z digest=sha256:35d263e3cadb2527f9efcda282f9440ed481fd31656b46c8b8689947c2e525e9

Observation 659592f8-6664-4c2f-be00-80e17c541768 · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:12.581524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:09.734728Z digest=sha256:124c15f052db5bdfd7ee2792484d7ed0b1c4317e4b7647aad0511934e6ce6e42

Observation 1d67c75a-fb13-4905-a6d1-1d8418547a8c · outbound

This paper cites PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.774670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.774670Z digest=sha256:b4cb386d4f795486ed1019a3a1ccf9d045f70843abff14c3ca3f9512106e1523

Observation f36ec859-0d57-46f0-b088-5c6fd51419b0 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.842625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.842625Z digest=sha256:7b71e60222790d5afdfc73e5ebf49b7ae2fa55fb589f10cdc1be3b59ab4d5509

Observation b553b981-26be-4d8a-92ae-a673c15ab3ab · outbound

This paper cites Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.922611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.922611Z digest=sha256:518255671d96257a80d9600f7f8c2d07c8d94fb452b98c1be883ad687ea42683

Observation 3d8147e4-c982-45cd-b27a-d515dac70a01 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.984945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.984945Z digest=sha256:157085c5f47a75184603aecf8c47e4435a6bb372ec9aaaf6342cef443abc56a4

Observation 7d459e2f-15cd-4280-880e-385a3e59541c · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:10.140582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:10.140582Z digest=sha256:8ba2e84ecf8d9f08bb9c006427414e3d6d3708a5fd3ff2cefabeb698255b741c

Observation 5c2172d6-9ddc-4205-957a-15206135aa6e · outbound

This paper cites roll" and.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models roll" and

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:12.329091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:10.230972Z digest=sha256:b32f55e3e44338e4429d0e4cff305b213adaf24a43d460d7aeefcc5a0048f273

Observation 3e2d31fb-a1f2-49c6-b874-6b9a9064bf35 · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:12.121997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:10.287475Z digest=sha256:5da2dbd4e4c40cc9c1b336503011bf6c9283a1e185f42445787a5e45243b3493

Observation 0deb7079-5f46-447c-9c7a-4d41c6f058df · outbound

This paper cites an unresolved cited work.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:45:11.905645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:10.354483Z digest=sha256:a341853e6c65219d252be0f54769415717c27a257c6bd5a7dd0c93e46809e6e2

Observation 7e39f078-fe85-4db1-9bc0-b6885f72ef55 · outbound

This paper cites Based on empirical evaluations in Table 6, we adopted the first strategy: processing all frames at a reduced resolution.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Based on empirical evaluations in Table 6, we adopted the first strategy: processing all frames at a reduced resolution

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:45:11.655548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:45:10.408876Z digest=sha256:9056452dfed905ce4674cc80cb1ad3b16f04af82886e3234d0ad69f3c731d3e9

Pith citing papers

Observation b5a4a27d-6d6a-47e5-9999-9b24a1c0b084 · inbound

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling cites this paper.

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T05:17:06.569867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T05:13:28.767788Z digest=sha256:6d26612b22e847c9e851d793e43cc1cf56d3a973f088e289c14ccd1092822b92

Observation 32b2452f-2ca2-4a75-a901-0b65018381a0 · inbound

"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models cites this paper.

"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:30:52.394202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:30:52.394202Z digest=sha256:2815849136380162b06207b82db7198083c40350666432453fb5700ff2ce2f39

Observation b39a40a2-396c-4989-9eeb-1d6698f25ca4 · inbound

Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI cites this paper.

Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:01:13.996334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T09:56:36.716680Z digest=sha256:d6e2ce9bf757bd3a89ee3614a3b12c663385d442ca96ae68d72373d77d116af5

Observation 347d065b-bf02-47cf-93e4-dc32cde8e675 · inbound

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models cites this paper.

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:27.218289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T17:57:57.263574Z digest=sha256:bc88237d02e84550053a71d04617ab98a29308516142c7da855c8dde33d758dd

Observation 69637c36-5420-458d-a69b-d39cf17d89a6 · inbound

Transition Matching Distillation for Fast Video Generation cites this paper.

Transition Matching Distillation for Fast Video Generation VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T10:35:06.650603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:35:06.650603Z digest=sha256:a467d2a3c4a9bf133b65a429013440f2b74d81e3c9e0d044194067f65b42139c

Observation fea40679-2fe9-42ba-a4d7-980a8fee9702 · inbound

Olaf-World: Orienting Latent Actions for Video World Modeling cites this paper.

Olaf-World: Orienting Latent Actions for Video World Modeling VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T01:20:05.776188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:20:05.776188Z digest=sha256:3aa9076b25db76d73e4faa5b5955a07b664dfea2568789e02bb9fb96fe87d62e

Observation f5a7bc71-0409-43f1-ab5f-f5fe57d6eed1 · inbound

Under One Sun: Multi-Object Generative Perception of Materials and Illumination cites this paper.

Under One Sun: Multi-Object Generative Perception of Materials and Illumination VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T22:08:47.493022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:08:47.493022Z digest=sha256:b863518db8daee0498422c7168b737a6a04f833f23378f591c6ef1edabee58c1

Observation b8350745-ef14-44b6-b0c1-2fb281f9fa2e · inbound

Human Cognition in Machines: A Unified Perspective of World Models cites this paper.

Human Cognition in Machines: A Unified Perspective of World Models VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 221

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:12:26.004914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T08:12:15.663761Z digest=sha256:72af1f80fcef3e32ea0670481d6bd5ad8a29a22d5079c9511735746140677b72

Observation ddfce7a8-fb29-4369-8fe4-64e89c4cea8f · inbound

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models cites this paper.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:46:02.916958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:51:12.604102Z digest=sha256:0efb7f5fde88844df436aeac0626458c366ddf0d042d3786e01b2cf86ce098c3

Observation cfba5c5f-3f70-4db3-a916-26d354c6e17d · inbound

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models cites this paper.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:27.925208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:13ca019e64278dadec1bf927d1b66e03065e6764ad431e8c0c54992015b12c76

Observation d844438a-eb74-4af3-bbda-a7a2a3b2ab97 · inbound

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models cites this paper.

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:25:54.965159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T02:21:52.861714Z digest=sha256:d002148df60a73f208e73f9c4e1c37aec3374e8acbb9643506eb6f778ecdd5b1

Observation bef583c8-5140-4de0-8c28-eb5ae6021bc8 · inbound

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models cites this paper.

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:15:07.975899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T23:13:31.195562Z digest=sha256:7669494d37f13bc3ff830011000a59c26d8d8f7dde939f88cfbd4f7c7d450dcb

Observation dee2aa1f-2305-4f4a-bc05-e1b9005fd7ac · inbound

Improved Baselines with Representation Autoencoders cites this paper.

Improved Baselines with Representation Autoencoders VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:43:15.208972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T11:40:14.358108Z digest=sha256:ef3b4d2bbe36bf3c687387e2eef8f5b0c3af871fff1b2521f91ab5499cd31d9a

Observation cf386c95-a14b-4de8-8545-7dfc4da47dfc · inbound

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation cites this paper.

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:08:13.552484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T11:06:09.367559Z digest=sha256:db8ae78d50f072a2716aa5e2f4642f562db9d4ba7b96e773c7b2b2a57c6d0fc8

Observation e4378f01-936c-4241-8c9f-b3310203a143 · inbound

Spatial Gram Alignment for Ultra-High-Resolution Image Synthesis cites this paper.

Spatial Gram Alignment for Ultra-High-Resolution Image Synthesis VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:29:39.647375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T05:26:29.331382Z digest=sha256:91d9d0fb85c0905ac9ed27471ff678c22d74a55e8a2fd00ce6bf2e0ad33a2115

Observation 5a6330e7-96b9-40a6-b95f-eeea49b81e39 · inbound

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation cites this paper.

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.673333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T05:34:03.684055Z digest=sha256:cc2d85d5ffc4b76cae7de6dfeab3170b06fc3a44db8b3c85abdd417c105a49b0

Observation 06816102-79a0-4900-aef0-2137740776ec · inbound

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation cites this paper.

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:59.244185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T16:45:15.955954Z digest=sha256:a4c07459d063640322fb84bba97565e9ae6bb7ab8468361dc3e526f4fb58871d

Observation 69bacd6a-cb6c-4e45-b530-b31d7f79b389 · inbound

Tempered Self-Similarity Alignment for Physically Plausible Video Generation cites this paper.

Tempered Self-Similarity Alignment for Physically Plausible Video Generation VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:44:38.419321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T11:39:06.597513Z digest=sha256:d5d411f182cdc95a2dfa25993d408b1b9f68da6f4ab223ee92e8f293dcfb6afa

Observation 8a7e44a6-7118-4dbf-aefb-16e6097fecee · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:26:17.292794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T15:26:21.284810Z digest=sha256:5f1fea59071f3c348efb6b77f403d80cb1c9ec3959bed391bf3b1b280a01157d

Observation 41590636-0f5c-4401-8ca2-229780b66fb2 · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T10:44:36.814378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T10:38:22.619277Z digest=sha256:4b3e3f03c648040d3810f57fee13c3c8d987139ffbfa9cb958c16260aa32b7b4

Observation 82558469-6c06-4310-a168-d1df9ffc5073 · inbound

Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment cites this paper.

Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:16:44.766781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T07:02:37.291472Z digest=sha256:d95a6d68930292f58014ff75035d3ae88dd179fcc9b0aec19e6b1ad69d7e9459

Observation fa4e91e3-0bc4-46b1-9ce4-87de4c77cc3e · inbound

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model cites this paper.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.711321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:d3bd0d7ba38ede6783a3fd9b94c9ad74932c5aa4507ecc5b3617da75ab2412a9

Observation a3d93cff-ae59-400c-9522-85b5c75c3b6a · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:59:58.035148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:ba58b4b094b3d77c5ec16b98457d22a29d3bdb722d5cc1d9330f66474c52cfad

Observation 0cf101dc-2c4b-4fee-ad47-32c22d7f180f · inbound

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation cites this paper.

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:19:51.068810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T05:16:53.011837Z digest=sha256:143d65d91c490df5f675b98aa0c64481e5906d74104d251b5bfd3130545392c6

Observation 41115953-6c40-4c0f-b9ca-6988c3897b11 · inbound

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation cites this paper.

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:56.910145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T04:34:38.286863Z digest=sha256:b5f1c414e9c7762b9613c86fee94d15795159a72080c985a7e751b94900f425e

Observation 9af6b0f9-eacb-41c7-a90e-2ac286a2674c · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:95a4193ff0b2a97b73c6e70a6465da96fc1311a85d60e323d27346dce84740aa

Observation 6d6d83d4-61b0-4f0e-a988-27a6bc4da859 · inbound

Enhancing Video Physical Consistency via Role-aware Joint Training and Modality-decoupled Denoising cites this paper.

Enhancing Video Physical Consistency via Role-aware Joint Training and Modality-decoupled Denoising VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 43

Resolution
malformed identifier
no resolver link, observed 2026-07-11T15:44:22.210969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:44:22.210969Z digest=sha256:5c4fbe8f5b37c9c5b4a5a5273d7e1fe086143c1f13673c9e2c8289f42028b25a

Observation 9b091439-78e3-4a07-94a1-48dbf571c7d3 · inbound

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders cites this paper.

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T02:51:47.212234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:51:47.212234Z digest=sha256:49c267a86b107f115db30b247a755f757b4f34412d6e31786b0d5fa85c0083b4

Observation 52c22c41-0074-482b-bf14-43e8adf5a7f5 · inbound

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment cites this paper.

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T05:27:20.293209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:27:20.293209Z digest=sha256:ce613cefde28612aa1ae3cd7374a4a3dd88c66c280ab3761dbe8c7c70a3b7d5a