Pith. sign in

Paper Citation Record · LEDGER

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models

As of 4 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 1 inbound Pith citation observation for arXiv:2605.01896.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.01896 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-02T23:56:39.865433Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T05:16:53.011837Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T13:19:51.035539Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact13
  • verified fuzzy22
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch28

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c4b3368-2e7e-443e-82d8-574ec1cd9ea5 · outbound

This paper cites Proceedings of the 3rd International Workshop on Rich Media With Generative AI, ACM (2025) 1.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Proceedings of the 3rd International Workshop on Rich Media With Generative AI, ACM (2025) 1

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.207281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:12967566f12d5196297e6f93aa24d6d65c8f4f89478573d9733d86bc56619f7e

Observation 8c3c0d46-2af4-40b7-92b5-d9299b936c3b · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.889960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:1158130ff85c6c1b4124fbf874d35ff27e50bb2c611511c77f903f32a80e0bd3

Observation 81b96c12-5d7d-4832-9ac5-7adee8e2471c · outbound

This paper cites Advances in Neural Information Processing Systems37, 24081–24125 (2024) 4, 5.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Advances in Neural Information Processing Systems37, 24081–24125 (2024) 4, 5

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.186154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:c2e68e819be3db9dc39fce5910f146c98b078f37d5454bf238315f29717a39de

Observation b048413b-4703-4966-b837-44b9b060478c · outbound

This paper cites DeepVerse: 4D Autoregressive Video Generation as a World Model.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models DeepVerse: 4D Autoregressive Video Generation as a World Model

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:27.850200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:bc8ec32c877c4fd6af2d337f20fe67efd0e9ceb0cf954c4e3d5211884549653e

Observation ef9fdfa1-6f09-40a9-aacd-c0b4f6bbafe3 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.219922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:2f9b3d759aeb103f32ca8654b3d60eb18e318ca34e8a269d8d855b76dce9656f

Observation 579803c9-6f1d-4c43-91c6-21b837ecb1a1 · outbound

This paper cites 4DNeX: Feed-Forward 4D Generative Modeling Made Easy.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models 4DNeX: Feed-Forward 4D Generative Modeling Made Easy

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:27.836478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:423dffd43b3bb8174ee346f332be7ec3289929986be370defc26732ee9f075d0

Observation cfd476e8-64a1-449c-b0b6-4002a1964e6f · outbound

This paper cites LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T23:57:27.884639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:2199923d4489292bf4283a37bf0967cebecf3dc6952a402f6b7aa5b40c8a3f8d

Observation 16a0e8cc-8836-4443-bbf3-b3e6f6b1d546 · outbound

This paper cites URL: https://oasis-model.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models URL: https://oasis-model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.156482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:3748244efa13c3ea8f437113b723a35b4ae44679553cd9dbdd52fa4c2f9fecf4

Observation d1b6e760-6972-4ed9-95b0-804f3c9ecc2f · outbound

This paper cites an unresolved cited work.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-07-05T17:31:22.144644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:232d2d466da5a58f6b69a12aa1d58ad3c5135b8d248065befdb42481b2f0d6ec

Observation 29c39a25-47de-4e12-a230-810c1c2cfdc1 · outbound

This paper cites Advances in Neural Information Processing Systems37, 91560–91596 (2024) 1.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Advances in Neural Information Processing Systems37, 91560–91596 (2024) 1

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.181661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:f5f09bc6c809651f73af0c2606684f64b83c76b5daa2e4341c7222f2a1842e21

Observation 961ad75d-00ea-4d21-919c-151c442086c5 · outbound

This paper cites World Models.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models World Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.871866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:5e4b3a9c066585f1eb7372aef8fb0a77c50a851749feeff733f67ceeca558ca7

Observation a1889bb7-857e-4e50-a879-5f0a64659e33 · outbound

This paper cites Matrix-game 2.0: An open-source real-time and streaming interactive world model.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Matrix-game 2.0: An open-source real-time and streaming interactive world model

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.811276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:853dee8dd1a91740f3fcdd9c4797b21f3d820b925ff812eebcf93b917ef7d76b

Observation 52ffbd6f-cd5a-453e-987f-cd469ff81c59 · outbound

This paper cites Advances in neural information processing systems33, 6840–6851 (2020) 1.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Advances in neural information processing systems33, 6840–6851 (2020) 1

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.216666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:5ae688264d13c46b9aad9dca65e8cb33a5fb9c1e1d4fd9e15cee4876b8c1484f

Observation fabfaf09-0c74-485d-8489-d5998ec7854a · outbound

This paper cites In: International Conference on Machine Learning.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models In: International Conference on Machine Learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.193283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:e648367f766031c59b7f3f5a728ff3b5862082477964dfb3bcf900ca4dec7fe4

Observation 721ed95e-5286-4734-a5cb-2b258d6903b4 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models GAIA-1: A Generative World Model for Autonomous Driving

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.902745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:ed5f6af5a5ffdef8c2c5bd0528662b96b2885cc40ac939a5ef9b5c7d537d8049

Observation e79db412-d805-4d9c-b1a5-bea832a1d79b · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.196333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:648aafbf715600c1d055c35d3141bea6fa77dd9a7f39d17f124b7e10b3e5525c

Observation 77af5a1f-d2b3-4504-83f1-d1c83da860b1 · outbound

This paper cites ACM Transactions on Graphics (TOG)44(6), 1–15 (2025) 2, 4.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models ACM Transactions on Graphics (TOG)44(6), 1–15 (2025) 2, 4

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.175011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:0e6f0a53fe799aab412401ecba599eb921943f161cf4a24f9f68bc5432a54863

Observation d6831395-87f5-447f-ae13-daa5a933fc9f · outbound

This paper cites Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.917885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:e9d8b4d1aa2285c990d054128f0f6d6cc9d214e753939a57fc41177049481c32

Observation e8bcbbe9-8751-4278-abd6-76fd8212487c · outbound

This paper cites Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:27.892610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:8f9ec951858b0d7ceb20ca547ce61cd3426b7eba0b7ec7e60623c56e811094bf

Observation 245c330c-a734-4f6f-9b36-b0be66ed5122 · outbound

This paper cites arXiv preprint arXiv:2412.11673 (2024).

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models arXiv preprint arXiv:2412.11673 (2024)

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:27.887021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:458fd487fd6260f202176f3c9efe80985ba9e4cc4e330ff148ab9c6d9d6ee9f0

Observation 5ee74a63-e93a-4d30-9f4d-1b1187feb07a · outbound

This paper cites Advances in Neural Information Processing Systems37, 89834–89868 (2024) 4.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Advances in Neural Information Processing Systems37, 89834–89868 (2024) 4

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.200148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:db7a3b16f1e9ff46dded9419dd10f475aee0202fd2fe1ee39ec771277a8d1b2b

Observation 13a13f52-b2b9-44e5-9a54-dd0b4a34351c · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.907794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:b9b10e652f7050051c3889f41016699e9b5c17f696482f7841fd1c48639ee5cf

Observation 07cad770-5ab8-42fc-9d5e-74945155b4ce · outbound

This paper cites In: International conference on machine learning.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models In: International conference on machine learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.166903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:71bbf308b2160982d885c6961f23f31178dffb5b9a75859e89d153db606d342d

Observation 8aba3b19-c2bf-422d-9d61-f82eec3c4288 · outbound

This paper cites an unresolved cited work.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-07-05T17:31:22.170846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:76fb1b596ee03a6b2e4334fef880434aca5fd346747d0d6b0b2029ffd5a8900b

Observation abcaa948-8913-4c45-9735-94c1b44736ba · outbound

This paper cites Flow Matching for Generative Modeling.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Flow Matching for Generative Modeling

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.897575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:edc46ad4864b120456aa6fb5beaf5f8d146acebcb8201544d47b342530ada923

Observation 55beb807-b9de-4f37-b79a-cb641e158126 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.920232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:7952d3c7db971270edb43ce605066a0e36a06ca532d88612c081de05eedb040f

Observation 165d0f4e-c3d6-4621-98f9-e5c9a7de1574 · outbound

This paper cites WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:27.831413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:54d05e185c60c9db9611a53771221856d35910984e5e6ad8abf21ba3d2a26bfa

Observation f13de194-e6aa-4da3-831d-a8a0d413d012 · outbound

This paper cites arXiv preprint arXiv:2510.03104 (2025) 2, 4.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models arXiv preprint arXiv:2510.03104 (2025) 2, 4

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:27.866731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:68593fa09ed03afa532b10301a64088e5bb0f645ad3a95aa4c8fff0990b0fdec

Observation ca22bd48-8b2f-43cc-8b6b-0ace4623e5c8 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models DINOv2: Learning Robust Visual Features without Supervision

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.913034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:b73c9299bb81c47435eb7236d7e8c431c5e50ef4cb70e6721b082a9f5c0b8e15

Observation 90d9f500-73d9-4be9-8e94-aed7fb809064 · outbound

This paper cites URL: https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/ (2025) 1, 4.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models URL: https://deepmind.google/blog/genie-3-a-new-frontier-for-world-models/ (2025) 1, 4

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.189550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:557ec1e25bb981943a994163921d1ba0b99c9009db5913e8e5ebd8a1f3dc320c

Observation 41c726b4-13ca-4e61-877e-b65b3507b5aa · outbound

This paper cites an unresolved cited work.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-07-05T17:31:22.159499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:17a1e630542c0cf3a6cf0eb6e85c6233ddbbfe02da62e3c1ecad93ad0e476560

Observation 1a510f4b-6676-4fb4-81d5-1073b85ecb3e · outbound

This paper cites arXiv preprint arXiv:2510.07313 (2025).

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models arXiv preprint arXiv:2510.07313 (2025)

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:27.895216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:5f2fc5359c81d06f196a2b84b90dbf30352cf14752afb66c966f719d35d07ade

Observation b6334d7d-f5bd-4cc7-8fd8-928cbeea3341 · outbound

This paper cites WorldSimBench: Towards Video Generation Models as World Simulators.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models WorldSimBench: Towards Video Generation Models as World Simulators

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:27.856576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:ddcebaa64e2c331385fa508d0e8d8e3acecab114c1b85cf2754b7f9435a3167e

Observation 002bb7aa-47bf-4735-a5fc-07a24fc34430 · outbound

This paper cites In: International conference on machine learning.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models In: International conference on machine learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.178199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:47f45d879fae37e7647a56fafc378b085c1afab31aac2adc933603d92f9d0018

Observation 0d394dd0-917a-4984-9021-542b0e0d63c5 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models SAM 2: Segment Anything in Images and Videos

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.927604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:148148d4036a86c803f1430f60d6f0ca176e7a15305f73ea933c23d2b85074e5

Observation bb32773c-3889-4717-895d-37ce5eceb954 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.210259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:8a93b94d666e90be7c6fe44cc06be3d499eb6587cf067f2f15facf5e751f51d8

Observation 19dc11d3-2859-4f34-a692-3dee7fefa261 · outbound

This paper cites GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.821747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:e075f81b9ad820053d33ed6f907c47dc15f35f536e2fc95bb5ec22cc7cf35875

Observation 9a68f8b6-04ce-4b5d-b718-21b84b9c8d21 · outbound

This paper cites DINOv3.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models DINOv3

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.882251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:031a1eec6784f204a62ca4f8f52b4193b8cd7ceb508f86646ed64583e6e8bc29

Observation 772a0fc2-c6f2-490b-969f-216e194fd766 · outbound

This paper cites History-Guided Video Diffusion.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models History-Guided Video Diffusion

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.814132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:08de0f27f4194869a492b3827dae13c1fe12d82b5cf553444c1e9356b8bb6d6c

Observation 45ec1105-9765-4fcd-8633-7a67ff17fc1c · outbound

This paper cites U-repa: Aligning diffusion u-nets to vits.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models U-repa: Aligning diffusion u-nets to vits

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:27.840448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:39867149d41a647e0e442dacf5340b91f70ced5ca33a7f791e23c6c7db56a3e0

Observation 2947dde1-9bae-4068-8951-eb976d099be4 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.852710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:259485e08ac841f55ade234568a260c9e5135d6aacc9f2da9c7004bbf3755aaa

Observation df4857d4-9714-4b5e-ae6a-829feedd2037 · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Diffusion Models Are Real-Time Game Engines

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.900051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:ece2555f8bfc1cb1d72b6783b8454a5fd0dcc7c6c587f8cce282024189a3a58a

Observation a56a3f20-9075-4b4e-b477-85b240e4a839 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.828238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:ca18711f9466fe10e77ea41abb5af18f1edc629cd36af22f0ef4602079b04443

Observation 015b2e87-e4ab-4650-b284-59d97ed7b0a0 · outbound

This paper cites Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025a.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025a

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:27.931051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:5bd011f65cad7c00b0c4a4ad865dde2feaddcdd0d1574e08f0ffe2ffbfc9f359

Observation 8c678ac9-0059-4779-abd1-9f2d5f1c37cb · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.203697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:680bd77d5083f15c313944d0b4f7e3d44ddc80a7921fa753f526e58b1e4737b0

Observation a155b283-2218-4850-9066-817fa20cbdb4 · outbound

This paper cites IEEE transactions on image processing 13(4), 600–612 (2004) 11.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models IEEE transactions on image processing 13(4), 600–612 (2004) 11

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.223632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:8d60e14718fb083459f09d73b03911d39167f463c671c454e8d13f7e6f69d1a0

Observation 89c631ad-c3c7-4c5f-8eeb-f103b5195d80 · outbound

This paper cites Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.922659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:8989cc341778fcac65ffd153aac61226d44f9eea0baa559add82110ab6bc379a

Observation 6fb69925-a0a0-43c6-a349-d17a6294520c · outbound

This paper cites arXiv preprint arXiv:2504.12369 , year=.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models arXiv preprint arXiv:2504.12369 , year=

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:27.910598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:39d7594beccc4f7172e8d193aebc68b1dff8bbf18c3f4d3e400cc0ac1cd4b1c5

Observation d15c71d8-514b-4a5e-8b14-294303f31870 · outbound

This paper cites Pixel-perfect depth with semantics-prompted diffusion transformers.arXiv preprint arXiv:2510.07316, 2025a.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Pixel-perfect depth with semantics-prompted diffusion transformers.arXiv preprint arXiv:2510.07316, 2025a

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:27.933570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:dc14d591345d05f589e9f5048bb232c4e95195c3d977c60e2b06a545ed22da05

Observation 066d2acb-eb48-426e-88ef-1c7ec9e4f0eb · outbound

This paper cites In: International Conference on Machine Learning.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models In: International Conference on Machine Learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.141214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:18364ffea1f32a205d1e0881dbba2e693f6191489a423d757f8f505cd2eba875

Observation 65faefb7-0699-42b9-9295-a268527dd94f · outbound

This paper cites Advances in Neural Information Processing Systems37, 21875–21911 (2024) 2, 3, 4, 7, 9, 10, 11, 12, 15.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Advances in Neural Information Processing Systems37, 21875–21911 (2024) 2, 3, 4, 7, 9, 10, 11, 12, 15

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.148428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:897a36b4bee560d2cd337d95341da014d50c38f24dc4dd3ad10a435debae4847

Observation 22667131-2698-4e39-bf24-1622af43da4f · outbound

This paper cites Learning Interactive Real-World Simulators.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Learning Interactive Real-World Simulators

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.878918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:9685ee480546cb370ce4ef0df48913e91e02f8b0004da47e8874eac51fefeb99

Observation 16b3b3ec-13db-46c2-b9a2-84c7af644737 · outbound

This paper cites Video as the New Language for Real-World Decision Making.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Video as the New Language for Real-World Decision Making

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:27.915621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:2589fcd64068f8103738f7c40bf45c96463ba30a14310da069253b2267089e14

Observation 16c89200-14bc-410e-bf2a-517831081b45 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.863725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:90623afc5a8cb07d4d0e4712cc77f628fab3a4d9bfd98b38fbed50c809be6480

Observation 1e71acb3-c509-4b6e-b028-1ca3acea4c95 · outbound

This paper cites In: RSS 2025 Workshop: Mobile Manipulation: Emerging Opportunities{\&}Con- temporary Challenges (2025) 4.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models In: RSS 2025 Workshop: Mobile Manipulation: Emerging Opportunities{\&}Con- temporary Challenges (2025) 4

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.162537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:93323e975c212f8bc062ff64e0318ba98e8beae6debe67888ab66e7f1cc84e0f

Observation 1e19d1fe-6548-48b4-ad6b-c9a91d735ac4 · outbound

This paper cites In: CVPR (2025) 4.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models In: CVPR (2025) 4

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.137895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:ed984365c523350d9b1432a6f8ffa642d1a4fe8d33ca5de558e13330f8c9b3d5

Observation 604b4393-89b0-4dd0-80e3-59990c0a952f · outbound

This paper cites Visual representation alignment for multimodal large language models.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Visual representation alignment for multimodal large language models

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:27.843983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:2f59511fb65f81628588d68ea2f46888349875b760270f69005a90a2d0ed60ea

Observation f87eac1d-a842-433b-b357-4fe221651a40 · outbound

This paper cites Gamefactory: Creating new games with generative interactive videos.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Gamefactory: Creating new games with generative interactive videos

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:27.875984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:be7af27e60eb1e08d335b6b61bb5febaa9359abe990d8577eaef617d3b119c0e

Observation 99e33726-fbba-4588-819d-aa1209f538a8 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:57:27.869426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:e868d55ca2c10bb9b43033b8b535564a50be8c3548cb5df761a95eca67ec9a34

Observation 133aa9b7-5b50-49dd-8e31-7eb46622d838 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.213161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:ca8d7203946345a9dcbff05123054f5672fa60e327d3658c3655f3573e36e0c3

Observation cfba5c5f-3f70-4db3-a916-26d354c6e17d · outbound

This paper cites VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:27.925208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:1473142dede870e6d550720c46efc3a7a71a98fc1e134000c5e279059383b2ab

Observation 57c266b1-ed0e-47ba-966c-c51feb0cb975 · outbound

This paper cites TesserAct: Learning 4D Embodied World Models.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models TesserAct: Learning 4D Embodied World Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:27.847549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:9eb113f6f985af5b5f851f3d99c4d7dec0793b396ab23824f8cf53c480a71a4b

Observation a850ea57-30e7-4ac6-a81e-7fb31ae08b30 · outbound

This paper cites Stereo Magnification: Learning View Synthesis using Multiplane Images.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Stereo Magnification: Learning View Synthesis using Multiplane Images

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-02T23:57:27.825294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:99782e3a262a5bb9fec5a8bcaf3940a90339833031a8233fd5dd1b8acbfdf622

Observation 9fd9cd23-4d78-4374-95c3-6c7bcb577864 · outbound

This paper cites Omniworld: A multi-domain and multi-modal dataset for 4d world modeling.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Omniworld: A multi-domain and multi-modal dataset for 4d world modeling

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:27.905387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:104a48e5f35c9bc0229df40dfd0b01237a4a308136f925a6bef95a32410a00ed

Observation 3381d885-ce7b-4b1e-900b-682b9d9311ff · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T17:31:22.153423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:71a6294c656dcd7207fe161d2f790a5b9483767a39bb85fbfe4e903ac6057795

Observation 66d3ea46-9a82-4ace-9628-db3227b8c675 · outbound

This paper cites Is sora a world simulator? A comprehensive survey on general world models and beyond.

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models Is sora a world simulator? A comprehensive survey on general world models and beyond

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:27.817748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T23:56:39.865433Z digest=sha256:b8f6d00d5288491de0144c76894458f96c20d55983f352e17a8f17c618c1207d

Pith citing papers

Observation b06947b8-23d3-4f84-9b9a-4efe674d0959 · inbound

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation cites this paper.

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:19:51.036949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T05:16:53.011837Z digest=sha256:99b68a8ced8f1c52df9d7454aa5c0e817920441e1c1719cf2520414160dc52f1