Pith. sign in

Paper Citation Record · LEDGER

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP

As of 21 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 0 inbound Pith citation observations for arXiv:2603.16100.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.16100 v2

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T23:59:10.664837Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

76 of 76 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved76
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b191f6c-b551-4616-9223-e989d8797b6c · outbound

This paper cites VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:12a7556a2439a553fab1f8870e0998cfcbcf7fb729ac07dec017fbe18753137e

Observation 029fcd65-4a84-4d40-9d22-c98fc5063885 · outbound

This paper cites Must3r: Multi-view network for stereo 3d reconstruction.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Must3r: Multi-view network for stereo 3d reconstruction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:14b5ed000ed5e69ff6965e1d2a3d485f685ad9cce34a87cbc1d1ccf038971f30

Observation 0d0af5a6-35f5-4d0e-8262-7177a648fe2a · outbound

This paper cites pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:b9b488bd6c5fa7f7b95b88aa9f11dd4e2dbbe3e926afaa6590a40db28061b3a4

Observation 80a75a05-95b1-43cd-a147-5cd79a5712b8 · outbound

This paper cites Aligning visual foundation encoders to tokenizers for diffusion models.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Aligning visual foundation encoders to tokenizers for diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:27c82f68089e0ce4b43ab2345a2df0f2675a392ddc700d98536a2447cec883a2

Observation 814399c9-b3ab-49e7-975e-696a3475917e · outbound

This paper cites Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:bba92ddd24622c821381420c26c7beca45de2c1123bacabd98f978b32d6e4fd3

Observation 0ae1eaf4-2b11-4fa0-8cbd-1fbaaea832b4 · outbound

This paper cites Mvsplat360: Feed-forward 360 scene synthesis from sparse views.Adv.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Mvsplat360: Feed-forward 360 scene synthesis from sparse views.Adv

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:fb61033d72a20e783060cf364761ec7fbd4f1b9e566b25aa0affa7c6c864206c

Observation e291774c-2c90-4813-beb4-9e100a05a10b · outbound

This paper cites LucidDreamer: Domain-free Generation of 3D Gaussian Splatting Scenes.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP LucidDreamer: Domain-free Generation of 3D Gaussian Splatting Scenes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:4e5a7718969ee582f64d009231971268f6a2be754030f3bdd35f24b5fd48460f

Observation e51f91fa-c2a0-4b31-b0ac-30a7af277477 · outbound

This paper cites Fantasyworld: Geometry- consistent world modeling via unified video and 3d prediction.arXiv preprint arXiv:2509.21657, 2025.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Fantasyworld: Geometry- consistent world modeling via unified video and 3d prediction.arXiv preprint arXiv:2509.21657, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:764c87e02d90b6b03e315b7686c131c92e64867e078328039150b08245f1e76f

Observation b0556425-314d-4603-b771-be678fa1efa4 · outbound

This paper cites Understanding world or predicting future? a comprehensive survey of world models.ACM Computing Surveys, 58(3):1–38, 2025.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Understanding world or predicting future? a comprehensive survey of world models.ACM Computing Surveys, 58(3):1–38, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:d759b253c864594d1ff5d0c57fce08ffbd0598715a5f2019f57b8395d5bc42b9

Observation 560faeba-434c-46bf-833b-269fe1ee157b · outbound

This paper cites Worldscore: A unified evaluation benchmark for world generation.arXiv preprint arXiv:2504.00983, 2025.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Worldscore: A unified evaluation benchmark for world generation.arXiv preprint arXiv:2504.00983, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:031ae23faf53a2ff58f1f8f42b2655dd08dd9cae225b0227ae187562ea2b25d0

Observation 3d53ddcd-6942-4b62-a14e-643cadecb456 · outbound

This paper cites Scaling rectified flow transform- ers for high-resolution image synthesis.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Scaling rectified flow transform- ers for high-resolution image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:05b4a7932fe9c51432ad73ca55a15ef334d7eec3fd5183b4b537cd8be102a970

Observation f5145c32-1329-4b85-b503-ac26695ac591 · outbound

This paper cites Scenescape: Text-driven consistent scene generation.Adv.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Scenescape: Text-driven consistent scene generation.Adv

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:2698275b9a61631f6de8fe0aef49d4142571c518ceec283639a0ec6b2a6f1c76

Observation 6f55b328-21ac-4ec4-8d5f-45dee050b867 · outbound

This paper cites CAT3D: Create Anything in 3D with Multi-View Diffusion Models.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP CAT3D: Create Anything in 3D with Multi-View Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:b1c555a9987afae4ac80ce4e1387c7ee25286d1dfff2b581a8ebb56086af2ae1

Observation ddbfb5d9-f222-44db-b50a-674e00cd5ce9 · outbound

This paper cites Vist3a: Text-to-3d by stitching a multi-view reconstruction network to a video generator.arXiv preprint arXiv:2510.13454, 2025.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Vist3a: Text-to-3d by stitching a multi-view reconstruction network to a video generator.arXiv preprint arXiv:2510.13454, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:446675c2d3ab238fc4b6d626e8f99dea285eadef2991b24d34d7f4bd2dc201ef

Observation 7355950e-28eb-408b-948c-e31c809e4972 · outbound

This paper cites Splatflow: Multi-view rectified flow model for 3d gaussian splatting synthesis.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Splatflow: Multi-view rectified flow model for 3d gaussian splatting synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:36f301eaaf32a757cd3101a7de0b2d2e1e8a378c854b448d6a6860c988375a91

Observation 71f69703-2e26-437f-a660-2156b4d42b74 · outbound

This paper cites VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:68495142505cda94770337658de07fafe84da8d0d4d863b75aaf313eeabaacaa

Observation 705d5836-de75-4a73-a082-ccc5d2e8c2e7 · outbound

This paper cites Seed1.5-VL Technical Report.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Seed1.5-VL Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:0a48b6164387558630a15511fcc9fd93338c6211ecfe6bd9f67d982c35da6a9d

Observation 8eaf07e2-437c-44bc-b307-3852179d821a · outbound

This paper cites GaussVideoDreamer: 3D Scene Generation with Video Diffusion and Inconsistency-Aware Gaussian Splatting.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP GaussVideoDreamer: 3D Scene Generation with Video Diffusion and Inconsistency-Aware Gaussian Splatting

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:6c248f4d55f8807bfb2ee0fe180a26283b81e85b4d4f614ceec24cc176b1b71f

Observation e9912137-1bb0-4f12-8215-2b49e58d0f3e · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:c90e13391072439350ffaff59b479271934cd0cd25882292e14d5652a34f6a05

Observation f30304b8-828d-4b70-ba03-61e8eeda9523 · outbound

This paper cites Gen3r: 3d scene generation meets feed-forward reconstruction.arXiv preprint arXiv:2601.04090, 2026.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Gen3r: 3d scene generation meets feed-forward reconstruction.arXiv preprint arXiv:2601.04090, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:09802b583c1cb0b0059a6ffd22dc4e3635562c407486cca50c743483457dd6a7

Observation 1bc43f44-ca18-4f7c-b32a-cf0f18da9a17 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Vbench: Comprehensive benchmark suite for video generative models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:0f79612661e0ac121c0a39bb2ddc3a1c2ab0efd286171e45d19ef735b46b3a46

Observation 352d9f85-c5e1-4f6d-90b5-c7628a5d1c16 · outbound

This paper cites Vbench++: Comprehensive and versatile benchmark suite for video generative models.IEEE Trans.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Vbench++: Comprehensive and versatile benchmark suite for video generative models.IEEE Trans

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:09507125fc1fcf525c040580f2e9305dc1fce85d034e415ec7c0eaf348943b0c

Observation dc999ada-a0ec-4ef1-bcb9-3e2afb9a144f · outbound

This paper cites Anysplat: Feed-forward 3d gaussian splatting from unconstrained views.ACM Transactions on Graphics (TOG), 44(6):1–16, 2025.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Anysplat: Feed-forward 3d gaussian splatting from unconstrained views.ACM Transactions on Graphics (TOG), 44(6):1–16, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:6d87d4c893dc79d9e78f6ad247bbe9dd09bdbbda62a15af9aa354f4bf0fa6bf2

Observation 991ec32d-c686-473a-96c2-50b3f7625870 · outbound

This paper cites Lvsm: A large view synthesis model with minimal 3d inductive bias.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Lvsm: A large view synthesis model with minimal 3d inductive bias

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:06cbcebfb4adc52a8cfe9448bcaf088564d612792100c1b9dc67f43569a88e5c

Observation de1c2ced-e61d-4a5f-ac97-08f8cdcd4d29 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.Advances in neural information processing systems, 35: 26565–26577, 2022.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Elucidating the design space of diffusion-based generative models.Advances in neural information processing systems, 35: 26565–26577, 2022

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:350ca7ea64d7eb2d19e9ac05e6c61c42e40b7c616e553abb85629c7b755d2590

Observation 02970360-4ced-4ff0-907e-1d63bd49cefb · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.ACM Trans.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP 3d gaussian splatting for real-time radiance field rendering.ACM Trans

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:2830371c40aee72fc318de47f6c4dd9070bbc7b1a8c35729082d3c4a3297e339

Observation 543fde11-da9e-41d5-a7b4-689f283f7f4d · outbound

This paper cites Auto-Encoding Variational Bayes.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Auto-Encoding Variational Bayes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:87a9efe7562f6f5233edd28c1052381e3670443cdff8f20c6b993f5f0b438851

Observation 403da686-007f-4955-816d-033c1afbb8b7 · outbound

This paper cites 3d and 4d world modeling: A survey.arXiv preprint arXiv:2509.07996, 2025.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP 3d and 4d world modeling: A survey.arXiv preprint arXiv:2509.07996, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:45b2c996caed7cefc73c6850c5a432539634ac0825a4e2b0292da005309bdd2f

Observation c2374c19-65bd-49ce-800b-ae2e3b25088a · outbound

This paper cites Grounding image matching in 3d with mast3r.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Grounding image matching in 3d with mast3r

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:578f6a3a82bdefebc20f6be95c6a2f5db5e059507a248c05db2e3e87a1e0f45b

Observation 604af6f0-5621-4631-88be-95e162737a4e · outbound

This paper cites Back to Basics: Let Denoising Generative Models Denoise.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Back to Basics: Let Denoising Generative Models Denoise

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:e926a06e789a3d9184c0d7e1f5dfa2acf4157e4a32b1e41cc2e3d7e2c6dafed9

Observation cef447f8-6916-4e8c-8ae6-83403d54ab8e · outbound

This paper cites Director3d: Real-world camera trajectory and 3d scene generation from text.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Director3d: Real-world camera trajectory and 3d scene generation from text

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:5bee00853281ddecc3ef58d0ad80d9429a8a340bac6bcf330ca4048b602efab9

Observation 95bc99e7-20d3-4b02-a727-df2a501705b1 · outbound

This paper cites Flashworld: High-quality 3d scene generation within seconds.arXiv preprint arXiv:2510.13678, 2025.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Flashworld: High-quality 3d scene generation within seconds.arXiv preprint arXiv:2510.13678, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:b7ded2bf9911f4c89c16ee13e1fb3c60abb774ad2a326738d01e6083371e2d39

Observation 30bfbe6f-5c10-42fc-b12c-2ac83b0861ba · outbound

This paper cites Magic3d: High-resolution text-to-3d content creation.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Magic3d: High-resolution text-to-3d content creation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:e6bccd4fcc6afe182e76dd4075d35da0c7cbe1e85c305867df6b687dc52d0218

Observation c37928b5-bcb3-4d24-b4d7-0a1adbe93de1 · outbound

This paper cites Depth Anything 3: Recovering the Visual Space from Any Views.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Depth Anything 3: Recovering the Visual Space from Any Views

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:c06cd4485b1e4cf27f6d67fbee0a30641b5ce1356b0278921abe12427a3cf4d6

Observation 3ba1740b-ac28-41d2-8153-63f9eeafa229 · outbound

This paper cites Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:4004f787ffd2a9dac925d9eb43f84b5b1d8150430ee279c4c5725ca1c89a7484

Observation 4c8cfda2-339a-4d03-be46-940c150ad8d7 · outbound

This paper cites ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:679a2426b3d88b511e857536763b2cf30035c9f3dbdc82146108261043d95f3e

Observation c8715932-6374-4356-9d36-cc8d31bd2254 · outbound

This paper cites Zero-1-to-3: Zero-shot one image to 3d object.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Zero-1-to-3: Zero-shot one image to 3d object

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:4050f1a965e53161746f50886decc4cf6f67ad1c33335f1332284704d37c6008

Observation a4bef872-e3bd-4c8f-9f17-63a523de2573 · outbound

This paper cites SyncDreamer: Generating Multiview-consistent Images from a Single-view Image.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:8d0119db7eb67664c32ab5285a8a41d627160d5463ee2472e75bcba3231efb61

Observation c4504ef3-9fa6-46d7-995f-2a263de570b3 · outbound

This paper cites Decoupled Weight Decay Regularization.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Decoupled Weight Decay Regularization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:89d08596942f1a10a607c3d6312bb30ccf1066509a70c83f5f9d1676ae7ca02f

Observation 42d91b02-966b-4f8f-b7a2-7089ef97822f · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:571d3c7633de3671c6015ed7c03768e8338edc38c8a493b9958ff2b7735477fd

Observation 241d1dd4-1b49-4482-b34e-d530dafd3bac · outbound

This paper cites an unresolved cited work.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:c010cd989eddf9d623a702b9b4bdaa13a979f2162b3ad32b305324887cf0b768

Observation 7405044f-6abc-43c6-ae42-c48bb46d52e7 · outbound

This paper cites Scalable diffusion models with transformers.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Scalable diffusion models with transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:b36e8f86b6082ec7e0e981bee14057461d6fd06221705fc285137c91058e4622

Observation 5bcb0028-f47c-47ae-8136-f85b288aec51 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:299c9d7e869ebe82091f19a8bfa8fe7b58fc7cd3c00fef22edbf04c14bd8baab

Observation cf64ae49-015d-46b8-b965-fb4962161767 · outbound

This paper cites DreamFusion: Text-to-3D using 2D Diffusion.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP DreamFusion: Text-to-3D using 2D Diffusion

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:729395210ba859ffb81ca4b93558952d68faf9ece33961f212cb0fbcd7ef7f87

Observation 291c1345-acf7-4652-9137-3668ebd08304 · outbound

This paper cites Gen3c: 3d-informed world- consistent video generation with precise camera control.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Gen3c: 3d-informed world- consistent video generation with precise camera control

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:0790211a590be214048e32e213c70f06063c87e13a62ae9102386327e250a231

Observation 00ee1852-2741-4697-9b90-4b033fdc00b4 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP High- resolution image synthesis with latent diffusion models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:5ad828c925169853e1fd81c9a87575e5e7a37938ffde6f71235f5a1c1c0daa47

Observation cb7738e9-891f-4075-a6e7-4cd629ccecd9 · outbound

This paper cites ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:fe36d7cd0524b9d422b6ee754c2773df713c82c737febeea3402b7f1326fc67c

Observation 9edce66f-3b3c-4440-a622-3b9af687f0cd · outbound

This paper cites A recipe for generating 3d worlds from a single image.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP A recipe for generating 3d worlds from a single image

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:0c058bfa15be00db78fca2b7664b0e59247f7cf7b0a8368ebe2e61ec12a1ee52

Observation 36a22438-d0e2-442f-bff5-9ccc878cb22b · outbound

This paper cites Latent diffusion model without variational autoencoder.arXiv preprint arXiv:2510.15301, 2025.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Latent diffusion model without variational autoencoder.arXiv preprint arXiv:2510.15301, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:608c454e32f5f0e66e46206d8853dad85c649c986714d098be23758518404dad

Observation febed8f3-c62e-47ea-8769-8b05c54db666 · outbound

This paper cites MVDream: Multi-view Diffusion for 3D Generation.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP MVDream: Multi-view Diffusion for 3D Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:1cc4a28eda6ae7c68940f8ca51745500b2b968651f45c58e16c0d974415e3b78

Observation fe3d860c-1d4f-42ee-aa0b-24ab1ed6299f · outbound

This paper cites Denoising diffusion implicit models.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Denoising diffusion implicit models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:efc9c717c29552663626f752f1ba5c90a69908571e3836fc08f1267930fa1151

Observation bd2bd376-8415-4c5e-88fa-c89ecafef589 · outbound

This paper cites Score-based generative modeling through stochastic differential equations.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Score-based generative modeling through stochastic differential equations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:15d95fcec03b315abd1454c1329e79867f2c7a40f3792d2b1eef74381d1f8f47

Observation 27b5c168-b5f7-401f-a108-e3ec61d1256b · outbound

This paper cites DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:c6d7b92c9ad063c6fb94c5223fbb2f5687c8175cfe25970463553d8cac82c452

Observation 5aca08e2-bdaf-46ff-9883-5b0a0f11bef1 · outbound

This paper cites DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:e9f9bbe768e7fda389d4911175e0a73a95f592186d8910460967efac1553c2f5

Observation 7f98a07d-8839-4662-8ff7-063652d095d7 · outbound

This paper cites Neural discrete representation learning.Adv.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Neural discrete representation learning.Adv

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:07176d4ab0d5d85062d13a395ce1686f5691cf6b6b27636b7284568f1b71b0dc

Observation 32ead8e7-b308-4339-8928-12e5a41ba991 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Wan: Open and Advanced Large-Scale Video Generative Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:e55c96ddb4e7621124a3f5732e8d5ad4879b5f0117e4ce4613df29cddf1d28d6

Observation c2bff267-4d36-429f-bc3a-27fcd2b12b62 · outbound

This paper cites Vggt: Visual geometry grounded transformer.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Vggt: Visual geometry grounded transformer

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:389b85219f12ff9b2255f9dd91c1f23bb2b6a40d079abea593baaf1998bdce6e

Observation cf136523-4b43-4ff6-8f5c-e477fe8921b5 · outbound

This paper cites Dust3r: Geometric 3d vision made easy.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Dust3r: Geometric 3d vision made easy

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:2a27596f6c740858f9507e9cc604f3297459f82d781ed00b6ddbc0d23be667c9

Observation 0aa97812-b7d4-4de0-815f-f634c3273e49 · outbound

This paper cites $\pi^3$: Permutation-Equivariant Visual Geometry Learning.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP $\pi^3$: Permutation-Equivariant Visual Geometry Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:07bf49622920aa7f531197ebb1a837b47aa97e9305529357aef495e19d77ff55

Observation be77dd03-3894-43a6-87e0-64dab91c0c2a · outbound

This paper cites Pro- lificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Pro- lificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:9a11a59f840e954156bbdecfa6532e16921105b051153caef6a0e001b94e12d3

Observation 17282fe3-a238-439f-baea-350ccc223453 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4): 600–612, 2004.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4): 600–612, 2004

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:fcbe3e8a83581acae019d2133bcb679a5515143e20a67bad7d0cccb9e5681b17

Observation e2107027-5ffa-4e39-86b5-b8019e1f3fc9 · outbound

This paper cites Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:fe052cdd074b21cb647dc21ce425feb7b4a7c82863fe50c8b77f3d0aa91e6b87

Observation 711fe5a6-4c24-4d32-8652-0ecd8bc9bee6 · outbound

This paper cites Reconfusion: 3d reconstruction with diffusion priors.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Reconfusion: 3d reconstruction with diffusion priors

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:059c60d7cb9865248273e972a121b5edcbb485f2a916b52ae6144bebbea8c520

Observation 4a062beb-da70-45c1-8f47-d1ef9a893ebc · outbound

This paper cites Depthsplat: Connecting gaussian splatting and depth.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Depthsplat: Connecting gaussian splatting and depth

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:98c771a52e9a0c473969a1b09ff546d1ad25c874f6acb530e5c141e4bf190189

Observation 3a096dff-995c-4aa7-a071-25b4c6a4bfc3 · outbound

This paper cites Prometheus: 3d-aware latent diffusion models for feed-forward text-to-3d scene generation.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Prometheus: 3d-aware latent diffusion models for feed-forward text-to-3d scene generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:3a4ea55a5795075734c9f193032b52d9e975fc5eb525b2ca55434d36a5525850

Observation 26ffedac-8ef4-4dac-b2c0-b45c464e7790 · outbound

This paper cites Reconstruction vs.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Reconstruction vs

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:c93f9483d19b16c8bfb72477cb87d6693fbe60ba371c8c2a328465a2c4acfb06

Observation 26f027ec-9d57-4819-a7b3-fc1fbdb273dd · outbound

This paper cites World Action Models are Zero-shot Policies.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP World Action Models are Zero-shot Policies

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:2c450d7049f1ca896003ab55bd7633daab3f1212927a1e87ff53ca0e5e7ca561

Observation d75d5af1-5cc0-4dd3-ab8a-bfa876b7c5e2 · outbound

This paper cites Wonderjourney: Going from anywhere to everywhere.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Wonderjourney: Going from anywhere to everywhere

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:4790d78f5eeb898d1a1f8adea6e34611d7f39e60befd6172732aa710757185a0

Observation 2839f287-de05-46e6-a958-7899040773b3 · outbound

This paper cites Wonder- world: Interactive 3d scene generation from a single image.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Wonder- world: Interactive 3d scene generation from a single image

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:095fc12a7e2c18650697c081e329b9fb518faca9b2c65cba4c0bffdc9b198d5c

Observation e58dabbf-9292-4cf5-be39-22d98aae12cc · outbound

This paper cites Advances in feed-forward 3d reconstruction and view synthesis: A survey.arXiv preprint arXiv:2507.14501, 2025.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Advances in feed-forward 3d reconstruction and view synthesis: A survey.arXiv preprint arXiv:2507.14501, 2025

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:1c619bfaebdbf649f69cfd4847433bd0d6e81ac7eaaebef603d8c9aae8d82b1b

Observation f7cf1882-91b8-4cfd-8dae-33ea18604bb4 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Adding conditional control to text-to-image diffusion models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:2b137600d6fe1f5a92bf185c3ab65b6cb5feab7fcf71f5e77624cfe2e263881f

Observation 02bfc550-75d2-45ff-9d7d-eaa126c4481c · outbound

This paper cites The unreason- able effectiveness of deep features as a perceptual metric.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP The unreason- able effectiveness of deep features as a perceptual metric

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:d3b9dd6fcd3dac09e65d7c38f9d4de06d860e042bb354615301b82ae04a53f96

Observation 9ad08f3d-be9b-4250-933c-3c717d39ad84 · outbound

This paper cites GenXD: Generating Any 3D and 4D Scenes.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP GenXD: Generating Any 3D and 4D Scenes

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:d80b219d300270cae4d28b7054dc5ccac9715e5f158fd206a6ebef900b126acd

Observation 000044d4-64f4-4ef9-bb60-4348ae130c8a · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Diffusion Transformers with Representation Autoencoders

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:9ec59141537126bec23c70b8cd2cefacac5b3e771cf0c13eb1250f0938f4ee94

Observation 96b7fa68-bc50-4e2a-a6bd-8e949685c8df · outbound

This paper cites Stereo magni- fication: learning view synthesis using multiplane images.ACM Trans.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Stereo magni- fication: learning view synthesis using multiplane images.ACM Trans

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:0dc7ff457fcbdbc0773fdfb6851714071243440ad5601d34c2f5026e21e6e643

Observation d8953c77-9d29-4025-8c7f-aa120883129a · outbound

This paper cites Aether: Geometric-aware unified world modeling.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Aether: Geometric-aware unified world modeling

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:4919619484e3d15e52a40bb17a1c377561c9704bdf5ea424fc68ca5834916665

Pith citing papers

No inbound Pith citation observations are available.