{"work":{"id":"0a54b500-1e9d-46c2-85eb-8e16cbac8461","openalex_id":"https://openalex.org/W7105740705","doi":"10.48550/arxiv.2511.10647","arxiv_id":"2511.10647","raw_key":null,"title":"Depth Anything 3: Recovering the Visual Space from Any Views","authors":null,"authors_text":"Haotong Lin, Sili Chen, Junhao Liew, Donny Y. Chen, Zhenyu Li, Guang Shi","year":2025,"venue":"cs.CV","abstract":"We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization, and a singular depth-ray prediction target obviates the need for complex multi-task learning. Through our teacher-student training paradigm, the model achieves a level of detail and generalization on par with Depth Anything 2 (DA2). We establish a new visual geometry benchmark covering camera pose estimation, any-view geometry and visual rendering. On this benchmark, DA3 sets a new state-of-the-art across all tasks, surpassing prior SOTA VGGT by an average of 44.3% in camera pose accuracy and 25.1% in geometric accuracy. Moreover, it outperforms DA2 in monocular depth estimation. All models are trained exclusively on public academic datasets.","external_url":"https://arxiv.org/abs/2511.10647","cited_by_count":0,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2511.10647","created_at":"2026-05-09T06:35:39.534275+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Depth Anything 3: Recovering the Visual Space from Any Views","render_title":"Depth Anything 3: Recovering the Visual Space from Any Views"},"hub":{"state":{"work_id":"0a54b500-1e9d-46c2-85eb-8e16cbac8461","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":217,"external_cited_by_count":0,"distinct_field_count":5,"first_pith_cited_at":"2025-05-27T14:51:34+00:00","last_pith_cited_at":"2026-07-08T14:22:05+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T21:09:45.757413+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":13},{"context_role":"method","n":13},{"context_role":"baseline","n":4},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"use_method","n":13},{"context_polarity":"background","n":12},{"context_polarity":"baseline","n":4},{"context_polarity":"unclear","n":1},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Depth Anything 3: Recovering the Visual Space from Any Views","claims":[{"claim_text":"We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization, and a singular depth-ray prediction target obviates the need for complex multi-task learning. Through our teacher-student training paradigm, the model achieves a level of detail and generalization on par with Depth Anything 2 (DA2). We establish a new","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Left:Given a text prompt, we synthesize a camera trajectory and embed it into the latent space via noise wrapping [34] to achieve implicit camera conditioning.Middle:The foundation model [3] generates candidate video clips.Right:We employ an comprehensive reward system: the generated video is lifted to a 3D Gaussian Splatting representation via a 3D Foundation Model [17]. We compute a 3D-aware reward score based on meta-view evaluation, rendering fidelity, and trajectory alignment, combined with","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"framework designed to balance local target sensitivity with global geometric coherence. Our key intuition is to combine the complementary priors of promptable segmentation and monocular depth foundation models: Segment Anything Model 3 (SAM3) [ 4] provides prompt-grounded spatial selectivity for identifying user-specified target regions, while the Depth Anything family (DAs), such as DA2 [ 34] and DA3 [ 16], provides a strong pretrained prior over dense scene geometry. However, directly fusing t","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"Equipped with these mechanisms, our model achieves highly persistent and long-horizon scene generation. Nevertheless, videos synthesized by diffusion models inevitably contain minor multi-view inconsistencies that easily break traditional 3D reconstruction models, causing floaters and noisy artifacts. To achieve reliable scene reconstruction, we employ a feed-forward 3D Gaussian Splatting (3DGS) pipeline [58]. Fine-tuned on our generated sequences, this feed-forward model leverages its learned m","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"The stitched image retains the reliable projected structure and incorporates generated content only where needed, yielding a more coherent final result. 3.1 Generate the Initial-View 4D Asset We first generate an initial-view 4D asset from the source video, which can be either observed or generated. For each frame t, we estimate the source-view camera and depth using DA3 [19]. We then construct a triangular mesh on the image lattice by connecting neighboring valid pixels into faces. Each mesh ve","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"into a full 3D objectx H faithful to the input data. We are motivated by the fact that current image-to-3D generators produce high-quality and detailed 3D shapes that are however not necessarily faithful to the input image, and/or cannot account for more than one image as input. On the other hand, large 3D reconstruction neural networks like VGGT [58] and many others [22,33,61] reconstruct the visible geometry as faithfully as possible from several views, resulting in 3D reconstructionsyL that a","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"and is adopted as the canonical satellite camera height. Metric depth and tri-view pairing.Each modality receives its dense metric depth through a dedicated pipeline (Fig. 11).Ground depth(Fig. 2(a)) starts from Street View, which ships an absolute-scale but coarse depth that captures the overall scene scale yet misses fine structure (distant buildings, street lamps); we therefore run DEPTHANYTHING V3 [ 21] on each ground frame to obtain a sharp butrelativedepth, and fuse the two throughPrior De","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Depth Anything 3: Recovering the Visual Space from Any Views because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (13 contexts).","role_counts":[{"n":13,"context_role":"background"},{"n":13,"context_role":"method"},{"n":4,"context_role":"baseline"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-06-29T13:59:02.502787+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"3ccd9062-70b9-44f6-bf93-a994692dd14c","orcid":null,"display_name":"Haotong Lin"},{"id":"8d38954f-6478-4411-8944-3b772ee015d1","orcid":null,"display_name":"Sili Chen"},{"id":"65e3efce-2691-4492-96b3-489f8edbd715","orcid":null,"display_name":"Junhao Liew"},{"id":"6222e6f2-99a3-4a8a-9aa5-bc506d878609","orcid":null,"display_name":"Donny Y. Chen"},{"id":"638ea729-1cb6-4282-9c38-41ac9500e61f","orcid":null,"display_name":"Zhenyu Li"},{"id":"3e5bbff8-b1ca-48ed-8e25-24fa6e542bca","orcid":null,"display_name":"Guang Shi"}]},"error":null,"updated_at":"2026-06-29T13:59:02.496047+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T11:39:50.837704+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"MapAnything: Universal Feed-Forward Metric 3D Reconstruction","work_id":"cb742a0f-0648-4cb0-904b-180184987f6b","shared_citers":15},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":15},{"title":"$\\pi^3$: Permutation-Equivariant Visual Geometry Learning","work_id":"8ab9cfd6-a60d-45e6-a572-9b4db82b8527","shared_citers":14},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":11},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":9},{"title":"arXiv2506.15442(2025) 10","work_id":"ee52f4d7-462f-4491-9549-4160820ae563","shared_citers":7},{"title":"DINOv3","work_id":"c8b07deb-8fe7-4e18-9620-f3569d3529ce","shared_citers":7},{"title":"SAM 3: Segment Anything with Concepts","work_id":"4a72a006-2592-4554-aad0-a9c41a9f952d","shared_citers":7},{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","work_id":"4f68eada-27e3-437a-a2fe-6e4ca524d0d3","shared_citers":7},{"title":"Virtual KITTI 2","work_id":"c0d9c030-aa25-44e7-9cc4-72d7403f1447","shared_citers":7},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":6},{"title":"Moge-2: Accurate monocular geometry with metric scale and sharp details.arXiv preprint arXiv:2507.02546","work_id":"4a287a38-b116-4830-99a0-306010a0c8e5","shared_citers":6},{"title":"Qwen-Image Technical Report","work_id":"d06d7ecc-7579-4f89-a60b-4278a0f3c562","shared_citers":6},{"title":"SAM 3D: 3Dfy Anything in Images","work_id":"dc22e9ff-fcf5-4069-8ace-35ae3a0bfd7c","shared_citers":6},{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":5},{"title":"Anysplat: Feed-forward 3d gaussian splatting from unconstrained views.arXiv preprint arXiv:2505.23716","work_id":"60813889-15bf-4c3b-b3a8-a141aab96316","shared_citers":5},{"title":"CameraCtrl: Enabling Camera Control for Text-to-Video Generation","work_id":"1c05c278-c023-4ef0-a359-25a41f1065eb","shared_citers":5},{"title":"Cosmos World Foundation Model Platform for Physical AI","work_id":"a2dba24c-318d-476a-8b21-4289c265810c","shared_citers":5},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":5},{"title":"Depth Anything V2","work_id":"5f5274d6-f7a5-4598-a3f6-d11f44520a2e","shared_citers":5},{"title":"Depth pro: Sharp monocular metric depth in less than a second","work_id":"0b67883b-1901-45f1-9d58-1ef7a928df23","shared_citers":5},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":5},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":5},{"title":"Lvsm: A large view synthesis model with minimal 3d inductive bias","work_id":"c5e225cf-22b3-4efa-a3c4-5346587616d0","shared_citers":5}],"time_series":[{"n":1,"year":2025},{"n":54,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T11:39:55.465096+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T11:39:59.630555+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Depth Anything 3: Recovering the Visual Space from Any Views","claims":[{"claim_text":"We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization, and a singular depth-ray prediction target obviates the need for complex multi-task learning. Through our teacher-student training paradigm, the model achieves a level of detail and generalization on par with Depth Anything 2 (DA2). We establish a new","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Left:Given a text prompt, we synthesize a camera trajectory and embed it into the latent space via noise wrapping [34] to achieve implicit camera conditioning.Middle:The foundation model [3] generates candidate video clips.Right:We employ an comprehensive reward system: the generated video is lifted to a 3D Gaussian Splatting representation via a 3D Foundation Model [17]. We compute a 3D-aware reward score based on meta-view evaluation, rendering fidelity, and trajectory alignment, combined with","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"framework designed to balance local target sensitivity with global geometric coherence. Our key intuition is to combine the complementary priors of promptable segmentation and monocular depth foundation models: Segment Anything Model 3 (SAM3) [ 4] provides prompt-grounded spatial selectivity for identifying user-specified target regions, while the Depth Anything family (DAs), such as DA2 [ 34] and DA3 [ 16], provides a strong pretrained prior over dense scene geometry. However, directly fusing t","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"Equipped with these mechanisms, our model achieves highly persistent and long-horizon scene generation. Nevertheless, videos synthesized by diffusion models inevitably contain minor multi-view inconsistencies that easily break traditional 3D reconstruction models, causing floaters and noisy artifacts. To achieve reliable scene reconstruction, we employ a feed-forward 3D Gaussian Splatting (3DGS) pipeline [58]. Fine-tuned on our generated sequences, this feed-forward model leverages its learned m","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"The stitched image retains the reliable projected structure and incorporates generated content only where needed, yielding a more coherent final result. 3.1 Generate the Initial-View 4D Asset We first generate an initial-view 4D asset from the source video, which can be either observed or generated. For each frame t, we estimate the source-view camera and depth using DA3 [19]. We then construct a triangular mesh on the image lattice by connecting neighboring valid pixels into faces. Each mesh ve","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"into a full 3D objectx H faithful to the input data. We are motivated by the fact that current image-to-3D generators produce high-quality and detailed 3D shapes that are however not necessarily faithful to the input image, and/or cannot account for more than one image as input. On the other hand, large 3D reconstruction neural networks like VGGT [58] and many others [22,33,61] reconstruct the visible geometry as faithfully as possible from several views, resulting in 3D reconstructionsyL that a","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"and is adopted as the canonical satellite camera height. Metric depth and tri-view pairing.Each modality receives its dense metric depth through a dedicated pipeline (Fig. 11).Ground depth(Fig. 2(a)) starts from Street View, which ships an absolute-scale but coarse depth that captures the overall scene scale yet misses fine structure (distant buildings, street lamps); we therefore run DEPTHANYTHING V3 [ 21] on each ground frame to obtain a sharp butrelativedepth, and fuse the two throughPrior De","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Depth Anything 3: Recovering the Visual Space from Any Views because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (13 contexts).","role_counts":[{"n":13,"context_role":"background"},{"n":13,"context_role":"method"},{"n":4,"context_role":"baseline"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-06-29T13:59:02.499984+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Depth Anything 3: Recovering the Visual Space from Any Views","claims":[{"claim_text":"We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization, and a singular depth-ray prediction target obviates the need for complex multi-task learning. Through our teacher-student training paradigm, the model achieves a level of detail and generalization on par with Depth Anything 2 (DA2). We establish a new","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Depth Anything 3: Recovering the Visual Space from Any Views because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T11:39:59.634947+00:00"}},"summary":{"title":"Depth Anything 3: Recovering the Visual Space from Any Views","claims":[{"claim_text":"We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization, and a singular depth-ray prediction target obviates the need for complex multi-task learning. Through our teacher-student training paradigm, the model achieves a level of detail and generalization on par with Depth Anything 2 (DA2). We establish a new","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Depth Anything 3: Recovering the Visual Space from Any Views because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"MapAnything: Universal Feed-Forward Metric 3D Reconstruction","work_id":"cb742a0f-0648-4cb0-904b-180184987f6b","shared_citers":15},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":15},{"title":"$\\pi^3$: Permutation-Equivariant Visual Geometry Learning","work_id":"8ab9cfd6-a60d-45e6-a572-9b4db82b8527","shared_citers":14},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":11},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":9},{"title":"arXiv2506.15442(2025) 10","work_id":"ee52f4d7-462f-4491-9549-4160820ae563","shared_citers":7},{"title":"DINOv3","work_id":"c8b07deb-8fe7-4e18-9620-f3569d3529ce","shared_citers":7},{"title":"SAM 3: Segment Anything with Concepts","work_id":"4a72a006-2592-4554-aad0-a9c41a9f952d","shared_citers":7},{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","work_id":"4f68eada-27e3-437a-a2fe-6e4ca524d0d3","shared_citers":7},{"title":"Virtual KITTI 2","work_id":"c0d9c030-aa25-44e7-9cc4-72d7403f1447","shared_citers":7},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":6},{"title":"Moge-2: Accurate monocular geometry with metric scale and sharp details.arXiv preprint arXiv:2507.02546","work_id":"4a287a38-b116-4830-99a0-306010a0c8e5","shared_citers":6},{"title":"Qwen-Image Technical Report","work_id":"d06d7ecc-7579-4f89-a60b-4278a0f3c562","shared_citers":6},{"title":"SAM 3D: 3Dfy Anything in Images","work_id":"dc22e9ff-fcf5-4069-8ace-35ae3a0bfd7c","shared_citers":6},{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":5},{"title":"Anysplat: Feed-forward 3d gaussian splatting from unconstrained views.arXiv preprint arXiv:2505.23716","work_id":"60813889-15bf-4c3b-b3a8-a141aab96316","shared_citers":5},{"title":"CameraCtrl: Enabling Camera Control for Text-to-Video Generation","work_id":"1c05c278-c023-4ef0-a359-25a41f1065eb","shared_citers":5},{"title":"Cosmos World Foundation Model Platform for Physical AI","work_id":"a2dba24c-318d-476a-8b21-4289c265810c","shared_citers":5},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":5},{"title":"Depth Anything V2","work_id":"5f5274d6-f7a5-4598-a3f6-d11f44520a2e","shared_citers":5},{"title":"Depth pro: Sharp monocular metric depth in less than a second","work_id":"0b67883b-1901-45f1-9d58-1ef7a928df23","shared_citers":5},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":5},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":5},{"title":"Lvsm: A large view synthesis model with minimal 3d inductive bias","work_id":"c5e225cf-22b3-4efa-a3c4-5346587616d0","shared_citers":5}],"time_series":[{"n":1,"year":2025},{"n":54,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"6222e6f2-99a3-4a8a-9aa5-bc506d878609","orcid":null,"display_name":"Donny Y. Chen","source":"manual","import_confidence":0.72},{"id":"3e5bbff8-b1ca-48ed-8e25-24fa6e542bca","orcid":null,"display_name":"Guang Shi","source":"manual","import_confidence":0.72},{"id":"3ccd9062-70b9-44f6-bf93-a994692dd14c","orcid":null,"display_name":"Haotong Lin","source":"manual","import_confidence":0.72},{"id":"65e3efce-2691-4492-96b3-489f8edbd715","orcid":null,"display_name":"Junhao Liew","source":"manual","import_confidence":0.72},{"id":"8d38954f-6478-4411-8944-3b772ee015d1","orcid":null,"display_name":"Sili Chen","source":"manual","import_confidence":0.72},{"id":"638ea729-1cb6-4282-9c38-41ac9500e61f","orcid":null,"display_name":"Zhenyu Li","source":"manual","import_confidence":0.72}]}}