{"work":{"id":"d0e5199d-8907-47b1-905a-07ab8b623a4c","openalex_id":"https://openalex.org/W4413155826","doi":"10.1109/cvpr52734.2025","arxiv_id":"2734.2025","raw_key":null,"title":"In: IEEE Conf","authors":null,"authors_text":"Wang, J","year":2025,"venue":null,"abstract":null,"external_url":"https://arxiv.org/abs/2734.2025","cited_by_count":86,"metadata_source":"arxiv_reference","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":null,"created_at":"2026-05-09T19:50:54.829229+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":false,"display_title":"Freeman, Frédo Durand, Eli Shechtman, and Xun Huang","render_title":"Freeman, Frédo Durand, Eli Shechtman, and Xun Huang"},"hub":{"state":{"work_id":"d0e5199d-8907-47b1-905a-07ab8b623a4c","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":194,"external_cited_by_count":86,"distinct_field_count":21,"first_pith_cited_at":"2025-09-28T08:46:11+00:00","last_pith_cited_at":"2026-07-08T16:04:05+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T17:49:17.727930+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":47},{"context_role":"baseline","n":6},{"context_role":"method","n":2},{"context_role":"dataset","n":1},{"context_role":"other","n":1}],"polarity_counts":[{"context_polarity":"background","n":45},{"context_polarity":"baseline","n":6},{"context_polarity":"unclear","n":3},{"context_polarity":"use_method","n":2},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Freeman, Frédo Durand, Eli Shechtman, and Xun Huang","claims":[{"claim_text":"camera parameters are implicitly inferred by the VFM, significantly reducing the complexity compared with conventional calibration workflows. The core idea is to leverage the reasoning capability of a VFM to infer 2D point trajectories and camera parameters from stereo videos, enabling robust 3D displacement reconstruction under real-world conditions. Specif- ically, we adopt VGGT [33], a VFM pretrained on large-scale and diverse datasets spanning multiple vision tasks, as a core component of th","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"speaking, these techniques required learning from labeled CAD datasets [35, 36] which were hard to come by and had sparse cover- age for many types of objects. This led to limitations in generating more complex and / or out-of-distribution objects [14, 39]. More recent work has focused on leveraging pretrained large language models (LLMs) to generate CAD models. CAD-LLama [18], CAD- GPT [33], and others [ 26, 37, 38] finetuned (multimodal) LLMs (MLLMs) to produce code that could be compiled to a","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"While CAD tools have become indispensable for manufacturing- increasing productivity, design precision, and product quality-they can still be tedious to use [ 8, 25]. As a result, the graphics and AI communities have sought to automate the generation of CAD models by training deep generative models to reverse engineer them from images, text descriptions, 3D point clouds, and more [3, 13, 14, 16, 18, 26, 37-39]. To overcome a lack of training data, more recent work has built agentic systems where","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"7 62.3 73.1 68.7 General Video Understanding Ca- pability.A natural concern is whether the perception-sensitivity regularizer, while improving causal discovery, compromises the model's general video understanding capability. To ad- dress this, we evaluate ADPO on a suite of standard video benchmarks, namely Video-MME [32], LongVideoBench [33], MMVU [34], and MVBench [35], covering perception and reasoning over dif- ferent video durations. As shown in Table 5, ADPO achieves better across all benc","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"gaze model for 40 epochs with a batch size (B) of 1,024, and use the AdamW optimizer with an initial learning rate of 1e-4. 6 Table 1: Comparison with the state-of-the-art temporal action detection approaches. We report AP at different tIoU. Method StOP? 0.1 0.2 0.3 0.4 0.5 Avg. ActionFormer [41] 12.63 12.63 12.63 5.79 1.22 8.98 TriDet [34] 23.46 13.37 13.35 13.27 1.87 13.06 TemporalMaxer [35] 16.04 14.36 13.72 13.53 12.38 14.01 Team-OR [6] 30.8520.78 20.66 20.58 20.52 22.68 Ours (off-the-shelf)","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Since these datasets lack ground-truth class labels, class unlearning uses pseudo-class clusters in the pretrained backbone embedding space. Counterparts.We compare our method against nine baselines from two categories. Federated unlearning methods: (1)FedEraser[ 27], (2)FedRecover[ 3], (3)Ferrari[ 28], (4)FedOSD[ 31], (5)NoT[ 20], (6)SoUL[ 17], (7)FFMU[ 5], (8)FUSED[ 60]; centralized unlearning adapted to the federated setting: (9)GradAscent[ 39]; and the retrain reference ˜w. Detailed descript","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Freeman, Frédo Durand, Eli Shechtman, and Xun Huang because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (45 contexts).","role_counts":[{"n":45,"context_role":"background"},{"n":6,"context_role":"baseline"},{"n":2,"context_role":"method"},{"n":1,"context_role":"dataset"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-06-26T09:44:52.325842+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"c1e3e831-3d1a-4c2b-b9a3-34f1265d3f80","orcid":null,"display_name":"Feng"},{"id":"add66286-4b95-4ada-9ef6-13b5057c6359","orcid":null,"display_name":"Chao and Chen"},{"id":"a7badba3-c075-4f5c-846e-abf8dc1c0c60","orcid":null,"display_name":"Ziyang and Ho"}]},"error":null,"updated_at":"2026-06-26T09:45:10.440348+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T11:09:44.180700+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":24},{"title":"& Vondrick, C","work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","shared_citers":18},{"title":"In: 2023 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR)","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":12},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":12},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":8},{"title":"URL https://doi.org/10.48550/arXiv","work_id":"5c2060c6-427c-4321-be22-49ccae439d80","shared_citers":8},{"title":"Masked autoencoders are scalable vision learners","work_id":"0a23d1b7-bd56-43cc-8a80-7c43ce994e1e","shared_citers":7},{"title":"Editing conditional radiance fields","work_id":"3820f598-11b0-45c3-8c99-0079181ac0a7","shared_citers":6},{"title":"In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","work_id":"7083a41e-5666-435b-ab26-c753f6490b9a","shared_citers":6},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":6},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":5},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":4},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":4},{"title":"doi:10.1109/CVPR46437","work_id":"ac0d5a4b-3ab1-462a-bed7-00be8d403f67","shared_citers":4},{"title":"GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning","work_id":"366607ba-e4ea-4726-98c3-63356e32351c","shared_citers":4},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":4},{"title":"IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =","work_id":"9da51225-b7bd-4032-b7db-ca577971dafe","shared_citers":4},{"title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","work_id":"b8f5e260-fff5-444e-bcf5-2c42cfefd83d","shared_citers":4},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","work_id":"3e7e65c5-5aed-4fe9-8414-2092bcb31cc7","shared_citers":4},{"title":"Towards Accurate Generative Models of Video: A New Metric & Challenges","work_id":"72f42543-17d5-49aa-ba5a-25d67ffbb88a","shared_citers":4},{"title":"why should I trust you?","work_id":"238df2e4-a3e5-46f3-860e-3ae2b0094b97","shared_citers":4},{"title":"Ai safety assurance for automated vehicles: A survey on research, standardization, regulation","work_id":"c7d0586e-892f-4740-9491-d608729bd748","shared_citers":3},{"title":"Bias for action: Video implicit neural representations with bias modulation","work_id":"f07730f8-a754-45f3-a4b9-1c9a68e55468","shared_citers":3},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":3}],"time_series":[{"n":58,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T11:09:33.431915+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T11:09:37.663847+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Freeman, Frédo Durand, Eli Shechtman, and Xun Huang","claims":[{"claim_text":"camera parameters are implicitly inferred by the VFM, significantly reducing the complexity compared with conventional calibration workflows. The core idea is to leverage the reasoning capability of a VFM to infer 2D point trajectories and camera parameters from stereo videos, enabling robust 3D displacement reconstruction under real-world conditions. Specif- ically, we adopt VGGT [33], a VFM pretrained on large-scale and diverse datasets spanning multiple vision tasks, as a core component of th","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"speaking, these techniques required learning from labeled CAD datasets [35, 36] which were hard to come by and had sparse cover- age for many types of objects. This led to limitations in generating more complex and / or out-of-distribution objects [14, 39]. More recent work has focused on leveraging pretrained large language models (LLMs) to generate CAD models. CAD-LLama [18], CAD- GPT [33], and others [ 26, 37, 38] finetuned (multimodal) LLMs (MLLMs) to produce code that could be compiled to a","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"While CAD tools have become indispensable for manufacturing- increasing productivity, design precision, and product quality-they can still be tedious to use [ 8, 25]. As a result, the graphics and AI communities have sought to automate the generation of CAD models by training deep generative models to reverse engineer them from images, text descriptions, 3D point clouds, and more [3, 13, 14, 16, 18, 26, 37-39]. To overcome a lack of training data, more recent work has built agentic systems where","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"7 62.3 73.1 68.7 General Video Understanding Ca- pability.A natural concern is whether the perception-sensitivity regularizer, while improving causal discovery, compromises the model's general video understanding capability. To ad- dress this, we evaluate ADPO on a suite of standard video benchmarks, namely Video-MME [32], LongVideoBench [33], MMVU [34], and MVBench [35], covering perception and reasoning over dif- ferent video durations. As shown in Table 5, ADPO achieves better across all benc","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"gaze model for 40 epochs with a batch size (B) of 1,024, and use the AdamW optimizer with an initial learning rate of 1e-4. 6 Table 1: Comparison with the state-of-the-art temporal action detection approaches. We report AP at different tIoU. Method StOP? 0.1 0.2 0.3 0.4 0.5 Avg. ActionFormer [41] 12.63 12.63 12.63 5.79 1.22 8.98 TriDet [34] 23.46 13.37 13.35 13.27 1.87 13.06 TemporalMaxer [35] 16.04 14.36 13.72 13.53 12.38 14.01 Team-OR [6] 30.8520.78 20.66 20.58 20.52 22.68 Ours (off-the-shelf)","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Since these datasets lack ground-truth class labels, class unlearning uses pseudo-class clusters in the pretrained backbone embedding space. Counterparts.We compare our method against nine baselines from two categories. Federated unlearning methods: (1)FedEraser[ 27], (2)FedRecover[ 3], (3)Ferrari[ 28], (4)FedOSD[ 31], (5)NoT[ 20], (6)SoUL[ 17], (7)FFMU[ 5], (8)FUSED[ 60]; centralized unlearning adapted to the federated setting: (9)GradAscent[ 39]; and the retrain reference ˜w. Detailed descript","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Freeman, Frédo Durand, Eli Shechtman, and Xun Huang because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (45 contexts).","role_counts":[{"n":45,"context_role":"background"},{"n":6,"context_role":"baseline"},{"n":2,"context_role":"method"},{"n":1,"context_role":"dataset"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-06-26T09:44:57.097143+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"MambaVision: A hybrid Mamba- Transformer vision backbone","claims":[],"why_cited":"Pith tracks MambaVision: A hybrid Mamba- Transformer vision backbone because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T11:09:42.278132+00:00"}},"summary":{"title":"MambaVision: A hybrid Mamba- Transformer vision backbone","claims":[],"why_cited":"Pith tracks MambaVision: A hybrid Mamba- Transformer vision backbone because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":24},{"title":"& Vondrick, C","work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","shared_citers":18},{"title":"In: 2023 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR)","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":12},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":12},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":8},{"title":"URL https://doi.org/10.48550/arXiv","work_id":"5c2060c6-427c-4321-be22-49ccae439d80","shared_citers":8},{"title":"Masked autoencoders are scalable vision learners","work_id":"0a23d1b7-bd56-43cc-8a80-7c43ce994e1e","shared_citers":7},{"title":"Editing conditional radiance fields","work_id":"3820f598-11b0-45c3-8c99-0079181ac0a7","shared_citers":6},{"title":"In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","work_id":"7083a41e-5666-435b-ab26-c753f6490b9a","shared_citers":6},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":6},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":5},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":4},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":4},{"title":"doi:10.1109/CVPR46437","work_id":"ac0d5a4b-3ab1-462a-bed7-00be8d403f67","shared_citers":4},{"title":"GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning","work_id":"366607ba-e4ea-4726-98c3-63356e32351c","shared_citers":4},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":4},{"title":"IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =","work_id":"9da51225-b7bd-4032-b7db-ca577971dafe","shared_citers":4},{"title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","work_id":"b8f5e260-fff5-444e-bcf5-2c42cfefd83d","shared_citers":4},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","work_id":"3e7e65c5-5aed-4fe9-8414-2092bcb31cc7","shared_citers":4},{"title":"Towards Accurate Generative Models of Video: A New Metric & Challenges","work_id":"72f42543-17d5-49aa-ba5a-25d67ffbb88a","shared_citers":4},{"title":"why should I trust you?","work_id":"238df2e4-a3e5-46f3-860e-3ae2b0094b97","shared_citers":4},{"title":"Ai safety assurance for automated vehicles: A survey on research, standardization, regulation","work_id":"c7d0586e-892f-4740-9491-d608729bd748","shared_citers":3},{"title":"Bias for action: Video implicit neural representations with bias modulation","work_id":"f07730f8-a754-45f3-a4b9-1c9a68e55468","shared_citers":3},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":3}],"time_series":[{"n":58,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"add66286-4b95-4ada-9ef6-13b5057c6359","orcid":null,"display_name":"Chao and Chen","source":"manual","import_confidence":0.72},{"id":"c1e3e831-3d1a-4c2b-b9a3-34f1265d3f80","orcid":null,"display_name":"Feng","source":"manual","import_confidence":0.72},{"id":"a7badba3-c075-4f5c-846e-abf8dc1c0c60","orcid":null,"display_name":"Ziyang and Ho","source":"manual","import_confidence":0.72}]}}