{"work":{"id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","openalex_id":"https://openalex.org/W4390874170","doi":"10.1109/iccv51070.2023","arxiv_id":"1070.2023","raw_key":null,"title":"Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross B","authors":null,"authors_text":"Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chlo ´e Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C","year":2023,"venue":null,"abstract":null,"external_url":"https://arxiv.org/abs/1070.2023","cited_by_count":360,"metadata_source":"arxiv_reference","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":null,"created_at":"2026-05-09T18:55:07.410796+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":false,"display_title":"In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV)","render_title":"In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV)"},"hub":{"state":{"work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":246,"external_cited_by_count":360,"distinct_field_count":22,"first_pith_cited_at":"2024-03-05T18:45:39+00:00","last_pith_cited_at":"2026-07-08T10:13:20+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-10T15:29:08.365887+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":44},{"context_role":"method","n":10},{"context_role":"baseline","n":2},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"background","n":44},{"context_polarity":"use_method","n":10},{"context_polarity":"baseline","n":2},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Ddp: Diffusion model for dense visual prediction","claims":[{"claim_text":"Video (R) [76], Top&Random (R) [77] ,Reddit Images (R) [78], UIV (C) [74] Emotions and Social Signals Affective Analysis Pitts Ads Dataset (C) [64], Video Emotion Dataset (C) [11], Ekman Emotion Dataset (C) [11], VAAD (C) [79], iMiGUE (C) [80], EALD (Q) [81], VCE (C) [82], V2V (R) [82], VEATIC (R) [83], MERR (C,Cap) [14], 3MASSIV (C) [70], LAMBDA (Q) [63], ArtEmis (C,Cap) [84], EmoSet (C) [85] Relationships SRIV (C) [86], ViSR (C) [87], PERR (C) [88], MovieGraphs (Q) [89], LVU (C) [66], VideoAds","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Qwen 2.5-VL.Bai et al . [8] enhanced the Qwen 2-VL architecture with a redesigned vision encoder that applies self-attention only to four transformer layers while employing windowed attention with 112 × 112-pixel windows (corresponding to 8× 8 patches) for the remaining layers. Additional optimizations include replacing layer normalization with RMSNorm [166] and standard activation functions with SwiGLU [34] in feed-forward layers. Visual features use the same M-RoPE positional encoding before p","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"AVRA distribution and corresponding structural patterns to capture clearly dis- tinct atrophy states, and our conclusions are not expected to depend strongly on small variations of these. Subjects with0.35<A VRA<0.8were labelled as N-intermediate. Two3Dimagesynthesisarchitectureswereevaluated:thewidelybench- marked conditional GAN pix2pix [6] and a latent diffusion model (LDM) [14,24], which has recently been applied to amyloid PET synthesis [12,16]. All models were trained using identical data ","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"horizon pretraining that blends self-supervision, language-image alignment, and promptable segmentation. Recent work such as DINOv2 [33], EVA-02 [9], I- JEPA [1], and InternImage [48] scales masked or predictive objectives to large corporaandyieldshighlytransferablefeatures.Vision-languageefforts(SigLIP[54], PaLI-3 [4]) and unified decoders (Florence-2 [50]) broaden this to multi-task in- terfaces, while Segment Anything [24] shows promptable segmentation at scale. Within this landscape, we posi","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Model-zoo and hyper-representation work showed that trained weights can be character- ized and embedded directly [6, 32, 34, 38]; neural-functional and graph-based readers then made structure and symmetry explicit modeling concerns [20, 22, 26, 42, 43]. The weight-space learning view consolidates understanding, representation, and generation directions [8, 11, 30, 33], but also exposes a tension: useful task structure in weights need not be raw coordinate geometry. Implicit neural representation","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"\"The Unreasonable Effectiveness of Deep Features as a Perceptual Metric\". In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2018, pp. 586- 595. [28] William Peebles and Saining Xie. \"Scalable Diffusion Models with Transformers\". In:Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023, pp. 4195-4205. [29] William Peebles and Saining Xie. \"Scalable Diffusion Models with Transformers\". In:Proceedings of the IEEE/CVF International Confere","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Ddp: Diffusion model for dense visual prediction because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (42 contexts).","role_counts":[{"n":42,"context_role":"background"},{"n":9,"context_role":"method"},{"n":2,"context_role":"baseline"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-05-20T15:22:01.356842+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"3a34b6a8-5a99-48f2-a01e-f01c89b034bc","orcid":null,"display_name":"Chandan Yeshwanth"},{"id":"2036f1c8-58bd-4058-a65f-6a326afd1185","orcid":null,"display_name":"Yueh-Cheng Liu"},{"id":"a325f85a-4f0c-41b5-90a9-1a7672185a56","orcid":null,"display_name":"Matthias Nießner"},{"id":"14b06e38-eff3-4a6a-9c9f-52bd43895a01","orcid":null,"display_name":"and Angela Dai"}]},"error":null,"updated_at":"2026-05-20T15:22:03.627046+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T18:49:50.125530+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"URL https://doi.org/10.1109/CVPR52733","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":24},{"title":"Barron, Ben Mildenhall, Mehdi S","work_id":"0a23d1b7-bd56-43cc-8a80-7c43ce994e1e","shared_citers":18},{"title":"OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations","work_id":"d0e5199d-8907-47b1-905a-07ab8b623a4c","shared_citers":18},{"title":"ImageBind One Embedding Space to Bind Them All","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":16},{"title":"Adabins: Depth estimation using adap- tive bins","work_id":"7083a41e-5666-435b-ab26-c753f6490b9a","shared_citers":13},{"title":"In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","work_id":"9da51225-b7bd-4032-b7db-ca577971dafe","shared_citers":12},{"title":"In: 2021 IEEE/CVF In- ternational Conference on Computer Vision (ICCV)","work_id":"3820f598-11b0-45c3-8c99-0079181ac0a7","shared_citers":8},{"title":"URLhttps://doi.org/10.48550/arXiv","work_id":"5c2060c6-427c-4321-be22-49ccae439d80","shared_citers":6},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":5},{"title":"Enforcing geometric constraints of vir- tual normal for depth prediction","work_id":"135418b1-cafd-49fd-803d-1ca6433d4b1b","shared_citers":5},{"title":"Learning Transferable Visual Models From Natural Language Supervision","work_id":"6de86bb5-27bd-4d5c-8b89-967ebfc52659","shared_citers":5},{"title":"Reid, and Silvio Savarese","work_id":"45b0bfd8-65dc-4252-b2ab-2f6b411d04d0","shared_citers":5},{"title":"Distilling the Knowledge in a Neural Network","work_id":"d927ab1f-17b8-4002-9d09-c3d55764fbad","shared_citers":4},{"title":"EW Dijkstra","work_id":"effdb28b-742e-4840-b3ca-d89502a6cd4d","shared_citers":4},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":4},{"title":"Lawrence Zitnick","work_id":"9a41dfbd-d7c8-4b32-943f-59c08fdd6db7","shared_citers":4},{"title":"Pattern Recognition 127 (2022), 108611","work_id":"238df2e4-a3e5-46f3-860e-3ae2b0094b97","shared_citers":4},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":4},{"title":"URLhttp://dx.doi.org/10.1109/CVPR.2016.90","work_id":"b353bda2-591d-479a-9c8b-22dfcba12431","shared_citers":4},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":3},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":3},{"title":"doi:10.1109/CVPR46437","work_id":"ac0d5a4b-3ab1-462a-bed7-00be8d403f67","shared_citers":3},{"title":"In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","work_id":"caaf86e4-4cdb-450e-80c7-d20d4375abae","shared_citers":3},{"title":"IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models","work_id":"98e51b10-54bd-4251-8a2d-f79bd6215c19","shared_citers":3}],"time_series":[{"n":1,"year":2024},{"n":1,"year":2025},{"n":57,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T18:49:42.652957+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T18:49:30.153785+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Ddp: Diffusion model for dense visual prediction","claims":[{"claim_text":"Video (R) [76], Top&Random (R) [77] ,Reddit Images (R) [78], UIV (C) [74] Emotions and Social Signals Affective Analysis Pitts Ads Dataset (C) [64], Video Emotion Dataset (C) [11], Ekman Emotion Dataset (C) [11], VAAD (C) [79], iMiGUE (C) [80], EALD (Q) [81], VCE (C) [82], V2V (R) [82], VEATIC (R) [83], MERR (C,Cap) [14], 3MASSIV (C) [70], LAMBDA (Q) [63], ArtEmis (C,Cap) [84], EmoSet (C) [85] Relationships SRIV (C) [86], ViSR (C) [87], PERR (C) [88], MovieGraphs (Q) [89], LVU (C) [66], VideoAds","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Qwen 2.5-VL.Bai et al . [8] enhanced the Qwen 2-VL architecture with a redesigned vision encoder that applies self-attention only to four transformer layers while employing windowed attention with 112 × 112-pixel windows (corresponding to 8× 8 patches) for the remaining layers. Additional optimizations include replacing layer normalization with RMSNorm [166] and standard activation functions with SwiGLU [34] in feed-forward layers. Visual features use the same M-RoPE positional encoding before p","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"AVRA distribution and corresponding structural patterns to capture clearly dis- tinct atrophy states, and our conclusions are not expected to depend strongly on small variations of these. Subjects with0.35<A VRA<0.8were labelled as N-intermediate. Two3Dimagesynthesisarchitectureswereevaluated:thewidelybench- marked conditional GAN pix2pix [6] and a latent diffusion model (LDM) [14,24], which has recently been applied to amyloid PET synthesis [12,16]. All models were trained using identical data ","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"horizon pretraining that blends self-supervision, language-image alignment, and promptable segmentation. Recent work such as DINOv2 [33], EVA-02 [9], I- JEPA [1], and InternImage [48] scales masked or predictive objectives to large corporaandyieldshighlytransferablefeatures.Vision-languageefforts(SigLIP[54], PaLI-3 [4]) and unified decoders (Florence-2 [50]) broaden this to multi-task in- terfaces, while Segment Anything [24] shows promptable segmentation at scale. Within this landscape, we posi","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Model-zoo and hyper-representation work showed that trained weights can be character- ized and embedded directly [6, 32, 34, 38]; neural-functional and graph-based readers then made structure and symmetry explicit modeling concerns [20, 22, 26, 42, 43]. The weight-space learning view consolidates understanding, representation, and generation directions [8, 11, 30, 33], but also exposes a tension: useful task structure in weights need not be raw coordinate geometry. Implicit neural representation","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"\"The Unreasonable Effectiveness of Deep Features as a Perceptual Metric\". In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2018, pp. 586- 595. [28] William Peebles and Saining Xie. \"Scalable Diffusion Models with Transformers\". In:Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023, pp. 4195-4205. [29] William Peebles and Saining Xie. \"Scalable Diffusion Models with Transformers\". In:Proceedings of the IEEE/CVF International Confere","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Ddp: Diffusion model for dense visual prediction because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (42 contexts).","role_counts":[{"n":42,"context_role":"background"},{"n":9,"context_role":"method"},{"n":2,"context_role":"baseline"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-05-20T15:22:01.360031+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Ddp: Diffusion model for dense visual prediction","claims":[],"why_cited":"Pith tracks Ddp: Diffusion model for dense visual prediction because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T18:49:42.595138+00:00"}},"summary":{"title":"Ddp: Diffusion model for dense visual prediction","claims":[],"why_cited":"Pith tracks Ddp: Diffusion model for dense visual prediction because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"URL https://doi.org/10.1109/CVPR52733","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":24},{"title":"Barron, Ben Mildenhall, Mehdi S","work_id":"0a23d1b7-bd56-43cc-8a80-7c43ce994e1e","shared_citers":18},{"title":"OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations","work_id":"d0e5199d-8907-47b1-905a-07ab8b623a4c","shared_citers":18},{"title":"ImageBind One Embedding Space to Bind Them All","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":16},{"title":"Adabins: Depth estimation using adap- tive bins","work_id":"7083a41e-5666-435b-ab26-c753f6490b9a","shared_citers":13},{"title":"In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","work_id":"9da51225-b7bd-4032-b7db-ca577971dafe","shared_citers":12},{"title":"In: 2021 IEEE/CVF In- ternational Conference on Computer Vision (ICCV)","work_id":"3820f598-11b0-45c3-8c99-0079181ac0a7","shared_citers":8},{"title":"URLhttps://doi.org/10.48550/arXiv","work_id":"5c2060c6-427c-4321-be22-49ccae439d80","shared_citers":6},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":5},{"title":"Enforcing geometric constraints of vir- tual normal for depth prediction","work_id":"135418b1-cafd-49fd-803d-1ca6433d4b1b","shared_citers":5},{"title":"Learning Transferable Visual Models From Natural Language Supervision","work_id":"6de86bb5-27bd-4d5c-8b89-967ebfc52659","shared_citers":5},{"title":"Reid, and Silvio Savarese","work_id":"45b0bfd8-65dc-4252-b2ab-2f6b411d04d0","shared_citers":5},{"title":"Distilling the Knowledge in a Neural Network","work_id":"d927ab1f-17b8-4002-9d09-c3d55764fbad","shared_citers":4},{"title":"EW Dijkstra","work_id":"effdb28b-742e-4840-b3ca-d89502a6cd4d","shared_citers":4},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":4},{"title":"Lawrence Zitnick","work_id":"9a41dfbd-d7c8-4b32-943f-59c08fdd6db7","shared_citers":4},{"title":"Pattern Recognition 127 (2022), 108611","work_id":"238df2e4-a3e5-46f3-860e-3ae2b0094b97","shared_citers":4},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":4},{"title":"URLhttp://dx.doi.org/10.1109/CVPR.2016.90","work_id":"b353bda2-591d-479a-9c8b-22dfcba12431","shared_citers":4},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":3},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":3},{"title":"doi:10.1109/CVPR46437","work_id":"ac0d5a4b-3ab1-462a-bed7-00be8d403f67","shared_citers":3},{"title":"In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","work_id":"caaf86e4-4cdb-450e-80c7-d20d4375abae","shared_citers":3},{"title":"IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models","work_id":"98e51b10-54bd-4251-8a2d-f79bd6215c19","shared_citers":3}],"time_series":[{"n":1,"year":2024},{"n":1,"year":2025},{"n":57,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"14b06e38-eff3-4a6a-9c9f-52bd43895a01","orcid":null,"display_name":"and Angela Dai","source":"manual","import_confidence":0.72},{"id":"3a34b6a8-5a99-48f2-a01e-f01c89b034bc","orcid":null,"display_name":"Chandan Yeshwanth","source":"manual","import_confidence":0.72},{"id":"a325f85a-4f0c-41b5-90a9-1a7672185a56","orcid":null,"display_name":"Matthias Nießner","source":"manual","import_confidence":0.72},{"id":"2036f1c8-58bd-4058-a65f-6a326afd1185","orcid":null,"display_name":"Yueh-Cheng Liu","source":"manual","import_confidence":0.72}]}}