{"work":{"id":"0a23d1b7-bd56-43cc-8a80-7c43ce994e1e","openalex_id":"https://openalex.org/W4312747027","doi":"10.1109/cvpr52688.2022","arxiv_id":"2688.2022","raw_key":null,"title":"In: 2022 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR)","authors":null,"authors_text":"Zhou, K","year":2022,"venue":null,"abstract":null,"external_url":"https://arxiv.org/abs/2688.2022","cited_by_count":725,"metadata_source":"arxiv_reference","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":null,"created_at":"2026-05-09T19:05:10.193085+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":false,"display_title":"A ConvNet for the 2020s","render_title":"A ConvNet for the 2020s"},"hub":{"state":{"work_id":"0a23d1b7-bd56-43cc-8a80-7c43ce994e1e","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":307,"external_cited_by_count":725,"distinct_field_count":32,"first_pith_cited_at":"2023-07-06T14:31:01+00:00","last_pith_cited_at":"2026-07-09T17:59:58+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-20T17:19:22.239315+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":43},{"context_role":"method","n":9},{"context_role":"baseline","n":6},{"context_role":"dataset","n":6}],"polarity_counts":[{"context_polarity":"background","n":42},{"context_polarity":"use_method","n":9},{"context_polarity":"baseline","n":6},{"context_polarity":"use_dataset","n":5},{"context_polarity":"support","n":1},{"context_polarity":"unclear","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Barron, Ben Mildenhall, Mehdi S","claims":[{"claim_text":"T= 16 consecutive frames to maintain the continuity of the frequency spectrum. Finally, these frames undergo a comprehensive video-level augmentation pipeline to simulate real-world variations. Baselines.For a comprehensive evaluation, we compare SpInShield with the following advanced and representative baselines, which are categorized into:Frame-level methods: SLADD [ 2], SBI [35], UCF [46], IID [ 16], LSDA [ 45], ProDet [ 3], and CDFA [ 23];Video-level methods: TALL [ 44], SLF [5], NACO [52], ","claim_type":"baseline","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"bones (also termed as blobs). Orange columns are skeleton- related, while Green columns are bone-related. Method Rigging Template Free Motion Adaptive Topology Correct Non-rigid Dynamic Robust [32] Skeleton ✓ ✓ ✗ ✗ TAVA [37] Skeleton ✗ ✗ ✓ ✗ BANMo [72] Bones ✓ ✓ ✗ ✓ RAC [73] Skeleton ✗ ✓ ✓ ✗ CAMM [30] Skeleton ✓ ✗ ✗ ✗ AP-NeRF [63] Skeleton ✓ ✗ ✗ ✗ WIM [49] Skeleton ✓ ✓ ✗ ✗ SC-GS [24] Bones ✓ ✗ ✗ ✓ DressRecon [58] SMPL,Bones✗ ✗ ✓ ✓ RigGS [75] Skeleton ✓ ✓ ✗ ✗ CANOR [21] Bones ✓ ✗ ✗ ✓ Ours Skelebo","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv, 2023. doi:10.48550/arXiv.2206.04615. [6] Sun, P., H. Kretzschmar, X. Dotiwalla, et al. Scalability in Perception for Autonomous Driving: Waymo Open Dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 2443-2451. doi:10.1109/CVPR42600.2020.00252. [7] Yu, H., Y. Luo, M. Shu, et al. DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Coopera","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"team communication detection, we train the spatial-temporal gaze model for 40 epochs with a batch size (B) of 1,024, and use the AdamW optimizer with an initial learning rate of 1e-4. 6 Table 1: Comparison with the state-of-the-art temporal action detection approaches. We report AP at different tIoU. Method StOP? 0.1 0.2 0.3 0.4 0.5 Avg. ActionFormer [41] 12.63 12.63 12.63 5.79 1.22 8.98 TriDet [34] 23.46 13.37 13.35 13.27 1.87 13.06 TemporalMaxer [35] 16.04 14.36 13.72 13.53 12.38 14.01 Team-OR","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"mechanisms at varying depths through explicit visualization of all model inputs, intermediate steps, and outputs. The Transformer Explainer [4] extends this approach with in-browser inference, ex- posing attention and token probabilities. Other tools focus on in- specting specific components of trained models, such as attention patterns [28, 33] and multimodal representations [1]. These tools focus on exploring inference-time behavior, with limited capacity to convey the training dynamics that a","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"task into several smaller and simpler sub-tasks, solves each sub- task individually, and then combines their solutions to reach the final solution for the original problem. Utilizing divide-to-conquer strategy makes great progress on various vision tasks, including visual generation [ 21, 46, 48], scene understanding [ 30, 38, 52], visual reasoning [5, 42, 50], and so on. Most works typically focus on exploring task-related representations from local regions and then generalize them to the entir","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Barron, Ben Mildenhall, Mehdi S because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (32 contexts).","role_counts":[{"n":32,"context_role":"background"},{"n":5,"context_role":"baseline"},{"n":4,"context_role":"method"},{"n":3,"context_role":"dataset"}]},"error":null,"updated_at":"2026-05-18T14:40:58.801533+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"c40c93be-497d-4f00-a601-d5081475c007","orcid":null,"display_name":"High-Resolution Image Synthesis with Latent Diffusion Models"}]},"error":null,"updated_at":"2026-05-18T14:40:58.795529+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T07:27:51.887766+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"In: 2023 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR)","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":19},{"title":"& Vondrick, C","work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","shared_citers":18},{"title":"IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =","work_id":"9da51225-b7bd-4032-b7db-ca577971dafe","shared_citers":16},{"title":"Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":14},{"title":"In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","work_id":"7083a41e-5666-435b-ab26-c753f6490b9a","shared_citers":14},{"title":"Editing conditional radiance fields","work_id":"3820f598-11b0-45c3-8c99-0079181ac0a7","shared_citers":9},{"title":"URLhttp://dx.doi.org/10.1109/CVPR.2016.90","work_id":"b353bda2-591d-479a-9c8b-22dfcba12431","shared_citers":9},{"title":"MambaVision: A hybrid Mamba- Transformer vision backbone","work_id":"d0e5199d-8907-47b1-905a-07ab8b623a4c","shared_citers":7},{"title":"URL https://doi.org/10.48550/arXiv","work_id":"5c2060c6-427c-4321-be22-49ccae439d80","shared_citers":7},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":6},{"title":"Tomasi and R","work_id":"135418b1-cafd-49fd-803d-1ca6433d4b1b","shared_citers":6},{"title":"doi: 10.18653/v1/N19-1423","work_id":"3e3c8ac8-b858-4b22-af32-393d98c883e0","shared_citers":5},{"title":"Foundation X: Integrating classification, localization, and segmentation through lock-release pretraining strategy for chest x-ray analysis","work_id":"498f8786-e863-4217-b210-bf4fe976c779","shared_citers":5},{"title":"Imagenet: A large-scale hierarchical image database","work_id":"effdb28b-742e-4840-b3ca-d89502a6cd4d","shared_citers":5},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":4},{"title":"Lawrence Zitnick","work_id":"9a41dfbd-d7c8-4b32-943f-59c08fdd6db7","shared_citers":4},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":4},{"title":"arXiv preprint arXiv:1903.11027 (May 2020)","work_id":"a687c611-43a9-4af4-bf00-36a2d9fa85a8","shared_citers":3},{"title":"Bovik, H.R","work_id":"a9ba8a9e-c00a-45ff-9d73-a0ae0919d283","shared_citers":3},{"title":"Clarke.Large Lan- guage Models for Software Engineering: A Systematic Mapping Study, page 64–79","work_id":"4a7d968d-ee43-47f9-9cca-832cf129af59","shared_citers":3},{"title":"Distilling the Knowledge in a Neural Network","work_id":"d927ab1f-17b8-4002-9d09-c3d55764fbad","shared_citers":3},{"title":"InInternational Joint Conference on Neural Networks, IJCNN 2024, Yokohama, Japan, June 30 - July 5","work_id":"813d43bf-9bad-4c36-bb4c-a82779adc003","shared_citers":3},{"title":"Recognizing indoor scenes","work_id":"45b0bfd8-65dc-4252-b2ab-2f6b411d04d0","shared_citers":3},{"title":"Toward Better Accuracy- EfficiencyTrade-Offs:DivideandCo-Training,in:Proceedingsofthe IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp","work_id":"3af1f7a7-0b03-44a7-8f18-d71b1a1d04b6","shared_citers":3}],"time_series":[{"n":1,"year":2024},{"n":1,"year":2025},{"n":75,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T07:27:47.506461+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T07:28:00.105068+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Barron, Ben Mildenhall, Mehdi S","claims":[{"claim_text":"T= 16 consecutive frames to maintain the continuity of the frequency spectrum. Finally, these frames undergo a comprehensive video-level augmentation pipeline to simulate real-world variations. Baselines.For a comprehensive evaluation, we compare SpInShield with the following advanced and representative baselines, which are categorized into:Frame-level methods: SLADD [ 2], SBI [35], UCF [46], IID [ 16], LSDA [ 45], ProDet [ 3], and CDFA [ 23];Video-level methods: TALL [ 44], SLF [5], NACO [52], ","claim_type":"baseline","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"bones (also termed as blobs). Orange columns are skeleton- related, while Green columns are bone-related. Method Rigging Template Free Motion Adaptive Topology Correct Non-rigid Dynamic Robust [32] Skeleton ✓ ✓ ✗ ✗ TAVA [37] Skeleton ✗ ✗ ✓ ✗ BANMo [72] Bones ✓ ✓ ✗ ✓ RAC [73] Skeleton ✗ ✓ ✓ ✗ CAMM [30] Skeleton ✓ ✗ ✗ ✗ AP-NeRF [63] Skeleton ✓ ✗ ✗ ✗ WIM [49] Skeleton ✓ ✓ ✗ ✗ SC-GS [24] Bones ✓ ✗ ✗ ✓ DressRecon [58] SMPL,Bones✗ ✗ ✓ ✓ RigGS [75] Skeleton ✓ ✓ ✗ ✗ CANOR [21] Bones ✓ ✗ ✗ ✓ Ours Skelebo","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv, 2023. doi:10.48550/arXiv.2206.04615. [6] Sun, P., H. Kretzschmar, X. Dotiwalla, et al. Scalability in Perception for Autonomous Driving: Waymo Open Dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 2443-2451. doi:10.1109/CVPR42600.2020.00252. [7] Yu, H., Y. Luo, M. Shu, et al. DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Coopera","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"team communication detection, we train the spatial-temporal gaze model for 40 epochs with a batch size (B) of 1,024, and use the AdamW optimizer with an initial learning rate of 1e-4. 6 Table 1: Comparison with the state-of-the-art temporal action detection approaches. We report AP at different tIoU. Method StOP? 0.1 0.2 0.3 0.4 0.5 Avg. ActionFormer [41] 12.63 12.63 12.63 5.79 1.22 8.98 TriDet [34] 23.46 13.37 13.35 13.27 1.87 13.06 TemporalMaxer [35] 16.04 14.36 13.72 13.53 12.38 14.01 Team-OR","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"mechanisms at varying depths through explicit visualization of all model inputs, intermediate steps, and outputs. The Transformer Explainer [4] extends this approach with in-browser inference, ex- posing attention and token probabilities. Other tools focus on in- specting specific components of trained models, such as attention patterns [28, 33] and multimodal representations [1]. These tools focus on exploring inference-time behavior, with limited capacity to convey the training dynamics that a","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"task into several smaller and simpler sub-tasks, solves each sub- task individually, and then combines their solutions to reach the final solution for the original problem. Utilizing divide-to-conquer strategy makes great progress on various vision tasks, including visual generation [ 21, 46, 48], scene understanding [ 30, 38, 52], visual reasoning [5, 42, 50], and so on. Most works typically focus on exploring task-related representations from local regions and then generalize them to the entir","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Barron, Ben Mildenhall, Mehdi S because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (32 contexts).","role_counts":[{"n":32,"context_role":"background"},{"n":5,"context_role":"baseline"},{"n":4,"context_role":"method"},{"n":3,"context_role":"dataset"}]},"error":null,"updated_at":"2026-05-18T14:40:58.805331+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Masked autoencoders are scalable vision learners","claims":[],"why_cited":"Pith tracks Masked autoencoders are scalable vision learners because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T07:27:47.513534+00:00"}},"summary":{"title":"Masked autoencoders are scalable vision learners","claims":[],"why_cited":"Pith tracks Masked autoencoders are scalable vision learners because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"In: 2023 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR)","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":19},{"title":"& Vondrick, C","work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","shared_citers":18},{"title":"IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =","work_id":"9da51225-b7bd-4032-b7db-ca577971dafe","shared_citers":16},{"title":"Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":14},{"title":"In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","work_id":"7083a41e-5666-435b-ab26-c753f6490b9a","shared_citers":14},{"title":"Editing conditional radiance fields","work_id":"3820f598-11b0-45c3-8c99-0079181ac0a7","shared_citers":9},{"title":"URLhttp://dx.doi.org/10.1109/CVPR.2016.90","work_id":"b353bda2-591d-479a-9c8b-22dfcba12431","shared_citers":9},{"title":"MambaVision: A hybrid Mamba- Transformer vision backbone","work_id":"d0e5199d-8907-47b1-905a-07ab8b623a4c","shared_citers":7},{"title":"URL https://doi.org/10.48550/arXiv","work_id":"5c2060c6-427c-4321-be22-49ccae439d80","shared_citers":7},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":6},{"title":"Tomasi and R","work_id":"135418b1-cafd-49fd-803d-1ca6433d4b1b","shared_citers":6},{"title":"doi: 10.18653/v1/N19-1423","work_id":"3e3c8ac8-b858-4b22-af32-393d98c883e0","shared_citers":5},{"title":"Foundation X: Integrating classification, localization, and segmentation through lock-release pretraining strategy for chest x-ray analysis","work_id":"498f8786-e863-4217-b210-bf4fe976c779","shared_citers":5},{"title":"Imagenet: A large-scale hierarchical image database","work_id":"effdb28b-742e-4840-b3ca-d89502a6cd4d","shared_citers":5},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":4},{"title":"Lawrence Zitnick","work_id":"9a41dfbd-d7c8-4b32-943f-59c08fdd6db7","shared_citers":4},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":4},{"title":"arXiv preprint arXiv:1903.11027 (May 2020)","work_id":"a687c611-43a9-4af4-bf00-36a2d9fa85a8","shared_citers":3},{"title":"Bovik, H.R","work_id":"a9ba8a9e-c00a-45ff-9d73-a0ae0919d283","shared_citers":3},{"title":"Clarke.Large Lan- guage Models for Software Engineering: A Systematic Mapping Study, page 64–79","work_id":"4a7d968d-ee43-47f9-9cca-832cf129af59","shared_citers":3},{"title":"Distilling the Knowledge in a Neural Network","work_id":"d927ab1f-17b8-4002-9d09-c3d55764fbad","shared_citers":3},{"title":"InInternational Joint Conference on Neural Networks, IJCNN 2024, Yokohama, Japan, June 30 - July 5","work_id":"813d43bf-9bad-4c36-bb4c-a82779adc003","shared_citers":3},{"title":"Recognizing indoor scenes","work_id":"45b0bfd8-65dc-4252-b2ab-2f6b411d04d0","shared_citers":3},{"title":"Toward Better Accuracy- EfficiencyTrade-Offs:DivideandCo-Training,in:Proceedingsofthe IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp","work_id":"3af1f7a7-0b03-44a7-8f18-d71b1a1d04b6","shared_citers":3}],"time_series":[{"n":1,"year":2024},{"n":1,"year":2025},{"n":75,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"c40c93be-497d-4f00-a601-d5081475c007","orcid":null,"display_name":"High-Resolution Image Synthesis with Latent Diffusion Models","source":"manual","import_confidence":0.72}]}}