{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2022:KOEQNG3KXBW5TI4FHWDZUY3UFO","short_pith_number":"pith:KOEQNG3K","schema_version":"1.0","canonical_sha256":"5389069b6ab86dd9a3853d879a63742b9a9a2f70dfc54a1974316d36a70b3688","source":{"kind":"arxiv","id":"2201.03545","version":2},"attestation_state":"computed","paper":{"title":"A ConvNet for the 2020s","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Chao-Yuan Wu, Christoph Feichtenhofer, Hanzi Mao, Saining Xie, Trevor Darrell, Zhuang Liu","submitted_at":"2022-01-10T18:59:10Z","abstract_excerpt":"The \"Roaring 20s\" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model. A vanilla ViT, on the other hand, faces difficulties when applied to general computer vision tasks such as object detection and semantic segmentation. It is the hierarchical Transformers (e.g., Swin Transformers) that reintroduced several ConvNet priors, making Transformers practically viable as a generic vision backbone and demonstrating remarkable performance on a wide variety of vision tasks. However, the eff"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2201.03545","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CV","submitted_at":"2022-01-10T18:59:10Z","cross_cats_sorted":[],"title_canon_sha256":"a310188c42445a0ab8da2f384b3257e37f5850432701bedf623f06261336a364","abstract_canon_sha256":"76b9481f5a8c8a1d44acc522584bc2331fe75035299e30cc85449a51b8471848"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T04:01:33.895441Z","signature_b64":"koBef04XAMmfMKdl7kMKky67Jb8y7q4gqhT32AOrbWCPQSK1nl3MeTNdv9ZV/DkLH/QBH3/rViuOTr/wwJpwAQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"5389069b6ab86dd9a3853d879a63742b9a9a2f70dfc54a1974316d36a70b3688","last_reissued_at":"2026-07-05T04:01:33.894967Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T04:01:33.894967Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"A ConvNet for the 2020s","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Chao-Yuan Wu, Christoph Feichtenhofer, Hanzi Mao, Saining Xie, Trevor Darrell, Zhuang Liu","submitted_at":"2022-01-10T18:59:10Z","abstract_excerpt":"The \"Roaring 20s\" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model. A vanilla ViT, on the other hand, faces difficulties when applied to general computer vision tasks such as object detection and semantic segmentation. It is the hierarchical Transformers (e.g., Swin Transformers) that reintroduced several ConvNet priors, making Transformers practically viable as a generic vision backbone and demonstrating remarkable performance on a wide variety of vision tasks. However, the eff"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2201.03545","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2201.03545/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2201.03545","created_at":"2026-07-05T04:01:33.895017+00:00"},{"alias_kind":"arxiv_version","alias_value":"2201.03545v2","created_at":"2026-07-05T04:01:33.895017+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2201.03545","created_at":"2026-07-05T04:01:33.895017+00:00"},{"alias_kind":"pith_short_12","alias_value":"KOEQNG3KXBW5","created_at":"2026-07-05T04:01:33.895017+00:00"},{"alias_kind":"pith_short_16","alias_value":"KOEQNG3KXBW5TI4F","created_at":"2026-07-05T04:01:33.895017+00:00"},{"alias_kind":"pith_short_8","alias_value":"KOEQNG3K","created_at":"2026-07-05T04:01:33.895017+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":22,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.26894","citing_title":"Modeling Local, Global, and Cross-Modal Context in Multimodal 3D MRI","ref_index":51,"is_internal_anchor":false},{"citing_arxiv_id":"2606.20561","citing_title":"TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living","ref_index":84,"is_internal_anchor":false},{"citing_arxiv_id":"2607.00228","citing_title":"Leveraging Multimodality for Real-Time Classification of Transients and Variables found by the Zwicky Transient Facility","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2606.03512","citing_title":"SPADE: Sketch-guided Path Planning Augmented with Diffusion Experts","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2606.01819","citing_title":"Hist2Style: Histogram-Guided Stylization with Bilateral Grids","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29723","citing_title":"ScaleAware-JEPA: Latent Representation for Discovery in Multiscale Physical Fields","ref_index":62,"is_internal_anchor":false},{"citing_arxiv_id":"2606.30108","citing_title":"LETT-NeXt: A Lightweight RECIST-Guided Model for 3D CT Lesion Segmentation","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2606.11315","citing_title":"Spiral arms across cosmic time: JWST measurements of the pitch angles of spiral galaxies at $z<3.5$","ref_index":55,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20250","citing_title":"Physics-informed convolutional neural networks for fluid flow through porous media","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06944","citing_title":"AIMIP Phase 1: systematic evaluations of AI weather and climate models","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2602.09524","citing_title":"HLGFA: High-Low Resolution Guided Feature Alignment for Unsupervised Anomaly Detection","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2602.20028","citing_title":"Descriptor: Parasitoid Wasps and Associated Hymenoptera Dataset (DAPWH)","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2604.10634","citing_title":"NTIRE 2026 The Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13555","citing_title":"Generating synthetic computed tomography for radiotherapy: SynthRAD2025 challenge report","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2312.14238","citing_title":"InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks","ref_index":96,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11107","citing_title":"Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2604.26324","citing_title":"Federated Medical Image Classification under Class and Domain Imbalance exploiting Synthetic Sample Generation","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2604.21780","citing_title":"Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative Models","ref_index":41,"is_internal_anchor":false},{"citing_arxiv_id":"2604.10634","citing_title":"NTIRE 2026 The Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2604.18067","citing_title":"Towards Real-Time ECG and EMG Modeling on $\\mu$NPUs","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05549","citing_title":"A Novel Graph-Regulated Disentangling Mamba Model with Sparse Tokens for Enhanced Tree Species Classification from MODIS Time Series","ref_index":78,"is_internal_anchor":false},{"citing_arxiv_id":"2604.24913","citing_title":"Generative diffusion models for spatiotemporal influenza forecasting","ref_index":12,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/KOEQNG3KXBW5TI4FHWDZUY3UFO","json":"https://pith.science/pith/KOEQNG3KXBW5TI4FHWDZUY3UFO.json","graph_json":"https://pith.science/api/pith-number/KOEQNG3KXBW5TI4FHWDZUY3UFO/graph.json","events_json":"https://pith.science/api/pith-number/KOEQNG3KXBW5TI4FHWDZUY3UFO/events.json","paper":"https://pith.science/paper/KOEQNG3K"},"agent_actions":{"view_html":"https://pith.science/pith/KOEQNG3KXBW5TI4FHWDZUY3UFO","download_json":"https://pith.science/pith/KOEQNG3KXBW5TI4FHWDZUY3UFO.json","view_paper":"https://pith.science/paper/KOEQNG3K","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2201.03545&json=true","fetch_graph":"https://pith.science/api/pith-number/KOEQNG3KXBW5TI4FHWDZUY3UFO/graph.json","fetch_events":"https://pith.science/api/pith-number/KOEQNG3KXBW5TI4FHWDZUY3UFO/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/KOEQNG3KXBW5TI4FHWDZUY3UFO/action/timestamp_anchor","attest_storage":"https://pith.science/pith/KOEQNG3KXBW5TI4FHWDZUY3UFO/action/storage_attestation","attest_author":"https://pith.science/pith/KOEQNG3KXBW5TI4FHWDZUY3UFO/action/author_attestation","sign_citation":"https://pith.science/pith/KOEQNG3KXBW5TI4FHWDZUY3UFO/action/citation_signature","submit_replication":"https://pith.science/pith/KOEQNG3KXBW5TI4FHWDZUY3UFO/action/replication_record"}},"created_at":"2026-07-05T04:01:33.895017+00:00","updated_at":"2026-07-05T04:01:33.895017+00:00"}