{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2026:7VITVVLL3U272IR42MMYWQ4ZLK","short_pith_number":"pith:7VITVVLL","schema_version":"1.0","canonical_sha256":"fd513ad56bdd35fd223cd3198b43995a84f4667c52311226ad625c8e25832ead","source":{"kind":"arxiv","id":"2604.20329","version":3},"attestation_state":"computed","paper":{"title":"Image Generators are Generalist Vision Learners","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Image generation pretraining builds general visual representations that reach SOTA on perception tasks when outputs are cast as RGB images.","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Howard Zhou, Huizhong Chen, Jean-Baptiste Alayrac, Jonathan T. Barron, Kaiming He, Karen Truong, Kyle Genova, Mandy Guo, Nithish Kannen, Oliver Wang, Paul Voigtlaender, Radu Soricut, Saining Xie, Shangbang Long, Sherry Ben, Shuyang Sun, Songyou Peng, Suhas Yogin, Thomas Funkhouser, Valentin Gabeur, Wenlei Zhou, Yanan Bao, Yandong Li, Yiming Gu, Zhicheng Wang","submitted_at":"2026-04-22T08:23:48Z","abstract_excerpt":"Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language understanding and reasoning from generative pretraining. While it has long been conjectured that the ability to create visual content implies an ability to understand it, there has been limited evidence that generative vision models have developed strong understanding capabilities. In this work, we demonstrate that image generation training serves a role similar to LLM pretraining, and lets models learn powerful and gener"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":true},"canonical_record":{"source":{"id":"2604.20329","kind":"arxiv","version":3},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2026-04-22T08:23:48Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"6e0eda6351a0995d80d3309c161471b9ef58e5ca7e667ed24766f606b6b05b72","abstract_canon_sha256":"7811065d4485055e432efffe682bf6adfc096c79f04da95cd16c4bf70f5273f2"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-06-05T00:13:46.610419Z","signature_b64":"OgjAfxWhHUzutkIh9Ak5G65VzjEJB8TqgeYTf4mTFUHVFFvTovMhaiKeAhxO7om2c73e9Mt+Q59HKFtQf0hVAQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"fd513ad56bdd35fd223cd3198b43995a84f4667c52311226ad625c8e25832ead","last_reissued_at":"2026-06-05T00:13:46.609853Z","signature_status":"signed_v1","first_computed_at":"2026-06-05T00:13:46.609853Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Image Generators are Generalist Vision Learners","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Image generation pretraining builds general visual representations that reach SOTA on perception tasks when outputs are cast as RGB images.","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Howard Zhou, Huizhong Chen, Jean-Baptiste Alayrac, Jonathan T. Barron, Kaiming He, Karen Truong, Kyle Genova, Mandy Guo, Nithish Kannen, Oliver Wang, Paul Voigtlaender, Radu Soricut, Saining Xie, Shangbang Long, Sherry Ben, Shuyang Sun, Songyou Peng, Suhas Yogin, Thomas Funkhouser, Valentin Gabeur, Wenlei Zhou, Yanan Bao, Yandong Li, Yiming Gu, Zhicheng Wang","submitted_at":"2026-04-22T08:23:48Z","abstract_excerpt":"Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language understanding and reasoning from generative pretraining. While it has long been conjectured that the ability to create visual content implies an ability to understand it, there has been limited evidence that generative vision models have developed strong understanding capabilities. In this work, we demonstrate that image generation training serves a role similar to LLM pretraining, and lets models learn powerful and gener"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"image generation training serves a role similar to LLM pretraining, and lets models learn powerful and general visual representations that enable SOTA performance on various vision tasks.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the reported SOTA results stem primarily from the generative pretraining rather than from the specific data mixture, evaluation protocol, or implicit leakage in the instruction-tuning stage.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Image generation pretraining builds generalist vision models that reach SOTA on 2D and 3D perception tasks by reframing them as RGB image outputs.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Image generation pretraining builds general visual representations that reach SOTA on perception tasks when outputs are cast as RGB images.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"3e93d2fa214876262b53ac05a60292c562b1302a20e8ffb158dd0b26bec91b71"},"source":{"id":"2604.20329","kind":"arxiv","version":3},"verdict":{"id":"e597e217-0af2-4723-ba4e-364b779295b1","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-15T07:37:38.564195Z","strongest_claim":"image generation training serves a role similar to LLM pretraining, and lets models learn powerful and general visual representations that enable SOTA performance on various vision tasks.","one_line_summary":"Image generation pretraining builds generalist vision models that reach SOTA on 2D and 3D perception tasks by reframing them as RGB image outputs.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the reported SOTA results stem primarily from the generative pretraining rather than from the specific data mixture, evaluation protocol, or implicit leakage in the instruction-tuning stage.","pith_extraction_headline":"Image generation pretraining builds general visual representations that reach SOTA on perception tasks when outputs are cast as RGB images."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2604.20329/integrity.json","findings":[],"available":true,"detectors_run":[{"name":"ai_meta_artifact","ran_at":"2026-05-21T14:43:15.756849Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"doi_compliance","ran_at":"2026-05-20T02:01:52.370703Z","status":"completed","version":"1.0.0","findings_count":0}],"snapshot_sha256":"51f9e0dedbf7544d0810452be47540271c7ca3f96e86caedf42c9efa40c8daea"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":2,"snapshot_sha256":"4bcadf2220deedb38cb2d319cc05421104f8d4cc3183003ed16acc6c73f37fb7"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2604.20329","created_at":"2026-06-05T00:13:46.609930+00:00"},{"alias_kind":"arxiv_version","alias_value":"2604.20329v3","created_at":"2026-06-05T00:13:46.609930+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2604.20329","created_at":"2026-06-05T00:13:46.609930+00:00"},{"alias_kind":"pith_short_12","alias_value":"7VITVVLL3U27","created_at":"2026-06-05T00:13:46.609930+00:00"},{"alias_kind":"pith_short_16","alias_value":"7VITVVLL3U272IR4","created_at":"2026-06-05T00:13:46.609930+00:00"},{"alias_kind":"pith_short_8","alias_value":"7VITVVLL","created_at":"2026-06-05T00:13:46.609930+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":13,"internal_anchor_count":13,"sample":[{"citing_arxiv_id":"2607.06553","citing_title":"From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models","ref_index":9,"is_internal_anchor":true},{"citing_arxiv_id":"2607.06560","citing_title":"Vision as Unified Multimodal Generation","ref_index":43,"is_internal_anchor":true},{"citing_arxiv_id":"2606.22660","citing_title":"Prompting Diffusion Models for Zero-Shot Instance Segmentation","ref_index":29,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19531","citing_title":"ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?","ref_index":33,"is_internal_anchor":true},{"citing_arxiv_id":"2606.13676","citing_title":"Modality Forcing for Scalable Spatial Generation","ref_index":15,"is_internal_anchor":true},{"citing_arxiv_id":"2606.09813","citing_title":"iMaC: Translating Actions into Motion and Contact Images for Embodied World Models","ref_index":45,"is_internal_anchor":true},{"citing_arxiv_id":"2605.07496","citing_title":"PathPainter: Transferring the Generalization Ability of Image Generation Models to Embodied Navigation","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"2606.00188","citing_title":"PaintBench: Deterministic Evaluation of Precise Visual Editing","ref_index":18,"is_internal_anchor":true},{"citing_arxiv_id":"2605.10588","citing_title":"Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence","ref_index":14,"is_internal_anchor":true},{"citing_arxiv_id":"2604.24575","citing_title":"Diffusion Model as a Generalist Segmentation Learner","ref_index":27,"is_internal_anchor":true},{"citing_arxiv_id":"2605.05627","citing_title":"Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping","ref_index":83,"is_internal_anchor":true},{"citing_arxiv_id":"2605.04566","citing_title":"Open-Source Image Editing Models Are Zero-Shot Vision Learners","ref_index":18,"is_internal_anchor":true},{"citing_arxiv_id":"2605.07496","citing_title":"PathPainter: Transferring the Generalization Ability of Image Generation Models to Embodied Navigation","ref_index":16,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":2,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/7VITVVLL3U272IR42MMYWQ4ZLK","json":"https://pith.science/pith/7VITVVLL3U272IR42MMYWQ4ZLK.json","graph_json":"https://pith.science/api/pith-number/7VITVVLL3U272IR42MMYWQ4ZLK/graph.json","events_json":"https://pith.science/api/pith-number/7VITVVLL3U272IR42MMYWQ4ZLK/events.json","paper":"https://pith.science/paper/7VITVVLL"},"agent_actions":{"view_html":"https://pith.science/pith/7VITVVLL3U272IR42MMYWQ4ZLK","download_json":"https://pith.science/pith/7VITVVLL3U272IR42MMYWQ4ZLK.json","view_paper":"https://pith.science/paper/7VITVVLL","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2604.20329&json=true","fetch_graph":"https://pith.science/api/pith-number/7VITVVLL3U272IR42MMYWQ4ZLK/graph.json","fetch_events":"https://pith.science/api/pith-number/7VITVVLL3U272IR42MMYWQ4ZLK/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/7VITVVLL3U272IR42MMYWQ4ZLK/action/timestamp_anchor","attest_storage":"https://pith.science/pith/7VITVVLL3U272IR42MMYWQ4ZLK/action/storage_attestation","attest_author":"https://pith.science/pith/7VITVVLL3U272IR42MMYWQ4ZLK/action/author_attestation","sign_citation":"https://pith.science/pith/7VITVVLL3U272IR42MMYWQ4ZLK/action/citation_signature","submit_replication":"https://pith.science/pith/7VITVVLL3U272IR42MMYWQ4ZLK/action/replication_record"}},"created_at":"2026-06-05T00:13:46.609930+00:00","updated_at":"2026-06-05T00:13:46.609930+00:00"}