{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:FN6NOUFFTYV34HLUJVZFARW3YI","short_pith_number":"pith:FN6NOUFF","schema_version":"1.0","canonical_sha256":"2b7cd750a59e2bbe1d744d725046dbc218dca98dc0145b8399499691567ff514","source":{"kind":"arxiv","id":"2501.09755","version":1},"attestation_state":"computed","paper":{"title":"Learnings from Scaling Visual Tokenizers for Reconstruction and Generation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Ching-Yao Chung, David Yan, Jialiang Wang, Orr Zohar, Peter Vajda, Philippe Hansen-Estruch, Sriram Vishwanath, Tao Xu, Tingbo Hou, Xinlei Chen","submitted_at":"2025-01-16T18:59:04Z","abstract_excerpt":"Visual tokenization via auto-encoding empowers state-of-the-art image and video generative models by compressing pixels into a latent space. Although scaling Transformer-based generators has been central to recent advances, the tokenizer component itself is rarely scaled, leaving open questions about how auto-encoder design choices influence both its objective of reconstruction and downstream generative performance. Our work aims to conduct an exploration of scaling in auto-encoders to fill in this blank. To facilitate this exploration, we replace the typical convolutional backbone with an enh"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2501.09755","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CV","submitted_at":"2025-01-16T18:59:04Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"798139dcea2b5fe2fada30bd83c19c07dcb5ac8bc5ac8b7877a2c47263d1c8ac","abstract_canon_sha256":"b46877db1112f0010d3fad9e27da22fd6cb8420390a2e30c1c1324f7a107d669"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:01:58.116335Z","signature_b64":"keUi8HRGJ8rxl34qLcjd6pL8pB8vcEVLyUv3NELe8zOqTRM6RsYFDpmTHE9ebtSZp2TiAA0yFBZtSC/rH+NnBQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"2b7cd750a59e2bbe1d744d725046dbc218dca98dc0145b8399499691567ff514","last_reissued_at":"2026-07-05T10:01:58.115859Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:01:58.115859Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Learnings from Scaling Visual Tokenizers for Reconstruction and Generation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Ching-Yao Chung, David Yan, Jialiang Wang, Orr Zohar, Peter Vajda, Philippe Hansen-Estruch, Sriram Vishwanath, Tao Xu, Tingbo Hou, Xinlei Chen","submitted_at":"2025-01-16T18:59:04Z","abstract_excerpt":"Visual tokenization via auto-encoding empowers state-of-the-art image and video generative models by compressing pixels into a latent space. Although scaling Transformer-based generators has been central to recent advances, the tokenizer component itself is rarely scaled, leaving open questions about how auto-encoder design choices influence both its objective of reconstruction and downstream generative performance. Our work aims to conduct an exploration of scaling in auto-encoders to fill in this blank. To facilitate this exploration, we replace the typical convolutional backbone with an enh"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2501.09755","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2501.09755/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2501.09755","created_at":"2026-07-05T10:01:58.115911+00:00"},{"alias_kind":"arxiv_version","alias_value":"2501.09755v1","created_at":"2026-07-05T10:01:58.115911+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2501.09755","created_at":"2026-07-05T10:01:58.115911+00:00"},{"alias_kind":"pith_short_12","alias_value":"FN6NOUFFTYV3","created_at":"2026-07-05T10:01:58.115911+00:00"},{"alias_kind":"pith_short_16","alias_value":"FN6NOUFFTYV34HLU","created_at":"2026-07-05T10:01:58.115911+00:00"},{"alias_kind":"pith_short_8","alias_value":"FN6NOUFF","created_at":"2026-07-05T10:01:58.115911+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":10,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.05352","citing_title":"Multiplayer Interactive World Models with Representation Autoencoders","ref_index":101,"is_internal_anchor":true},{"citing_arxiv_id":"2606.05552","citing_title":"Balancing Image Compression and Generation with Bootstrapped Tokenization","ref_index":53,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18390","citing_title":"Vision Foundation Models as Generalist Tokenizers for Image Generation","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2509.18611","citing_title":"Flow marching for a generative PDE foundation model","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09981","citing_title":"Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06628","citing_title":"LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05331","citing_title":"ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2604.16479","citing_title":"Latent-Compressed Variational Autoencoder for Video Diffusion Models","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2604.07340","citing_title":"TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07915","citing_title":"What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion","ref_index":27,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/FN6NOUFFTYV34HLUJVZFARW3YI","json":"https://pith.science/pith/FN6NOUFFTYV34HLUJVZFARW3YI.json","graph_json":"https://pith.science/api/pith-number/FN6NOUFFTYV34HLUJVZFARW3YI/graph.json","events_json":"https://pith.science/api/pith-number/FN6NOUFFTYV34HLUJVZFARW3YI/events.json","paper":"https://pith.science/paper/FN6NOUFF"},"agent_actions":{"view_html":"https://pith.science/pith/FN6NOUFFTYV34HLUJVZFARW3YI","download_json":"https://pith.science/pith/FN6NOUFFTYV34HLUJVZFARW3YI.json","view_paper":"https://pith.science/paper/FN6NOUFF","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2501.09755&json=true","fetch_graph":"https://pith.science/api/pith-number/FN6NOUFFTYV34HLUJVZFARW3YI/graph.json","fetch_events":"https://pith.science/api/pith-number/FN6NOUFFTYV34HLUJVZFARW3YI/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/FN6NOUFFTYV34HLUJVZFARW3YI/action/timestamp_anchor","attest_storage":"https://pith.science/pith/FN6NOUFFTYV34HLUJVZFARW3YI/action/storage_attestation","attest_author":"https://pith.science/pith/FN6NOUFFTYV34HLUJVZFARW3YI/action/author_attestation","sign_citation":"https://pith.science/pith/FN6NOUFFTYV34HLUJVZFARW3YI/action/citation_signature","submit_replication":"https://pith.science/pith/FN6NOUFFTYV34HLUJVZFARW3YI/action/replication_record"}},"created_at":"2026-07-05T10:01:58.115911+00:00","updated_at":"2026-07-05T10:01:58.115911+00:00"}