{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2026:J2SUWQ2VUT6GSTP7Z7NLPYX6G2","short_pith_number":"pith:J2SUWQ2V","schema_version":"1.0","canonical_sha256":"4ea54b4355a4fc694dffcfdab7e2fe36ade01ef7c164645fae48d7b0fb70a20f","source":{"kind":"arxiv","id":"2605.14333","version":1},"attestation_state":"computed","paper":{"title":"InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"InsightTok uses localized content-aware perceptual losses to improve text and face fidelity in discrete image tokenizers.","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Dong Chen, Fangyun Wei, Gao Huang, Jiayi Guo, Ji Li, Jinjing Zhao, Lei Shi, Li Chen, Tianyu He, Yang Yue, Yue Dong, Zanlin Ni, Zeyu Liu","submitted_at":"2026-05-14T03:57:25Z","abstract_excerpt":"Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregressive generators built on discrete tokenization. A central bottleneck is the tokenizer: aggressive downsampling and quantization often discard the fine-grained structures needed to preserve readable glyphs and distinctive facial features. We attribute this gap to standard discrete-tokenizer objectives being weakly aligned with text legibility and facial fidelity, as these objectives typically optimize generic reconstruction while compressing d"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"2605.14333","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CV","submitted_at":"2026-05-14T03:57:25Z","cross_cats_sorted":[],"title_canon_sha256":"bc892d40b41de6034c550b6e8f46c1fc03a36e85ee9702295e73845da32ae016","abstract_canon_sha256":"4ef55016f8f82c403eac1edf7860c2366b261416c6e96612c37953b8096c777d"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-05-17T23:39:08.269203Z","signature_b64":"3ZxPDxpQtcQjQrPmdQB83jpy8SmeX1ohfmyxvA5reTvnOapNrc/tQINJmDJlWPl4n96h0UJRnBH7Q1raJNGeDw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"4ea54b4355a4fc694dffcfdab7e2fe36ade01ef7c164645fae48d7b0fb70a20f","last_reissued_at":"2026-05-17T23:39:08.268499Z","signature_status":"signed_v1","first_computed_at":"2026-05-17T23:39:08.268499Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"InsightTok uses localized content-aware perceptual losses to improve text and face fidelity in discrete image tokenizers.","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Dong Chen, Fangyun Wei, Gao Huang, Jiayi Guo, Ji Li, Jinjing Zhao, Lei Shi, Li Chen, Tianyu He, Yang Yue, Yue Dong, Zanlin Ni, Zeyu Liu","submitted_at":"2026-05-14T03:57:25Z","abstract_excerpt":"Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregressive generators built on discrete tokenization. A central bottleneck is the tokenizer: aggressive downsampling and quantization often discard the fine-grained structures needed to preserve readable glyphs and distinctive facial features. We attribute this gap to standard discrete-tokenizer objectives being weakly aligned with text legibility and facial fidelity, as these objectives typically optimize generic reconstruction while compressing d"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"With a compact 16k codebook and a 16x downsampling rate, InsightTok significantly outperforms prior tokenizers in text and face reconstruction without compromising general reconstruction quality. These gains consistently transfer to autoregressive image generation in InsightAR, producing images with clearer text and more faithful facial details.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That localized content-aware perceptual losses will reliably capture fine-grained text legibility and facial fidelity across diverse images without introducing new artifacts or requiring extensive hyperparameter tuning for each domain.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"InsightTok improves text and face fidelity in discrete image tokenization via content-aware perceptual losses, with gains transferring to autoregressive generation.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"InsightTok uses localized content-aware perceptual losses to improve text and face fidelity in discrete image tokenizers.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"9e53ecc16e9da1302a70903439b00f381b09cd976d424b0b8d43e1f40d86e818"},"source":{"id":"2605.14333","kind":"arxiv","version":1},"verdict":{"id":"b3f63a77-4e97-4e17-93bb-80b81b739377","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-15T02:06:31.462898Z","strongest_claim":"With a compact 16k codebook and a 16x downsampling rate, InsightTok significantly outperforms prior tokenizers in text and face reconstruction without compromising general reconstruction quality. These gains consistently transfer to autoregressive image generation in InsightAR, producing images with clearer text and more faithful facial details.","one_line_summary":"InsightTok improves text and face fidelity in discrete image tokenization via content-aware perceptual losses, with gains transferring to autoregressive generation.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That localized content-aware perceptual losses will reliably capture fine-grained text legibility and facial fidelity across diverse images without introducing new artifacts or requiring extensive hyperparameter tuning for each domain.","pith_extraction_headline":"InsightTok uses localized content-aware perceptual losses to improve text and face fidelity in discrete image tokenizers."},"references":{"count":60,"sample":[{"doi":"","year":2014,"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","ref_index":1,"cited_arxiv_id":"1412.6980","is_internal_anchor":true},{"doi":"","year":2025,"title":"Cosmos World Foundation Model Platform for Physical AI","work_id":"a2dba24c-318d-476a-8b21-4289c265810c","ref_index":2,"cited_arxiv_id":"2501.03575","is_internal_anchor":true},{"doi":"","year":2025,"title":"Flextok: Resampling images into 1d token sequences of flexible length","work_id":"004bfe14-b2bd-4b36-8676-294f29202bcb","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2022,"title":"Scene text recognition with permuted autoregressive sequence models","work_id":"5041cb99-6c4c-4496-82c6-377ecfae1452","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2009,"title":"Faces and text attract gaze independent of the task: Experimental data and computer model.Journal of vision, 9(12):10–10, 2009","work_id":"682615f7-c5c6-4017-b092-7d54e60dc079","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":60,"snapshot_sha256":"49c8b3fc8aad26cd5ddaf3c70bcd9d63fb69b84c7ba2de4c48c1bfeaf364e34b","internal_anchors":14},"formal_canon":{"evidence_count":2,"snapshot_sha256":"fdc73ac68b20068506aae2ee16359469135b10f5e8f8f87f78d50506aeef68bf"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2605.14333","created_at":"2026-05-17T23:39:08.268618+00:00"},{"alias_kind":"arxiv_version","alias_value":"2605.14333v1","created_at":"2026-05-17T23:39:08.268618+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2605.14333","created_at":"2026-05-17T23:39:08.268618+00:00"},{"alias_kind":"pith_short_12","alias_value":"J2SUWQ2VUT6G","created_at":"2026-05-18T12:33:37.589309+00:00"},{"alias_kind":"pith_short_16","alias_value":"J2SUWQ2VUT6GSTP7","created_at":"2026-05-18T12:33:37.589309+00:00"},{"alias_kind":"pith_short_8","alias_value":"J2SUWQ2V","created_at":"2026-05-18T12:33:37.589309+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":0,"internal_anchor_count":0,"sample":[]},"formal_canon":{"evidence_count":2,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/J2SUWQ2VUT6GSTP7Z7NLPYX6G2","json":"https://pith.science/pith/J2SUWQ2VUT6GSTP7Z7NLPYX6G2.json","graph_json":"https://pith.science/api/pith-number/J2SUWQ2VUT6GSTP7Z7NLPYX6G2/graph.json","events_json":"https://pith.science/api/pith-number/J2SUWQ2VUT6GSTP7Z7NLPYX6G2/events.json","paper":"https://pith.science/paper/J2SUWQ2V"},"agent_actions":{"view_html":"https://pith.science/pith/J2SUWQ2VUT6GSTP7Z7NLPYX6G2","download_json":"https://pith.science/pith/J2SUWQ2VUT6GSTP7Z7NLPYX6G2.json","view_paper":"https://pith.science/paper/J2SUWQ2V","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2605.14333&json=true","fetch_graph":"https://pith.science/api/pith-number/J2SUWQ2VUT6GSTP7Z7NLPYX6G2/graph.json","fetch_events":"https://pith.science/api/pith-number/J2SUWQ2VUT6GSTP7Z7NLPYX6G2/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/J2SUWQ2VUT6GSTP7Z7NLPYX6G2/action/timestamp_anchor","attest_storage":"https://pith.science/pith/J2SUWQ2VUT6GSTP7Z7NLPYX6G2/action/storage_attestation","attest_author":"https://pith.science/pith/J2SUWQ2VUT6GSTP7Z7NLPYX6G2/action/author_attestation","sign_citation":"https://pith.science/pith/J2SUWQ2VUT6GSTP7Z7NLPYX6G2/action/citation_signature","submit_replication":"https://pith.science/pith/J2SUWQ2VUT6GSTP7Z7NLPYX6G2/action/replication_record"}},"created_at":"2026-05-17T23:39:08.268618+00:00","updated_at":"2026-05-17T23:39:08.268618+00:00"}