{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:5YPLY3O4LHGJZW3YZ6BQJN3OPY","short_pith_number":"pith:5YPLY3O4","schema_version":"1.0","canonical_sha256":"ee1ebc6ddc59cc9cdb78cf8304b76e7e351471bd972f2a658f6e159b7f025e2d","source":{"kind":"arxiv","id":"2410.07073","version":2},"attestation_state":"computed","paper":{"title":"Pixtral 12B","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Pixtral-12B outperforms similar and larger open multimodal models by processing images at their native resolution and aspect ratio.","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Albert Q. Jiang, Alexandre Sablayrolles, Am\\'elie H\\'eliou, Andy Lo, Arthur Mensch, Baptiste Bout, Baptiste Rozi\\`ere, Baudouin De Monicault, Devendra Chaplot, Diego Las Casas, Diogo Costa, Emma Bou Hanna, Guillaume Lample, Jessica Chudnovsky, Joachim Studnia, Kartik Khandelwal, Lawrence Stewart, Louis Martin, Lucile Saulnier, Marie Pellat, Nikhil Raghuraman, Patrick von Platen, Paul Jacob, Pavankumar Muddireddy, Pierre Stock, Pravesh Agrawal, Romain Sauvestre, Roman Soletskyi, Sagar Vaze, Sandeep Subramanian, Saurabh Garg, Soham Ghosh, Sophia Yang, Szymon Antoniak, Teven Le Scao, Theophile Gervet, Thibaut Lavril, Thomas Wang, Timoth\\'ee Lacroix, Valera Nemychnikova, Wendy Shang, William Marshall","submitted_at":"2024-10-09T17:16:22Z","abstract_excerpt":"We introduce Pixtral-12B, a 12--billion-parameter multimodal language model. Pixtral-12B is trained to understand both natural images and documents, achieving leading performance on various multimodal benchmarks, surpassing a number of larger models. Unlike many open-source models, Pixtral is also a cutting-edge text model for its size, and does not compromise on natural language performance to excel in multimodal tasks. Pixtral uses a new vision encoder trained from scratch, which allows it to ingest images at their natural resolution and aspect ratio. This gives users flexibility on the numb"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"2410.07073","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CV","submitted_at":"2024-10-09T17:16:22Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"3f972522ad7f9c9b8d3bfef4a659fa191a919883ed15c79909f74eaec56c52da","abstract_canon_sha256":"5ccc3122e53fc8ad76ad168b34e3934d70248678eb2a61aad11041f1ce918560"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-05-17T23:39:19.791985Z","signature_b64":"eXY8B8urN3yvterDOgvm21rPUKXC0lLcEwkoP7c8fKyJQM6iaSfKrOwRBJHetqBh3wBan4cWEdtJV1viHKQFCw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"ee1ebc6ddc59cc9cdb78cf8304b76e7e351471bd972f2a658f6e159b7f025e2d","last_reissued_at":"2026-05-17T23:39:19.791252Z","signature_status":"signed_v1","first_computed_at":"2026-05-17T23:39:19.791252Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Pixtral 12B","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Pixtral-12B outperforms similar and larger open multimodal models by processing images at their native resolution and aspect ratio.","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Albert Q. Jiang, Alexandre Sablayrolles, Am\\'elie H\\'eliou, Andy Lo, Arthur Mensch, Baptiste Bout, Baptiste Rozi\\`ere, Baudouin De Monicault, Devendra Chaplot, Diego Las Casas, Diogo Costa, Emma Bou Hanna, Guillaume Lample, Jessica Chudnovsky, Joachim Studnia, Kartik Khandelwal, Lawrence Stewart, Louis Martin, Lucile Saulnier, Marie Pellat, Nikhil Raghuraman, Patrick von Platen, Paul Jacob, Pavankumar Muddireddy, Pierre Stock, Pravesh Agrawal, Romain Sauvestre, Roman Soletskyi, Sagar Vaze, Sandeep Subramanian, Saurabh Garg, Soham Ghosh, Sophia Yang, Szymon Antoniak, Teven Le Scao, Theophile Gervet, Thibaut Lavril, Thomas Wang, Timoth\\'ee Lacroix, Valera Nemychnikova, Wendy Shang, William Marshall","submitted_at":"2024-10-09T17:16:22Z","abstract_excerpt":"We introduce Pixtral-12B, a 12--billion-parameter multimodal language model. Pixtral-12B is trained to understand both natural images and documents, achieving leading performance on various multimodal benchmarks, surpassing a number of larger models. Unlike many open-source models, Pixtral is also a cutting-edge text model for its size, and does not compromise on natural language performance to excel in multimodal tasks. Pixtral uses a new vision encoder trained from scratch, which allows it to ingest images at their natural resolution and aspect ratio. This gives users flexibility on the numb"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Pixtral-12B substantially outperforms other open models of similar sizes (Llama-3.2 11B & Qwen-2-VL 7B). It also outperforms much larger open models like Llama-3.2 90B while being 7x smaller.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the reported benchmark scores reflect fair, standardized evaluation without undisclosed differences in training data scale, filtering, or test-set contamination.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Pixtral-12B is a 12B multimodal LLM with a custom vision encoder that ingests images at native resolution and aspect ratio, achieving leading benchmark results among open models while preserving text capabilities.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Pixtral-12B outperforms similar and larger open multimodal models by processing images at their native resolution and aspect ratio.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"8e93a681ad426f1045b8e052da915ba33f4826cbb09e7bd84c6fa845a296c126"},"source":{"id":"2410.07073","kind":"arxiv","version":2},"verdict":{"id":"f8247641-e4d8-48b5-94f4-ac0fc00009ba","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-14T23:48:54.453453Z","strongest_claim":"Pixtral-12B substantially outperforms other open models of similar sizes (Llama-3.2 11B & Qwen-2-VL 7B). It also outperforms much larger open models like Llama-3.2 90B while being 7x smaller.","one_line_summary":"Pixtral-12B is a 12B multimodal LLM with a custom vision encoder that ingests images at native resolution and aspect ratio, achieving leading benchmark results among open models while preserving text capabilities.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the reported benchmark scores reflect fair, standardized evaluation without undisclosed differences in training data scale, filtering, or test-set contamination.","pith_extraction_headline":"Pixtral-12B outperforms similar and larger open multimodal models by processing images at their native resolution and aspect ratio."},"references":{"count":26,"sample":[{"doi":"","year":2024,"title":"The Claude 3 Model Family: Opus, Sonnet, Haiku","work_id":"03133174-6436-4e6e-98c7-7c3fa2b54b81","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2023,"title":"Bavishi, R., Elsen, E., Hawthorne, C., Nye, M., Odena, A., Somani, A., and Ta¸ sırlar, S. (2023). Fuyu-8b: A multimodal architecture for ai agents","work_id":"db0ba251-6b07-4faf-8d07-2c042af3b492","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2024,"title":"Dehghani, M., Mustafa, B., Djolonga, J., Heek, J., Minderer, M., Caron, M., Steiner, A., Puigcerver, J., Geirhos, R., Alabdulmohsin, I. M., et al. (2024). Patch n’pack: Navit, a vision transformer for","work_id":"881f9b4a-5aca-40b8-8634-f70d08430bd1","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2024,"title":"Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models","work_id":"b6841af7-2b88-4597-9d0e-54e96aa7a368","ref_index":4,"cited_arxiv_id":"2409.17146","is_internal_anchor":true},{"doi":"","year":2020,"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","ref_index":5,"cited_arxiv_id":"2010.11929","is_internal_anchor":true}],"resolved_work":26,"snapshot_sha256":"0c16f85e45ba5a5aaa8e9cbf8e892372216c662ed3cb1a06fb53f5da1ee34201","internal_anchors":11},"formal_canon":{"evidence_count":2,"snapshot_sha256":"3097732f83e3d70ccfbf3f548f487b885f9c1b4b5840a90675633b689328260c"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2410.07073","created_at":"2026-05-17T23:39:19.791379+00:00"},{"alias_kind":"arxiv_version","alias_value":"2410.07073v2","created_at":"2026-05-17T23:39:19.791379+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2410.07073","created_at":"2026-05-17T23:39:19.791379+00:00"},{"alias_kind":"pith_short_12","alias_value":"5YPLY3O4LHGJ","created_at":"2026-05-18T12:33:37.589309+00:00"},{"alias_kind":"pith_short_16","alias_value":"5YPLY3O4LHGJZW3Y","created_at":"2026-05-18T12:33:37.589309+00:00"},{"alias_kind":"pith_short_8","alias_value":"5YPLY3O4","created_at":"2026-05-18T12:33:37.589309+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":46,"internal_anchor_count":46,"sample":[{"citing_arxiv_id":"2606.26923","citing_title":"GAVEL: Grounded Caption Error Verification and Localization","ref_index":53,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26552","citing_title":"Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection","ref_index":21,"is_internal_anchor":true},{"citing_arxiv_id":"2607.02089","citing_title":"ESC: Emotional Self-Correction for Reliable Vision-Language Models","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2606.17030","citing_title":"Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation","ref_index":193,"is_internal_anchor":true},{"citing_arxiv_id":"2606.31257","citing_title":"Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning","ref_index":29,"is_internal_anchor":true},{"citing_arxiv_id":"2606.30393","citing_title":"SADL: What to Ignore? A Benchmark for Subject-Aware Distractor Localization","ref_index":23,"is_internal_anchor":true},{"citing_arxiv_id":"2605.27932","citing_title":"When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2606.00535","citing_title":"DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation","ref_index":96,"is_internal_anchor":true},{"citing_arxiv_id":"2502.13923","citing_title":"Qwen2.5-VL Technical Report","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2605.16269","citing_title":"Train the Trainers -- An Agentic AI Framework for Peer-Based Mental Health Support in Battlefield Environments","ref_index":51,"is_internal_anchor":true},{"citing_arxiv_id":"2605.15876","citing_title":"Unlocking Dense Metric Depth Estimation in VLMs","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2605.08146","citing_title":"VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning","ref_index":56,"is_internal_anchor":true},{"citing_arxiv_id":"2605.15876","citing_title":"Unlocking Dense Metric Depth Estimation in VLMs","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2604.02812","citing_title":"Learning Structured Robot Policies from Vision-Language Models via Synthetic Neuro-Symbolic Supervision","ref_index":25,"is_internal_anchor":true},{"citing_arxiv_id":"2506.06856","citing_title":"Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning","ref_index":2,"is_internal_anchor":true},{"citing_arxiv_id":"2507.13941","citing_title":"Shared representations in brains and models reveal a two-route cortical organization during scene perception","ref_index":92,"is_internal_anchor":true},{"citing_arxiv_id":"2508.10635","citing_title":"ChatENV: An Interactive Vision-Language Model for Sensor-Guided Environmental Monitoring and Scenario Simulation","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"2509.26272","citing_title":"PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection","ref_index":43,"is_internal_anchor":true},{"citing_arxiv_id":"2501.00321","citing_title":"OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning","ref_index":54,"is_internal_anchor":true},{"citing_arxiv_id":"2412.14164","citing_title":"MetaMorph: Multimodal Understanding and Generation via Instruction Tuning","ref_index":274,"is_internal_anchor":true},{"citing_arxiv_id":"2505.05472","citing_title":"Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2603.18373","citing_title":"To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"2605.12034","citing_title":"Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation","ref_index":51,"is_internal_anchor":true},{"citing_arxiv_id":"2409.17146","citing_title":"Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"2603.23607","citing_title":"LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset","ref_index":3,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":2,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/5YPLY3O4LHGJZW3YZ6BQJN3OPY","json":"https://pith.science/pith/5YPLY3O4LHGJZW3YZ6BQJN3OPY.json","graph_json":"https://pith.science/api/pith-number/5YPLY3O4LHGJZW3YZ6BQJN3OPY/graph.json","events_json":"https://pith.science/api/pith-number/5YPLY3O4LHGJZW3YZ6BQJN3OPY/events.json","paper":"https://pith.science/paper/5YPLY3O4"},"agent_actions":{"view_html":"https://pith.science/pith/5YPLY3O4LHGJZW3YZ6BQJN3OPY","download_json":"https://pith.science/pith/5YPLY3O4LHGJZW3YZ6BQJN3OPY.json","view_paper":"https://pith.science/paper/5YPLY3O4","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2410.07073&json=true","fetch_graph":"https://pith.science/api/pith-number/5YPLY3O4LHGJZW3YZ6BQJN3OPY/graph.json","fetch_events":"https://pith.science/api/pith-number/5YPLY3O4LHGJZW3YZ6BQJN3OPY/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/5YPLY3O4LHGJZW3YZ6BQJN3OPY/action/timestamp_anchor","attest_storage":"https://pith.science/pith/5YPLY3O4LHGJZW3YZ6BQJN3OPY/action/storage_attestation","attest_author":"https://pith.science/pith/5YPLY3O4LHGJZW3YZ6BQJN3OPY/action/author_attestation","sign_citation":"https://pith.science/pith/5YPLY3O4LHGJZW3YZ6BQJN3OPY/action/citation_signature","submit_replication":"https://pith.science/pith/5YPLY3O4LHGJZW3YZ6BQJN3OPY/action/replication_record"}},"created_at":"2026-05-17T23:39:19.791379+00:00","updated_at":"2026-05-17T23:39:19.791379+00:00"}