{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:CMLQZHZF52ZYEA6EEZ5KOO334M","short_pith_number":"pith:CMLQZHZF","schema_version":"1.0","canonical_sha256":"13170c9f25eeb38203c4267aa73b7be32448a922051493abea9a59b0c8340f3d","source":{"kind":"arxiv","id":"2410.10855","version":4},"attestation_state":"computed","paper":{"title":"Core Knowledge Deficits in Multi-Modal Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CV"],"primary_cat":"cs.CL","authors_text":"Bingyang Wang, Dezhi Luo, Haiyun Lyu, Haoran Sun, Hokin Deng, Nuno Vasconcelos, Qingying Gao, Robert D. Hawkins, Tal Golan, Tianwei Zhao, Yijiang Li","submitted_at":"2024-10-06T20:13:11Z","abstract_excerpt":"While Multi-modal Large Language Models (MLLMs) demonstrate impressive abilities over high-level perception and reasoning, their robustness in the wild remains limited, often falling short on tasks that are intuitive and effortless for humans. We examine the hypothesis that these deficiencies stem from the absence of core knowledge--rudimentary cognitive abilities innate to humans from early childhood. To explore the core knowledge representation in MLLMs, we introduce CoreCognition, a large-scale benchmark encompassing 12 core knowledge concepts grounded in developmental cognitive science. We"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2410.10855","kind":"arxiv","version":4},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2024-10-06T20:13:11Z","cross_cats_sorted":["cs.AI","cs.CV"],"title_canon_sha256":"a13e86d11bef95b541ddd34b0700c54768f8a46bbc64943c486eb0521eb243e0","abstract_canon_sha256":"a2830070db4ab5ee2a04680a4058e059efce8462757749bb0713badfcdd29088"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:24:07.216935Z","signature_b64":"VDWdgiGZSGwBquu/yJOXSpjdV+QbynrYF63I8bSyMak+PCMzTqJCqS+o41FvFfi6nt3VvcDV7nBqIuaBPQTtCA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"13170c9f25eeb38203c4267aa73b7be32448a922051493abea9a59b0c8340f3d","last_reissued_at":"2026-07-05T11:24:07.216410Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:24:07.216410Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Core Knowledge Deficits in Multi-Modal Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CV"],"primary_cat":"cs.CL","authors_text":"Bingyang Wang, Dezhi Luo, Haiyun Lyu, Haoran Sun, Hokin Deng, Nuno Vasconcelos, Qingying Gao, Robert D. Hawkins, Tal Golan, Tianwei Zhao, Yijiang Li","submitted_at":"2024-10-06T20:13:11Z","abstract_excerpt":"While Multi-modal Large Language Models (MLLMs) demonstrate impressive abilities over high-level perception and reasoning, their robustness in the wild remains limited, often falling short on tasks that are intuitive and effortless for humans. We examine the hypothesis that these deficiencies stem from the absence of core knowledge--rudimentary cognitive abilities innate to humans from early childhood. To explore the core knowledge representation in MLLMs, we introduce CoreCognition, a large-scale benchmark encompassing 12 core knowledge concepts grounded in developmental cognitive science. We"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2410.10855","kind":"arxiv","version":4},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2410.10855/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2410.10855","created_at":"2026-07-05T11:24:07.216475+00:00"},{"alias_kind":"arxiv_version","alias_value":"2410.10855v4","created_at":"2026-07-05T11:24:07.216475+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2410.10855","created_at":"2026-07-05T11:24:07.216475+00:00"},{"alias_kind":"pith_short_12","alias_value":"CMLQZHZF52ZY","created_at":"2026-07-05T11:24:07.216475+00:00"},{"alias_kind":"pith_short_16","alias_value":"CMLQZHZF52ZYEA6E","created_at":"2026-07-05T11:24:07.216475+00:00"},{"alias_kind":"pith_short_8","alias_value":"CMLQZHZF","created_at":"2026-07-05T11:24:07.216475+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":4,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.07872","citing_title":"VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2606.05497","citing_title":"LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, \"Is Your VLM Smarter Than a 5th Grader?\")","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2605.31148","citing_title":"SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2509.02547","citing_title":"The Landscape of Agentic Reinforcement Learning for LLMs: A Survey","ref_index":215,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/CMLQZHZF52ZYEA6EEZ5KOO334M","json":"https://pith.science/pith/CMLQZHZF52ZYEA6EEZ5KOO334M.json","graph_json":"https://pith.science/api/pith-number/CMLQZHZF52ZYEA6EEZ5KOO334M/graph.json","events_json":"https://pith.science/api/pith-number/CMLQZHZF52ZYEA6EEZ5KOO334M/events.json","paper":"https://pith.science/paper/CMLQZHZF"},"agent_actions":{"view_html":"https://pith.science/pith/CMLQZHZF52ZYEA6EEZ5KOO334M","download_json":"https://pith.science/pith/CMLQZHZF52ZYEA6EEZ5KOO334M.json","view_paper":"https://pith.science/paper/CMLQZHZF","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2410.10855&json=true","fetch_graph":"https://pith.science/api/pith-number/CMLQZHZF52ZYEA6EEZ5KOO334M/graph.json","fetch_events":"https://pith.science/api/pith-number/CMLQZHZF52ZYEA6EEZ5KOO334M/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/CMLQZHZF52ZYEA6EEZ5KOO334M/action/timestamp_anchor","attest_storage":"https://pith.science/pith/CMLQZHZF52ZYEA6EEZ5KOO334M/action/storage_attestation","attest_author":"https://pith.science/pith/CMLQZHZF52ZYEA6EEZ5KOO334M/action/author_attestation","sign_citation":"https://pith.science/pith/CMLQZHZF52ZYEA6EEZ5KOO334M/action/citation_signature","submit_replication":"https://pith.science/pith/CMLQZHZF52ZYEA6EEZ5KOO334M/action/replication_record"}},"created_at":"2026-07-05T11:24:07.216475+00:00","updated_at":"2026-07-05T11:24:07.216475+00:00"}