{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:7LPG2P7FHIKLHNUM5TVFEZTVDO","short_pith_number":"pith:7LPG2P7F","schema_version":"1.0","canonical_sha256":"fade6d3fe53a14b3b68cecea5266751b9233045b9ff189c54afa683d896f6b79","source":{"kind":"arxiv","id":"2308.16890","version":2},"attestation_state":"computed","paper":{"title":"TouchStone: Evaluating Vision-Language Models by Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Chang Zhou, Jingren Zhou, Jinze Bai, Junyang Lin, Peng Wang, Shuai Bai, Shusheng Yang, Xinggang Wang, Xingxuan Zhang","submitted_at":"2023-08-31T17:52:04Z","abstract_excerpt":"Large vision-language models (LVLMs) have recently witnessed rapid advancements, exhibiting a remarkable capacity for perceiving, understanding, and processing visual information by connecting visual receptor with large language models (LLMs). However, current assessments mainly focus on recognizing and reasoning abilities, lacking direct evaluation of conversational skills and neglecting visual storytelling abilities. In this paper, we propose an evaluation method that uses strong LLMs as judges to comprehensively evaluate the various abilities of LVLMs. Firstly, we construct a comprehensive "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2308.16890","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2023-08-31T17:52:04Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"6cd7866d4802fdd89b61cc435fba2eb31aa66fc45fb3a46824d5553d2d78e859","abstract_canon_sha256":"868ec35b652d689479364c29ceea2d5dbe0150e6f6e8e5db270dde8d25e2dd2c"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T06:47:28.569978Z","signature_b64":"Pd8SQ8FydS9Ocrqczff6tVgo8DOUQi2Gmujfi7R4uyxGfBELcNVjLuSU4qU/lfHSUL/hD7Tg17uhoYBeToJrDw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"fade6d3fe53a14b3b68cecea5266751b9233045b9ff189c54afa683d896f6b79","last_reissued_at":"2026-07-05T06:47:28.569456Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T06:47:28.569456Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"TouchStone: Evaluating Vision-Language Models by Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Chang Zhou, Jingren Zhou, Jinze Bai, Junyang Lin, Peng Wang, Shuai Bai, Shusheng Yang, Xinggang Wang, Xingxuan Zhang","submitted_at":"2023-08-31T17:52:04Z","abstract_excerpt":"Large vision-language models (LVLMs) have recently witnessed rapid advancements, exhibiting a remarkable capacity for perceiving, understanding, and processing visual information by connecting visual receptor with large language models (LLMs). However, current assessments mainly focus on recognizing and reasoning abilities, lacking direct evaluation of conversational skills and neglecting visual storytelling abilities. In this paper, we propose an evaluation method that uses strong LLMs as judges to comprehensively evaluate the various abilities of LVLMs. Firstly, we construct a comprehensive "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2308.16890","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2308.16890/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2308.16890","created_at":"2026-07-05T06:47:28.569522+00:00"},{"alias_kind":"arxiv_version","alias_value":"2308.16890v2","created_at":"2026-07-05T06:47:28.569522+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2308.16890","created_at":"2026-07-05T06:47:28.569522+00:00"},{"alias_kind":"pith_short_12","alias_value":"7LPG2P7FHIKL","created_at":"2026-07-05T06:47:28.569522+00:00"},{"alias_kind":"pith_short_16","alias_value":"7LPG2P7FHIKLHNUM","created_at":"2026-07-05T06:47:28.569522+00:00"},{"alias_kind":"pith_short_8","alias_value":"7LPG2P7F","created_at":"2026-07-05T06:47:28.569522+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.17030","citing_title":"Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation","ref_index":94,"is_internal_anchor":false},{"citing_arxiv_id":"2606.06217","citing_title":"DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments","ref_index":41,"is_internal_anchor":false},{"citing_arxiv_id":"2403.00476","citing_title":"TempCompass: Do Video LLMs Really Understand Videos?","ref_index":73,"is_internal_anchor":false},{"citing_arxiv_id":"2408.13257","citing_title":"MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13173","citing_title":"OxyEcomBench: Benchmarking Multimodal Foundation Models across E-Commerce Ecosystems","ref_index":15,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/7LPG2P7FHIKLHNUM5TVFEZTVDO","json":"https://pith.science/pith/7LPG2P7FHIKLHNUM5TVFEZTVDO.json","graph_json":"https://pith.science/api/pith-number/7LPG2P7FHIKLHNUM5TVFEZTVDO/graph.json","events_json":"https://pith.science/api/pith-number/7LPG2P7FHIKLHNUM5TVFEZTVDO/events.json","paper":"https://pith.science/paper/7LPG2P7F"},"agent_actions":{"view_html":"https://pith.science/pith/7LPG2P7FHIKLHNUM5TVFEZTVDO","download_json":"https://pith.science/pith/7LPG2P7FHIKLHNUM5TVFEZTVDO.json","view_paper":"https://pith.science/paper/7LPG2P7F","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2308.16890&json=true","fetch_graph":"https://pith.science/api/pith-number/7LPG2P7FHIKLHNUM5TVFEZTVDO/graph.json","fetch_events":"https://pith.science/api/pith-number/7LPG2P7FHIKLHNUM5TVFEZTVDO/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/7LPG2P7FHIKLHNUM5TVFEZTVDO/action/timestamp_anchor","attest_storage":"https://pith.science/pith/7LPG2P7FHIKLHNUM5TVFEZTVDO/action/storage_attestation","attest_author":"https://pith.science/pith/7LPG2P7FHIKLHNUM5TVFEZTVDO/action/author_attestation","sign_citation":"https://pith.science/pith/7LPG2P7FHIKLHNUM5TVFEZTVDO/action/citation_signature","submit_replication":"https://pith.science/pith/7LPG2P7FHIKLHNUM5TVFEZTVDO/action/replication_record"}},"created_at":"2026-07-05T06:47:28.569522+00:00","updated_at":"2026-07-05T06:47:28.569522+00:00"}