{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:WVHAAXEGZ2A234FEDEVHW3FPGR","short_pith_number":"pith:WVHAAXEG","schema_version":"1.0","canonical_sha256":"b54e005c86ce81adf0a4192a7b6caf3461c1d0769932918777975be7a08d098e","source":{"kind":"arxiv","id":"2408.07981","version":1},"attestation_state":"computed","paper":{"title":"LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Brian R Quaranto, Garrett Skinner, Gene Yang, Jiajie Li, Jinjun Xiong, Peter C W Kim, Steven D Schwaitzberg","submitted_at":"2024-08-15T07:00:20Z","abstract_excerpt":"Multimodal large language models (LLMs) have achieved notable success across various domains, while research in the medical field has largely focused on unimodal images. Meanwhile, current general-domain multimodal models for videos still lack the capabilities to understand and engage in conversations about surgical videos. One major contributing factor is the absence of datasets in the surgical field. In this paper, we create a new dataset, Surg-QA, consisting of 102,000 surgical video-instruction pairs, the largest of its kind so far. To build such a dataset, we propose a novel two-stage que"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2408.07981","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.CV","submitted_at":"2024-08-15T07:00:20Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"c6c5238885a0be948e23f6ce80a916bbed3bb82c929c04cf30c9e5ee9f36a834","abstract_canon_sha256":"1fdb33be028877fd68aa19463810fda1a40cbbc8fda15bf5426eadfcf9d64aec"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:55:42.742887Z","signature_b64":"2no0R7/46/1+N0p93fjTaDS5V2JmEGJg5HhQbGGNwV9GwQh9X0ma7ZSQdyKn3DDvn+yH5lb8wMeKjkZJIBSWCQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"b54e005c86ce81adf0a4192a7b6caf3461c1d0769932918777975be7a08d098e","last_reissued_at":"2026-07-05T08:55:42.742480Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:55:42.742480Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Brian R Quaranto, Garrett Skinner, Gene Yang, Jiajie Li, Jinjun Xiong, Peter C W Kim, Steven D Schwaitzberg","submitted_at":"2024-08-15T07:00:20Z","abstract_excerpt":"Multimodal large language models (LLMs) have achieved notable success across various domains, while research in the medical field has largely focused on unimodal images. Meanwhile, current general-domain multimodal models for videos still lack the capabilities to understand and engage in conversations about surgical videos. One major contributing factor is the absence of datasets in the surgical field. In this paper, we create a new dataset, Surg-QA, consisting of 102,000 surgical video-instruction pairs, the largest of its kind so far. To build such a dataset, we propose a novel two-stage que"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2408.07981","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2408.07981/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2408.07981","created_at":"2026-07-05T08:55:42.742537+00:00"},{"alias_kind":"arxiv_version","alias_value":"2408.07981v1","created_at":"2026-07-05T08:55:42.742537+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2408.07981","created_at":"2026-07-05T08:55:42.742537+00:00"},{"alias_kind":"pith_short_12","alias_value":"WVHAAXEGZ2A2","created_at":"2026-07-05T08:55:42.742537+00:00"},{"alias_kind":"pith_short_16","alias_value":"WVHAAXEGZ2A234FE","created_at":"2026-07-05T08:55:42.742537+00:00"},{"alias_kind":"pith_short_8","alias_value":"WVHAAXEG","created_at":"2026-07-05T08:55:42.742537+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":6,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2607.01751","citing_title":"MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07433","citing_title":"Watch, Remember, Reason: Human-View Video Understanding with MLLMs","ref_index":265,"is_internal_anchor":false},{"citing_arxiv_id":"2606.27484","citing_title":"Fine-tuning a multimodal large language model for clinician-grade autism behavioral scoring from short home videos","ref_index":49,"is_internal_anchor":false},{"citing_arxiv_id":"2605.21132","citing_title":"SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2603.16024","citing_title":"Speak, Segment, Track, Navigate: An Interactive System for Video-Guided Skull-Base Surgery","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08712","citing_title":"From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation","ref_index":42,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/WVHAAXEGZ2A234FEDEVHW3FPGR","json":"https://pith.science/pith/WVHAAXEGZ2A234FEDEVHW3FPGR.json","graph_json":"https://pith.science/api/pith-number/WVHAAXEGZ2A234FEDEVHW3FPGR/graph.json","events_json":"https://pith.science/api/pith-number/WVHAAXEGZ2A234FEDEVHW3FPGR/events.json","paper":"https://pith.science/paper/WVHAAXEG"},"agent_actions":{"view_html":"https://pith.science/pith/WVHAAXEGZ2A234FEDEVHW3FPGR","download_json":"https://pith.science/pith/WVHAAXEGZ2A234FEDEVHW3FPGR.json","view_paper":"https://pith.science/paper/WVHAAXEG","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2408.07981&json=true","fetch_graph":"https://pith.science/api/pith-number/WVHAAXEGZ2A234FEDEVHW3FPGR/graph.json","fetch_events":"https://pith.science/api/pith-number/WVHAAXEGZ2A234FEDEVHW3FPGR/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/WVHAAXEGZ2A234FEDEVHW3FPGR/action/timestamp_anchor","attest_storage":"https://pith.science/pith/WVHAAXEGZ2A234FEDEVHW3FPGR/action/storage_attestation","attest_author":"https://pith.science/pith/WVHAAXEGZ2A234FEDEVHW3FPGR/action/author_attestation","sign_citation":"https://pith.science/pith/WVHAAXEGZ2A234FEDEVHW3FPGR/action/citation_signature","submit_replication":"https://pith.science/pith/WVHAAXEGZ2A234FEDEVHW3FPGR/action/replication_record"}},"created_at":"2026-07-05T08:55:42.742537+00:00","updated_at":"2026-07-05T08:55:42.742537+00:00"}