{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:D3Y2XKCAE2G2CXCGDZWDUF5CX6","short_pith_number":"pith:D3Y2XKCA","schema_version":"1.0","canonical_sha256":"1ef1aba840268da15c461e6c3a17a2bfbf949e3703cfe0c4de17869d1f68799b","source":{"kind":"arxiv","id":"2311.01813","version":3},"attestation_state":"computed","paper":{"title":"FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Lei Li, Lu Hou, Rundong Gao, Shicheng Li, Shuhuai Ren, Sishuo Chen, Xu Sun, Yuanxin Liu","submitted_at":"2023-11-03T09:46:05Z","abstract_excerpt":"Recently, open-domain text-to-video (T2V) generation models have made remarkable progress. However, the promising results are mainly shown by the qualitative cases of generated videos, while the quantitative evaluation of T2V models still faces two critical problems. Firstly, existing studies lack fine-grained evaluation of T2V models on different categories of text prompts. Although some benchmarks have categorized the prompts, their categorization either only focuses on a single aspect or fails to consider the temporal information in video generation. Secondly, it is unclear whether the auto"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2311.01813","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CV","submitted_at":"2023-11-03T09:46:05Z","cross_cats_sorted":[],"title_canon_sha256":"32391124ddc7dd6f3191b0dfe617fea97907e747fdfe4e9eec792c07f7955f48","abstract_canon_sha256":"7cf33b7f4638c2152f7ace16438fe69a220ba82b87d09c37960b42ea87b1275c"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:27:52.322512Z","signature_b64":"3jsuUTiXRO5d8wKa6S5hQoNBWkAVFbbvGUZKh4OxGW6765lFlq4BpHmRGNvr8ohNPmUa3JYyA0YcoQyImfJaAA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"1ef1aba840268da15c461e6c3a17a2bfbf949e3703cfe0c4de17869d1f68799b","last_reissued_at":"2026-07-05T07:27:52.321777Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:27:52.321777Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Lei Li, Lu Hou, Rundong Gao, Shicheng Li, Shuhuai Ren, Sishuo Chen, Xu Sun, Yuanxin Liu","submitted_at":"2023-11-03T09:46:05Z","abstract_excerpt":"Recently, open-domain text-to-video (T2V) generation models have made remarkable progress. However, the promising results are mainly shown by the qualitative cases of generated videos, while the quantitative evaluation of T2V models still faces two critical problems. Firstly, existing studies lack fine-grained evaluation of T2V models on different categories of text prompts. Although some benchmarks have categorized the prompts, their categorization either only focuses on a single aspect or fails to consider the temporal information in video generation. Secondly, it is unclear whether the auto"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2311.01813","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2311.01813/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2311.01813","created_at":"2026-07-05T07:27:52.321859+00:00"},{"alias_kind":"arxiv_version","alias_value":"2311.01813v3","created_at":"2026-07-05T07:27:52.321859+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2311.01813","created_at":"2026-07-05T07:27:52.321859+00:00"},{"alias_kind":"pith_short_12","alias_value":"D3Y2XKCAE2G2","created_at":"2026-07-05T07:27:52.321859+00:00"},{"alias_kind":"pith_short_16","alias_value":"D3Y2XKCAE2G2CXCG","created_at":"2026-07-05T07:27:52.321859+00:00"},{"alias_kind":"pith_short_8","alias_value":"D3Y2XKCA","created_at":"2026-07-05T07:27:52.321859+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":4,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.26244","citing_title":"LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2605.26918","citing_title":"Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2504.17180","citing_title":"We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback","ref_index":57,"is_internal_anchor":false},{"citing_arxiv_id":"2604.19193","citing_title":"How Far Are Video Models from True Multimodal Reasoning?","ref_index":46,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/D3Y2XKCAE2G2CXCGDZWDUF5CX6","json":"https://pith.science/pith/D3Y2XKCAE2G2CXCGDZWDUF5CX6.json","graph_json":"https://pith.science/api/pith-number/D3Y2XKCAE2G2CXCGDZWDUF5CX6/graph.json","events_json":"https://pith.science/api/pith-number/D3Y2XKCAE2G2CXCGDZWDUF5CX6/events.json","paper":"https://pith.science/paper/D3Y2XKCA"},"agent_actions":{"view_html":"https://pith.science/pith/D3Y2XKCAE2G2CXCGDZWDUF5CX6","download_json":"https://pith.science/pith/D3Y2XKCAE2G2CXCGDZWDUF5CX6.json","view_paper":"https://pith.science/paper/D3Y2XKCA","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2311.01813&json=true","fetch_graph":"https://pith.science/api/pith-number/D3Y2XKCAE2G2CXCGDZWDUF5CX6/graph.json","fetch_events":"https://pith.science/api/pith-number/D3Y2XKCAE2G2CXCGDZWDUF5CX6/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/D3Y2XKCAE2G2CXCGDZWDUF5CX6/action/timestamp_anchor","attest_storage":"https://pith.science/pith/D3Y2XKCAE2G2CXCGDZWDUF5CX6/action/storage_attestation","attest_author":"https://pith.science/pith/D3Y2XKCAE2G2CXCGDZWDUF5CX6/action/author_attestation","sign_citation":"https://pith.science/pith/D3Y2XKCAE2G2CXCGDZWDUF5CX6/action/citation_signature","submit_replication":"https://pith.science/pith/D3Y2XKCAE2G2CXCGDZWDUF5CX6/action/replication_record"}},"created_at":"2026-07-05T07:27:52.321859+00:00","updated_at":"2026-07-05T07:27:52.321859+00:00"}