Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:55:41.094698Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2507.04289.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:55:41.094698Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T01:20:57.490329Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T13:36:09.422158Z
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8e558589-f8de-4e52-94e3-437561ebe7c9 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Video-language understanding: A survey from model architecture, model training, and data perspectives
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 88150503-529d-4efc-a184-218b1afd31e3 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 61e1c16c-3f56-426a-b742-6e40a45e8ac0 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Dense-captioning events in videos
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ebc40eaf-9248-4aeb-bb84-07ca8691b617 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Towards answering health-related questions from medical videos: Datasets and approaches
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f22ac4af-4eff-423e-8976-e0ea004ceb77 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 18 Wangyu Wu, Siqi Song, Xianglin Qiu, Xiaowei Huang, Fei Ma, and Jimin Xiao
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9ccbc03a-aaf9-430f-88c4-4edd46f6cdb5 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 19 Zhen Yao, Jiawei Xu, Shuhang Hou, and Mooi Choo Chuah
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 09da3542-cf8d-43d1-b446-3fa285a370ac · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Overview of the nlpcc 2023 shared task: Chinese medical instructional video question answering
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 77563452-5055-4816-bebd-e27259f64448 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 24 Bin Li, Yixuan Weng, Qiya Song, Lianhui Liang, Xianwen Min, and Shoujun Zhou
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9117250d-5134-4312-8b97-ee9f67768324 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 25 Bin Li, Shenxi Liu, Yixuan Weng, Yue Du, Yuhang Tian, and Shoujun Zhou
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afd4d8b0-394e-4e57-bc29-8d4255a74d3d · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding FCMR: Robust evaluation of financial cross-modal multi-hop reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fd775883-6eaa-4203-9851-fdd29d733145 · outbound
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation abcad994-5377-4ac4-beee-cece2c0fd05e · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Learning to Locate Visual Answer in Video Corpus Using Question
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 959f91af-2469-4faa-9d74-095e315f53d7 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6e19e462-a9b6-4018-9333-7a59fa32c54d · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47f8b8c1-3793-4033-abbe-22c06b1e2e3b · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b47a3d3b-b6cf-498d-824e-6139ffeb4266 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 43 Zhongzhi Chen, Guang Liu, Bo-Wen Zhang, Qinghong Yang, and Ledell Wu
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1dbacb54-5993-45b6-bffd-af70585129a2 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Learning to segment actions from visual and language instructions via differentiable weak sequence alignment
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2c5f6c14-1043-4637-aa52-287ddd85ba81 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 29 Yixuan Weng and Bin Li
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 473ed0fc-616c-46b1-a6b0-c80447ec3472 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e78c0636-7dc6-45ad-8a77-2beb9fe0786f · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Movienet: A holistic dataset for movie understanding
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b1f7998a-f06c-4a11-8ce7-f89296260b84 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Buch, Cristobal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fec5053b-15a4-418f-8b56-792bc45af0a3 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Sci China Inf Sci 16 5 Xingjian Diao, Chunhui Zhang, Weiyi Wu, Zhongyu Ouyang, Peijun Qing, Ming Cheng, Soroush Vosoughi, and Jiang Gui
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ae9b021d-1696-414d-a9a4-0abf85f939f4 · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Learning to Unify Audio, Visual and Text for Audio-Enhanced Multilingual Visual Answer Localization
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f3f3491b-2274-4d88-86c8-02bebe49370f · outbound
M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933f59fc-0aec-49c4-84c3-d75423ceffb0 · inbound
SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.