Pith. sign in

Paper Citation Record · LEDGER

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding

As of 8 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2507.04289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04289 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:55:41.094698Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T01:20:57.490329Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T13:36:09.422158Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e558589-f8de-4e52-94e3-437561ebe7c9 · outbound

This paper cites Video-language understanding: A survey from model architecture, model training, and data perspectives.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Video-language understanding: A survey from model architecture, model training, and data perspectives

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.767844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.848927Z digest=sha256:0eee236456b97960a0c7519fb10e28f7ed81763a0be9b17fad38905ea9c650f2

Observation 88150503-529d-4efc-a184-218b1afd31e3 · outbound

This paper cites Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:55:41.566893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.858859Z digest=sha256:adccb5155512582b8108e6d5756f84c67abc5fb87be2e03224c05c616876471b

Observation 61e1c16c-3f56-426a-b742-6e40a45e8ac0 · outbound

This paper cites Dense-captioning events in videos.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Dense-captioning events in videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.559761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.864125Z digest=sha256:8a81b12122dfd13475812d67304281103d6414b43c2d3e9780705a09de63f5fe

Observation ebc40eaf-9248-4aeb-bb84-07ca8691b617 · outbound

This paper cites Towards answering health-related questions from medical videos: Datasets and approaches.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Towards answering health-related questions from medical videos: Datasets and approaches

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.021480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.886100Z digest=sha256:a05274a3f5f491c0ff0ac11d9379d8514bad1ec2d75fe356ad26d68209d72220

Observation f22ac4af-4eff-423e-8976-e0ea004ceb77 · outbound

This paper cites 18 Wangyu Wu, Siqi Song, Xianglin Qiu, Xiaowei Huang, Fei Ma, and Jimin Xiao.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 18 Wangyu Wu, Siqi Song, Xianglin Qiu, Xiaowei Huang, Fei Ma, and Jimin Xiao

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.777750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.894079Z digest=sha256:85077b0c79ed91416b4958c844b7ef93b48f59bef1b10293f5220fdf4f7f9fb5

Observation 9ccbc03a-aaf9-430f-88c4-4edd46f6cdb5 · outbound

This paper cites 19 Zhen Yao, Jiawei Xu, Shuhang Hou, and Mooi Choo Chuah.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 19 Zhen Yao, Jiawei Xu, Shuhang Hou, and Mooi Choo Chuah

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.646010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.904604Z digest=sha256:c069be69c2354fc889e2c88d47a19fb6f116c564b522455ca43b91f3f64a62f9

Observation 09da3542-cf8d-43d1-b446-3fa285a370ac · outbound

This paper cites Overview of the nlpcc 2023 shared task: Chinese medical instructional video question answering.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Overview of the nlpcc 2023 shared task: Chinese medical instructional video question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.298365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.931997Z digest=sha256:dcc613345e34902ae6d5ea6241048dc3b5bb22c67d72525643d70963feae2192

Observation 77563452-5055-4816-bebd-e27259f64448 · outbound

This paper cites 24 Bin Li, Yixuan Weng, Qiya Song, Lianhui Liang, Xianwen Min, and Shoujun Zhou.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 24 Bin Li, Yixuan Weng, Qiya Song, Lianhui Liang, Xianwen Min, and Shoujun Zhou

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.049156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.939915Z digest=sha256:3d5dc349e331fc79ee905dc3b08479450b7ff9def5bda614933ebf1a43562212

Observation 9117250d-5134-4312-8b97-ee9f67768324 · outbound

This paper cites 25 Bin Li, Shenxi Liu, Yixuan Weng, Yue Du, Yuhang Tian, and Shoujun Zhou.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 25 Bin Li, Shenxi Liu, Yixuan Weng, Yue Du, Yuhang Tian, and Shoujun Zhou

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.846058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.955363Z digest=sha256:00830db6165a2adccdc96757dd856155202c4277d8a20d504c5055860fba2065

Observation afd4d8b0-394e-4e57-bc29-8d4255a74d3d · outbound

This paper cites FCMR: Robust evaluation of financial cross-modal multi-hop reasoning.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding FCMR: Robust evaluation of financial cross-modal multi-hop reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.669923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.968361Z digest=sha256:7035014c01a1bfb810c359f31f9d753c1e60644d5d9c413ad44e2f7f929f7231

Observation fd775883-6eaa-4203-9851-fdd29d733145 · outbound

This paper cites Rosenblum.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Rosenblum

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.503693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.981325Z digest=sha256:a2bded735a719fa7cba82a423857c11eb785c3fe770fd1ecae3b207e8b6e663b

Observation abcad994-5377-4ac4-beee-cece2c0fd05e · outbound

This paper cites Learning to Locate Visual Answer in Video Corpus Using Question.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Learning to Locate Visual Answer in Video Corpus Using Question

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:55:41.243092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:41.006696Z digest=sha256:21ace3706edbe88587d58289eb5a68dd18c2cefc675fc103b2f6683438de8e3e

Observation 959f91af-2469-4faa-9d74-095e315f53d7 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.197129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:41.016707Z digest=sha256:a85118b1e34315adef70212ed29d3183c191df052841ec43991e89a89cf0fce4

Observation 6e19e462-a9b6-4018-9333-7a59fa32c54d · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:41.025430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:41.025430Z digest=sha256:e84a52168331d9af4f2dc784c7b14916642e3cc9a74be69787297dc96c570d47

Observation 47f8b8c1-3793-4033-abbe-22c06b1e2e3b · outbound

This paper cites Sentence-BERT: Sentence embeddings using Siamese BERT-networks.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Sentence-BERT: Sentence embeddings using Siamese BERT-networks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.046479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:41.065943Z digest=sha256:cb9be1323820977cfed8ce178b24a8feae5f5117513a9da00cde8b82fa45c328

Observation b47a3d3b-b6cf-498d-824e-6139ffeb4266 · outbound

This paper cites 43 Zhongzhi Chen, Guang Liu, Bo-Wen Zhang, Qinghong Yang, and Ledell Wu.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 43 Zhongzhi Chen, Guang Liu, Bo-Wen Zhang, Qinghong Yang, and Ledell Wu

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:41.876846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:41.094698Z digest=sha256:20f8f3994edc7b7e5b7f89049855177c4a15f742f49f25edf4b8e769b749fe50

Observation 1dbacb54-5993-45b6-bffd-af70585129a2 · outbound

This paper cites Learning to segment actions from visual and language instructions via differentiable weak sequence alignment.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Learning to segment actions from visual and language instructions via differentiable weak sequence alignment

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.376483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.869419Z digest=sha256:e804366a7b817ddfb395c81fc918114823f6835bf663bdb12f80da37ce0d7a99

Observation 2c5f6c14-1043-4637-aa52-287ddd85ba81 · outbound

This paper cites 29 Yixuan Weng and Bin Li.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding 29 Yixuan Weng and Bin Li

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:42.371022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.997756Z digest=sha256:a794ffac39a8cec15a33b400d8d99d325e87809a6d92c1141ec7dad2d4202996

Observation 473ed0fc-616c-46b1-a6b0-c80447ec3472 · outbound

This paper cites NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding

Reference 2020

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:55:41.397567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.879906Z digest=sha256:6e461905d4751ec500c1db4b76f9db9d8230d9350414563c60b7f877d33598f7

Observation e78c0636-7dc6-45ad-8a77-2beb9fe0786f · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Movienet: A holistic dataset for movie understanding

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.194492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.874654Z digest=sha256:18e49d8002c6237d2aae6b4228d41e5126a1c6dbdde1149065b0419216d43a45

Observation b1f7998a-f06c-4a11-8ce7-f89296260b84 · outbound

This paper cites Buch, Cristobal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Buch, Cristobal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:43.500908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.918432Z digest=sha256:faa1600e4636f99c0325fa2d4a77156f2e02dc1dd30907c5eb4b1f4603d51d48

Observation fec5053b-15a4-418f-8b56-792bc45af0a3 · outbound

This paper cites Sci China Inf Sci 16 5 Xingjian Diao, Chunhui Zhang, Weiyi Wu, Zhongyu Ouyang, Peijun Qing, Ming Cheng, Soroush Vosoughi, and Jiang Gui.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Sci China Inf Sci 16 5 Xingjian Diao, Chunhui Zhang, Weiyi Wu, Zhongyu Ouyang, Peijun Qing, Ming Cheng, Soroush Vosoughi, and Jiang Gui

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:44.923865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.837430Z digest=sha256:f7069a9a801d8a6261af73caa478ef5964c2cf6b91aca133d125c60d00185281

Observation ae9b021d-1696-414d-a9a4-0abf85f939f4 · outbound

This paper cites Learning to Unify Audio, Visual and Text for Audio-Enhanced Multilingual Visual Answer Localization.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Learning to Unify Audio, Visual and Text for Audio-Enhanced Multilingual Visual Answer Localization

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:55:41.732422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:55:40.853738Z digest=sha256:2cb3dd4355377044ac7ce790f11faf9e00eb9b2fa75895744963a83535ea18a5

Observation f3f3491b-2274-4d88-86c8-02bebe49370f · outbound

This paper cites Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:40.843380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:40.843380Z digest=sha256:a4e56c897ead179b2847fc4b6dc4d754911997f73ea9b77d660e554cde166b28

Pith citing papers

Observation 933f59fc-0aec-49c4-84c3-d75423ceffb0 · inbound

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark cites this paper.

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.424530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:20:57.490329Z digest=sha256:4267e4feb8cb95717ed29551892bd57478632b02cfa1224484558a6a465b2fac