Pith. sign in

Paper Citation Record · LEDGER

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection

As of 12 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2501.10692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10692 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:05:48.211356Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact3
  • verified fuzzy22
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db8af9e1-8c39-4713-8d3c-e623fa29db71 · outbound

This paper cites an unresolved cited work.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:05:49.166365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:47.988030Z digest=sha256:a08f70cb54504c3e63db3f5f11a017450b867b54b6525b0f95aa7e9025d663cd

Observation 643381ef-c2b0-4fa5-9f2d-59eb4ecdd82c · outbound

This paper cites an unresolved cited work.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:05:49.136619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:47.994781Z digest=sha256:0f12151356e10a5786a15e52f2b1dff9092099d640d35122229e7c545a4cea88

Observation c43ec5cb-6aac-43bb-ba86-362b02fd13de · outbound

This paper cites 2 Related work Most previous MR&HD approaches [5, 11, 12] only em- ploy image and text inputs.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection 2 Related work Most previous MR&HD approaches [5, 11, 12] only em- ploy image and text inputs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.115294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.001085Z digest=sha256:6c42baa6600f18b20f70793f5187dab0131400fff535582aece488b25973217b

Observation a2afd7b4-877e-4000-bd7f-f1507ba26e52 · outbound

This paper cites Effectiveness of each module in MRNet on QVHigh- lights val split.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Effectiveness of each module in MRNet on QVHigh- lights val split

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.087752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.008082Z digest=sha256:20f29cb1acd3598e70a51d8e5fdd94515d6755ce06599414243d0650cab08216

Observation a007fb0a-100a-4d32-971e-d10c7194861e · outbound

This paper cites The main reason is that Moment-DETR only utilizes RGB, which fails to fully under- Table 5.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection The main reason is that Moment-DETR only utilizes RGB, which fails to fully under- Table 5

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.054624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.015150Z digest=sha256:cbb5b5ed6b8a9bcaa41786930ccc00303c450718b521562f95c4ffccd0bfa5a7

Observation 2a1b707c-b3d1-49a0-b54c-cb00fc211b4e · outbound

This paper cites Localizing moments in video with natural lan- guage,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Localizing moments in video with natural lan- guage,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.034835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.020785Z digest=sha256:d8986f64a249fbd3331c899a74c421b1e8ba270685ba075f4a73a7f76938b467

Observation c00b4893-4ba6-4451-b5f7-87942f4596af · outbound

This paper cites Less is more: Learning highlight detection from video duration,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Less is more: Learning highlight detection from video duration,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.013948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.026677Z digest=sha256:b2c0a5d6b08d19c52ea8734b0e3e1ba6cd13d63961a0e2bcea8a8b3efa14b184

Observation e4951687-383d-45cd-8e87-f5c8c75176d6 · outbound

This paper cites Detect- ing Moments and Highlights in Videos via Natural Lan- guage Queries,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Detect- ing Moments and Highlights in Videos via Natural Lan- guage Queries,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.990075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.033453Z digest=sha256:1d480be5ff760deff8b1deab9fe59356a5e114ee9ff41aded0dd2f03dfc4b223

Observation 5c3af061-9e22-46ab-bc1b-32290657fc3e · outbound

This paper cites UMT: Unified Multi-modal Transformers for Joint Video Moment Re- trieval and Highlight Detection,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection UMT: Unified Multi-modal Transformers for Joint Video Moment Re- trieval and Highlight Detection,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.959393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.038646Z digest=sha256:627bef4e1077c3f41c9525bf14c84bbe13825ff8e38f85063f6e36a2d8e244f6

Observation 9c8c0675-fbc7-4842-ab7a-46be91c2877a · outbound

This paper cites MomentDiff: Generative Video Moment Retrieval from Random to Real.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection MomentDiff: Generative Video Moment Retrieval from Random to Real

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:05:48.545907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.046112Z digest=sha256:97a96e48d67f1134302046fda160cd3789ca57e2cac8338567e994a63088a612

Observation 8c05860d-5739-4c27-9f29-94bb911e5446 · outbound

This paper cites Gmflow: Learning optical flow via global matching,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Gmflow: Learning optical flow via global matching,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.934703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.061247Z digest=sha256:f30268204ef6a245f2490ba58451d8139373744e57c614b259b5b08a7966beff

Observation 0e1917f8-3afa-4cdc-8c29-bfc7eb45f657 · outbound

This paper cites Depth- cooperated trimodal network for video salient object de- tection,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Depth- cooperated trimodal network for video salient object de- tection,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.914079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.071804Z digest=sha256:423cbc95a84dc286864e2d91cf60ff200e527642744181a7f29c8d3f7d8efdf1

Observation 9fb4ab9b-c1aa-4395-a6fc-23887b50d271 · outbound

This paper cites Pyramid Feature Attention Network for Monocular Depth Pre- diction,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Pyramid Feature Attention Network for Monocular Depth Pre- diction,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.888167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.084100Z digest=sha256:a7004eab281548a5afef06715855e024768c0cf493dbf92152dd458ff05c93ef

Observation 29f57026-dfaf-4e31-9962-5bfcc611925f · outbound

This paper cites How hierarchical is language use?,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection How hierarchical is language use?,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.864654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.091445Z digest=sha256:07d23d74c56ddcf7b718a9b79142d86abfe919c0036a678d457c4b2fa206b17a

Observation 1abb676d-9dce-4296-bce8-ef2645e45d9b · outbound

This paper cites The emergence of hierarchical structure in human language,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection The emergence of hierarchical structure in human language,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.843168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.097892Z digest=sha256:7d45d679a786236e302893f02b55ea9183df59bae42b1abca05fa1d9b663a4e3

Observation eb7e6f6a-cb38-494f-8377-fc6072587a61 · outbound

This paper cites GPTSee: Enhancing moment retrieval and highlight detection via description-based similarity features,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection GPTSee: Enhancing moment retrieval and highlight detection via description-based similarity features,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.812275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.106113Z digest=sha256:b867cc09f629d32b7b7a75141c4bae5735a7dc9940abf83cc63910d08c026006

Observation 7f5b4600-e35f-4a04-85b9-265ac4897fc6 · outbound

This paper cites MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.113507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.113507Z digest=sha256:f1d462cc4ae5306e3a69128683f310bad923c5e9bae69b6532ca7d40e57e9c61

Observation 008f428b-8dcf-4a60-9f1b-351eb1da1fa2 · outbound

This paper cites An empirical study of end-to-end video-language transformers with masked visual modeling,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection An empirical study of end-to-end video-language transformers with masked visual modeling,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.781442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.119107Z digest=sha256:08c605955741dbc7e5e906e602f63cc35737ddb2347f238611b67333d3a5873a

Observation db86f254-899d-4769-902c-65d7b044870c · outbound

This paper cites End-to-end object detection with transformers,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection End-to-end object detection with transformers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.758557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.124164Z digest=sha256:23fcee65eed746d9a21edaf28c455efd7c7f1200f3e10614e6abffe9e422b4fc

Observation d51e2849-68c4-4d12-b15e-08fc9156fa70 · outbound

This paper cites Shifting more attention to visual backbone: Query-modulated re- finement networks for end-to-end visual grounding,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Shifting more attention to visual backbone: Query-modulated re- finement networks for end-to-end visual grounding,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.732109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.131702Z digest=sha256:affbcd0031032d08357453d0d8c039b9ce6ca85f8b075797c70a154afce6baa9

Observation e3fd3cd1-dd33-4563-a411-8c65390926cd · outbound

This paper cites Re- thinking transformer-based set prediction for object de- tection,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Re- thinking transformer-based set prediction for object de- tection,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.710933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.138847Z digest=sha256:c4e696bf63f322495ae30c08304bc0746d9bbd2d6c3a35459bc6192d4297fc8d

Observation 539371cd-619f-48a9-86e5-08a1ae89813d · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.144753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.144753Z digest=sha256:3ed39cb692b14cd11fcfba7589ce8758857e87f18856ac39d9eb21471dc30bd1

Observation 85873ffe-579d-4028-88f7-4c0ea041b52d · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Learning transferable visual models from natural lan- guage supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.685286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.150146Z digest=sha256:de43103e2ea944d600b26b9e585ee9162379b8d95a192042748e692791432e36

Observation ef50378e-ec48-4585-a08d-3416706516b6 · outbound

This paper cites Recognizing American Sign Language Manual Signs from RGB-D Videos.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Recognizing American Sign Language Manual Signs from RGB-D Videos

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:05:48.435257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.157147Z digest=sha256:36c80992b6f82523d0c060093221fcd708c8b26132e1ae2e7d827372d4d252b4

Observation fe17b1a6-13ed-47b1-8643-697bde5ccf45 · outbound

This paper cites Layer Normalization.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Layer Normalization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.162480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.162480Z digest=sha256:900a54db1b6daf9289e198d1bbeb353bd4b2a5aaf5ad25838e820ebcfdbaadad

Observation 978ceb9b-bade-4c84-8886-774fdcd1944e · outbound

This paper cites Temporal Sentence Grounding in Videos: A Survey and Future Directions.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Temporal Sentence Grounding in Videos: A Survey and Future Directions

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:05:48.364200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.169189Z digest=sha256:b183d8dcd52862c312129434247b40d081f5c9f57246dcf09dfdaea3ed329282

Observation e48aeb59-c1af-459e-8ac0-3ffd49040fc8 · outbound

This paper cites Attention is all you need,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Attention is all you need,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.654863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.179780Z digest=sha256:9c4cca43bfd2c3843e29d70f5b1cd972b3a87a459631deb50579c357b7d2a51a

Observation c1646409-cf5d-4aca-9303-6a1b637f9e74 · outbound

This paper cites RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.635334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.186680Z digest=sha256:1d3d281b22e684c7fec386fcf7621093bb14a7354b60c17578f2deeb9792a8ab

Observation b48827d7-79ec-4974-b1be-47e0abd404e2 · outbound

This paper cites Tall: Temporal activity localization via language query,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Tall: Temporal activity localization via language query,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.604469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.192444Z digest=sha256:9d8e0649cf39849f800ebb944d76487d4c61040d56719b3337d06367c7b8fdc1

Observation 324d2ccf-197e-459a-be21-cfa0132c1e6e · outbound

This paper cites Decoupled Weight Decay Regularization.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Decoupled Weight Decay Regularization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.198444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.198444Z digest=sha256:e9f54da975b07ab4759ba25a7d17fa071d9ca0388775e7fe7a295708ddcadfee

Observation 03d7fa82-70c5-4433-a111-b76be3f778ac · outbound

This paper cites Self-Chained Image-Language Model for Video Localization and Question Answering.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.204282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.204282Z digest=sha256:d9411409cda36485523d3f3c53b688570f680659e41ace9f76d2f145282777ce

Observation 8d26cd0a-6015-493c-a3e8-088fb680396d · outbound

This paper cites Learning 2d temporal adjacent networks for moment localization with natural language,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Learning 2d temporal adjacent networks for moment localization with natural language,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.577349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:05:48.211356Z digest=sha256:e2097c0cd37e45bc20fff023cdd4af54fed3c4677fa8cb879b684b49ce71014f

Pith citing papers

No inbound Pith citation observations are available.