Pith. sign in

Paper Citation Record · LEDGER

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection

As of 12 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2501.10692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10692 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:05:48.211356Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact3
  • verified fuzzy22
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db8af9e1-8c39-4713-8d3c-e623fa29db71 · outbound

This paper cites an unresolved cited work.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:05:49.166365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:47.988030Z digest=sha256:a9ab769b79b197c226f8a3cb3fda57e56945f39e96c474642283362cb28f408f

Observation 643381ef-c2b0-4fa5-9f2d-59eb4ecdd82c · outbound

This paper cites an unresolved cited work.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:05:49.136619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:47.994781Z digest=sha256:4aab79696756a47f3582b0502ae9451585ce94c4cba2b3ec5816ad2d57c8639d

Observation c43ec5cb-6aac-43bb-ba86-362b02fd13de · outbound

This paper cites 2 Related work Most previous MR&HD approaches [5, 11, 12] only em- ploy image and text inputs.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection 2 Related work Most previous MR&HD approaches [5, 11, 12] only em- ploy image and text inputs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.115294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.001085Z digest=sha256:770d7b41f7425764007cce29312686ff11e8e8e604d9d3a3e8677d7a047525e7

Observation a2afd7b4-877e-4000-bd7f-f1507ba26e52 · outbound

This paper cites Effectiveness of each module in MRNet on QVHigh- lights val split.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Effectiveness of each module in MRNet on QVHigh- lights val split

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.087752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.008082Z digest=sha256:c10079e2bf28c2b9f2c4144110b4022b9168063562c8935aa7d3f51c43b0cb17

Observation a007fb0a-100a-4d32-971e-d10c7194861e · outbound

This paper cites The main reason is that Moment-DETR only utilizes RGB, which fails to fully under- Table 5.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection The main reason is that Moment-DETR only utilizes RGB, which fails to fully under- Table 5

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.054624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.015150Z digest=sha256:2866ce7b83d5e955d5826b3370a2f7dbd7e38ded9c6de4a9887932519339e29b

Observation 2a1b707c-b3d1-49a0-b54c-cb00fc211b4e · outbound

This paper cites Localizing moments in video with natural lan- guage,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Localizing moments in video with natural lan- guage,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.034835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.020785Z digest=sha256:554127bc5479dd1cf3cd15bc39bd34371c673fef8e5f8198712ac84361afe748

Observation c00b4893-4ba6-4451-b5f7-87942f4596af · outbound

This paper cites Less is more: Learning highlight detection from video duration,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Less is more: Learning highlight detection from video duration,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:49.013948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.026677Z digest=sha256:de7250dfe1d7def8d5e7b36377318b6c163b5d600b7bb97ea4377d1a1a512502

Observation e4951687-383d-45cd-8e87-f5c8c75176d6 · outbound

This paper cites Detect- ing Moments and Highlights in Videos via Natural Lan- guage Queries,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Detect- ing Moments and Highlights in Videos via Natural Lan- guage Queries,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.990075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.033453Z digest=sha256:e855fc5cd1c5b39b70d63af21c906096ccefb7eb5d6dc648cabcba13c8382ea6

Observation 5c3af061-9e22-46ab-bc1b-32290657fc3e · outbound

This paper cites UMT: Unified Multi-modal Transformers for Joint Video Moment Re- trieval and Highlight Detection,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection UMT: Unified Multi-modal Transformers for Joint Video Moment Re- trieval and Highlight Detection,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.959393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.038646Z digest=sha256:338f603e7992c5bb910713ff5eec5ef9c863b7ad3308efffbe4f6352ae889562

Observation 9c8c0675-fbc7-4842-ab7a-46be91c2877a · outbound

This paper cites MomentDiff: Generative Video Moment Retrieval from Random to Real.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection MomentDiff: Generative Video Moment Retrieval from Random to Real

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:05:48.545907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.046112Z digest=sha256:97c1d48f967439f6d9eff3dccfe677328c654e76b9658c61b132543d85ea1620

Observation 8c05860d-5739-4c27-9f29-94bb911e5446 · outbound

This paper cites Gmflow: Learning optical flow via global matching,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Gmflow: Learning optical flow via global matching,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.934703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.061247Z digest=sha256:0c7c7bc11638a1fa0005f7ddbef1ba22746494bc2ea95e37cc85f5f727108c00

Observation 0e1917f8-3afa-4cdc-8c29-bfc7eb45f657 · outbound

This paper cites Depth- cooperated trimodal network for video salient object de- tection,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Depth- cooperated trimodal network for video salient object de- tection,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.914079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.071804Z digest=sha256:4f790a427bf08cd7bfca0effa4912cf31ad08a744e864bf45db156d73e2562a9

Observation 9fb4ab9b-c1aa-4395-a6fc-23887b50d271 · outbound

This paper cites Pyramid Feature Attention Network for Monocular Depth Pre- diction,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Pyramid Feature Attention Network for Monocular Depth Pre- diction,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.888167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.084100Z digest=sha256:2b8f7fb0c796f97dc34a28df90235004c21565cbbfc4680ab78380cc3700075d

Observation 29f57026-dfaf-4e31-9962-5bfcc611925f · outbound

This paper cites How hierarchical is language use?,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection How hierarchical is language use?,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.864654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.091445Z digest=sha256:048c4f46899c06e0bd25cb193847eee76be66e78a02408d075cef2499154bc94

Observation 1abb676d-9dce-4296-bce8-ef2645e45d9b · outbound

This paper cites The emergence of hierarchical structure in human language,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection The emergence of hierarchical structure in human language,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.843168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.097892Z digest=sha256:e09d9fb7d28970be2c5a68592a357f94855ae8f0075efa16fecedc4dd6ff24e4

Observation eb7e6f6a-cb38-494f-8377-fc6072587a61 · outbound

This paper cites GPTSee: Enhancing moment retrieval and highlight detection via description-based similarity features,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection GPTSee: Enhancing moment retrieval and highlight detection via description-based similarity features,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.812275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.106113Z digest=sha256:0852966a7e746c6cbc812e6d65b2496bdef29fcad143fef032126b3c6b937e29

Observation 7f5b4600-e35f-4a04-85b9-265ac4897fc6 · outbound

This paper cites MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.113507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.113507Z digest=sha256:f1d462cc4ae5306e3a69128683f310bad923c5e9bae69b6532ca7d40e57e9c61

Observation 008f428b-8dcf-4a60-9f1b-351eb1da1fa2 · outbound

This paper cites An empirical study of end-to-end video-language transformers with masked visual modeling,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection An empirical study of end-to-end video-language transformers with masked visual modeling,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.781442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.119107Z digest=sha256:0a6b36cca21b5b6fdcaf6c4540c83a79ddcfc185b7dbb9b0bc382e39cb158878

Observation db86f254-899d-4769-902c-65d7b044870c · outbound

This paper cites End-to-end object detection with transformers,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection End-to-end object detection with transformers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.758557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.124164Z digest=sha256:bfd43d04753188cd3b040f8b48dd40890c3820ac68290b1287c239cd8f1b376d

Observation d51e2849-68c4-4d12-b15e-08fc9156fa70 · outbound

This paper cites Shifting more attention to visual backbone: Query-modulated re- finement networks for end-to-end visual grounding,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Shifting more attention to visual backbone: Query-modulated re- finement networks for end-to-end visual grounding,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.732109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.131702Z digest=sha256:82ad0b61d7e731b53aea1ac23c6a74ea71a7e2e128b3317313891f42a9529015

Observation e3fd3cd1-dd33-4563-a411-8c65390926cd · outbound

This paper cites Re- thinking transformer-based set prediction for object de- tection,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Re- thinking transformer-based set prediction for object de- tection,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.710933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.138847Z digest=sha256:9daffc5f7bd584af2bfb23e612b22d39586370391c3fdbc11540d6875f8844cd

Observation 539371cd-619f-48a9-86e5-08a1ae89813d · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.144753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.144753Z digest=sha256:3ed39cb692b14cd11fcfba7589ce8758857e87f18856ac39d9eb21471dc30bd1

Observation 85873ffe-579d-4028-88f7-4c0ea041b52d · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Learning transferable visual models from natural lan- guage supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.685286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.150146Z digest=sha256:ec786b909ddddff1f70da4454bf77ddadf1771760c6e3c1736cb9a9626cc997c

Observation ef50378e-ec48-4585-a08d-3416706516b6 · outbound

This paper cites Recognizing American Sign Language Manual Signs from RGB-D Videos.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Recognizing American Sign Language Manual Signs from RGB-D Videos

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:05:48.435257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.157147Z digest=sha256:b66777158152da4494d0aab5d7585b74386cf619161b9a972e4c661c8c74c09f

Observation fe17b1a6-13ed-47b1-8643-697bde5ccf45 · outbound

This paper cites Layer Normalization.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Layer Normalization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.162480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.162480Z digest=sha256:900a54db1b6daf9289e198d1bbeb353bd4b2a5aaf5ad25838e820ebcfdbaadad

Observation 978ceb9b-bade-4c84-8886-774fdcd1944e · outbound

This paper cites Temporal Sentence Grounding in Videos: A Survey and Future Directions.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Temporal Sentence Grounding in Videos: A Survey and Future Directions

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:05:48.364200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.169189Z digest=sha256:71dc737519d3fc08695898c7ccb8510624de59a2a8faf0fa91003e02a3411f5e

Observation e48aeb59-c1af-459e-8ac0-3ffd49040fc8 · outbound

This paper cites Attention is all you need,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Attention is all you need,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.654863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.179780Z digest=sha256:7bbfaa3df51437cf09ecf7de4086df96967215c01e5114ba36d1ecf6266c3c8f

Observation c1646409-cf5d-4aca-9303-6a1b637f9e74 · outbound

This paper cites RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.635334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.186680Z digest=sha256:7b3ab1f9dc53a37b7bd24873bcfc74dfbda8f77215353d910648ab8153f234bb

Observation b48827d7-79ec-4974-b1be-47e0abd404e2 · outbound

This paper cites Tall: Temporal activity localization via language query,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Tall: Temporal activity localization via language query,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.604469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.192444Z digest=sha256:e64f82ffd708354fb9e9385a78db0b13be1bdc3c1edc7d4bc93643c830e532a8

Observation 324d2ccf-197e-459a-be21-cfa0132c1e6e · outbound

This paper cites Decoupled Weight Decay Regularization.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Decoupled Weight Decay Regularization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.198444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.198444Z digest=sha256:e9f54da975b07ab4759ba25a7d17fa071d9ca0388775e7fe7a295708ddcadfee

Observation 03d7fa82-70c5-4433-a111-b76be3f778ac · outbound

This paper cites Self-Chained Image-Language Model for Video Localization and Question Answering.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:48.204282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:48.204282Z digest=sha256:d9411409cda36485523d3f3c53b688570f680659e41ace9f76d2f145282777ce

Observation 8d26cd0a-6015-493c-a3e8-088fb680396d · outbound

This paper cites Learning 2d temporal adjacent networks for moment localization with natural language,.

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection Learning 2d temporal adjacent networks for moment localization with natural language,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:05:48.577349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T19:05:48.211356Z digest=sha256:8cb0eb832e8b7e5cf452c7e3cf6ab7a7739538c58974fa5dc04452f7cd08b6da

Pith citing papers

No inbound Pith citation observations are available.