Pith. sign in

Paper Citation Record · LEDGER

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding

As of 6 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2604.25886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.25886 v3

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T08:41:31.700367Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact16
  • verified fuzzy44
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 140600bd-b3cc-4483-88b0-50acf0d6d1a7 · outbound

This paper cites Univtg: Towards unified video-language temporal grounding.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Univtg: Towards unified video-language temporal grounding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.930723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:bf0a206efc9d8445bbcf134dad730b94897c93835c47584438321f144ac950a0

Observation 641867e7-b10f-4d3c-85b0-b29f92d5e7ed · outbound

This paper cites Context-aware biaffine localizing network for temporal sentence grounding.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Context-aware biaffine localizing network for temporal sentence grounding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.928956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:953a93044c5eaed4ce7cad3425f2022c466dcdf350e1ac7531a61af7ffb3005e

Observation dcf9ca09-3fc6-4e26-8af9-3d6855a4d2f3 · outbound

This paper cites Internvl: Scaling up vision foun- dation models and aligning for generic visual-linguistic tasks.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Internvl: Scaling up vision foun- dation models and aligning for generic visual-linguistic tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.889678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:56870a7f44bd10a6015833fd1033ff7299366be65033754320856f0d0b2878c1

Observation c3f5d603-35ae-48d7-a3ee-45be3ec592d0 · outbound

This paper cites Chatvtg: Video temporal grounding via chat with video dialogue large language models.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Chatvtg: Video temporal grounding via chat with video dialogue large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.891842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:1c260332cbb18438dda88a23794c941b5f4515e43842d98f7e72512b3dfe2fba

Observation c33fb6a7-93b8-410b-ab79-b58bbd373e39 · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.897136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:98a49ac0bbf40cc9250cdf97485faa8613d55aec2fcbd8c4e192515cd75afbbb

Observation f81d4d1d-90c9-4b7f-b36d-96c878701b03 · outbound

This paper cites TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.796582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:c3b8f9de2959422e77cd74f72906536170b0c875aff3a5b7cca3b8d5a9ed4377

Observation c36caa39-d471-4c62-b1ff-4fbdbf1e6be9 · outbound

This paper cites On the consistency of video large language models in temporal comprehension.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding On the consistency of video large language models in temporal comprehension

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.900672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:1af00496984990bfab53da5a297ab9dd409164f1619a0f642109fe9ad64af2f2

Observation 21bb46de-5cb1-4227-bbed-d1a16fe0d7e2 · outbound

This paper cites Bridging the gap: A unified video comprehension framework for moment retrieval and highlight detection.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Bridging the gap: A unified video comprehension framework for moment retrieval and highlight detection

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.907559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:b48c420913266bf52633c8518b4d1f0e277c2fed81637884351bb6fb831a5b44

Observation a90dd20a-34cf-41a0-a0e1-d0642cc72b6e · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.921714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:a809400e1c9774c7fc57705404b0371501a4e4950c650d97a926ae66eccfdea4

Observation c2e25f4e-d662-4fb6-9e07-2e55f82db39b · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding LLaVA-OneVision: Easy Visual Task Transfer

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:45:34.799308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:98e7c983fb712ad1c3caaed3c9b243bbfb14adfeb62d804b042420a5c466c312

Observation d61d4b4f-e110-4244-8274-137c203bae06 · outbound

This paper cites Zero-shot video moment retrieval via off-the-shelf multimodal large language models.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Zero-shot video moment retrieval via off-the-shelf multimodal large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.914505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:8d34e6a3c35263b858123aebdfa4302474b1f374502810715481e39ba1638c74

Observation 8032a6fe-45fa-4386-82b3-ab100cc10788 · outbound

This paper cites Background-aware moment detection for video moment retrieval.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Background-aware moment detection for video moment retrieval

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.916373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:36002ecbc85cd4046b134485a6c5ca86e1d32a6585cad4a84acaff2a9019163c

Observation 965e3597-bd40-45ac-a1c0-7bcbec4c4cac · outbound

This paper cites Dense video captioning: A survey of techniques, datasets and evaluation protocols.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Dense video captioning: A survey of techniques, datasets and evaluation protocols

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.919893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:3734cf260db54fbd2251bee8f35c50505fadcab2d82bb9d3a3fbfbead513fe93

Observation f8718ead-6a14-44fa-9856-a073aa767d69 · outbound

This paper cites Dense video captioning using unsupervised semantic information.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Dense video captioning using unsupervised semantic information

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.918001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:61bea7bed64c491a011630a52d3e63b3865b80242ea67c1dc52d806d44c283c5

Observation a26210ed-34d9-437f-a76d-b29c279b15a3 · outbound

This paper cites Lvd-2m: A long-take video dataset with temporally dense captions.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Lvd-2m: A long-take video dataset with temporally dense captions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.925196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:6337197d8d980d45c0e4a1c7f5685d9c1c8be5479a19364133478668b2bf4a20

Observation eba31263-66a4-4c5a-83fa-12ba049f3164 · outbound

This paper cites Unsupervised video highlight detection by learning from audio and visual recurrence.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Unsupervised video highlight detection by learning from audio and visual recurrence

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.909388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:fae316c7f21ac3d586d827166b06334e863e42d5bc3b0ec780ea38b0a2e3d7a2

Observation f1752f48-a684-4897-a918-b42a02f1af8d · outbound

This paper cites Less is more: Learning highlight detection from video duration.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Less is more: Learning highlight detection from video duration

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.893588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:37a8bbcd076dc2675cad723400dc0c4059b391724a34eba4bd75846d18b6f4d2

Observation f69f8d2e-543c-4534-9bb4-8d5d1e5d34fa · outbound

This paper cites Contrastive learn- ing for unsupervised video highlight detection.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Contrastive learn- ing for unsupervised video highlight detection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.912767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:5e71c749e41de64e1030c4f656ea43d020b86456e6fae5e82af703374bc4572f

Observation 6815b9c0-d40d-49b5-9b76-deabc2207972 · outbound

This paper cites Query-dependent video representation for moment retrieval and highlight detection.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Query-dependent video representation for moment retrieval and highlight detection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.850890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:a91d74bd12fc26f1294e9fe2a3ded0af58361293e90ee4f3928ab63aae62633e

Observation 97f42890-b4db-45fd-936a-b333dc3cdd4d · outbound

This paper cites Tvqa+: Spatio-temporal grounding for video question answering.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Tvqa+: Spatio-temporal grounding for video question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.860159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:06c686278a7e83563753cb3ea7bda225670bb0944db2fd5fdf9572436fc27f7d

Observation 7610431d-2804-432a-ba4d-cfd0c90a941d · outbound

This paper cites Timecraft: Navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Timecraft: Navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.878740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:8a7b1b1608cd752e3069bcdf0b7ab274c50bfb920e208b964b2c7451969548f6

Observation 063e4e36-3f3d-4609-889c-187dfc7381af · outbound

This paper cites Grounded question-answering in long egocentric videos.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Grounded question-answering in long egocentric videos

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.882280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:9f0e8feb188a00ac088ab59b340608d32dfb5829f0b0e8575efb5f454fce4819

Observation 7dc83e2b-e036-4bdf-99a3-5ca86ee44eb5 · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Can i trust your answer? visually grounded video question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.852676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:56eccb1eca253f233f8df630275f13cc43d2b0a5c2a8efe9d66fab0044f5fa06

Observation 5be6577f-0360-4df6-bcda-85ac9fc5504d · outbound

This paper cites A survey on video temporal grounding with multimodal large lan- guage model.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding A survey on video temporal grounding with multimodal large lan- guage model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.854562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:5f0070d11c4177aaa0656c8be3cda2700c40b61f93cdad0dbc2d8e6c88d67691

Observation d1ffa588-fc59-40b6-9d71-a39139421e77 · outbound

This paper cites TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:45:34.785938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:c07077442150a2629862ae344d21879934aba56f0e873ed6069445c6be937d16

Observation d6a29860-7a1a-422e-9de5-4fb838b158e3 · outbound

This paper cites Alvarez, Lei Zhang, and Zhiding Yu.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Alvarez, Lei Zhang, and Zhiding Yu

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:45:34.751361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:8f086e4015679e0e3d3d4c184d4fae41d0300e45a9166c4478b7412a2f367936

Observation 90111710-33c5-404b-9afb-80e8eb15db3c · outbound

This paper cites Towards visual-prompt temporal answer grounding in instructional video.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Towards visual-prompt temporal answer grounding in instructional video

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.874971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:1ebbdb89eb76a785af8446981c2f077559d9652f6e920ed930390d2d0ca072b7

Observation 4ffc359c-2fbf-4611-a7e5-00a4a5286d97 · outbound

This paper cites Number it: Temporal grounding videos like flipping manga.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Number it: Temporal grounding videos like flipping manga

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.880594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:e79c7e4870a96a0bedbbe79c53ea44d2f4b007bbd94332d46b55b8122e77522f

Observation 03fad860-5072-48ff-bc0f-6a89113d0702 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Vtimellm: Empower llm to grasp video moments

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.887923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:d4ee4711a1a0d817bf803b1403a3146a11e5e84a1e4f691f598d59b760aba707

Observation 54bd5be4-ee64-4b75-afce-bad464169b8b · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.898954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:50989ffa647a26251b73255f9f6977a7e39c5e61471c5e45d617ee05c6cccd9c

Observation 37509041-2377-4856-9312-547eb491ec15 · outbound

This paper cites Training-free video temporal grounding using large-scale pre-trained models.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Training-free video temporal grounding using large-scale pre-trained models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.904056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:5642248baf4d8ecd63bc556973fcfd5c081f71cfa8f7d4f31ab1f382217adb36

Observation 055d4cd6-6130-414a-b394-1959201a0180 · outbound

This paper cites TRACE: Temporal Grounding Video LLM via Causal Event Modeling.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.789428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:a0cd698152e95d8f3d03c9f3dad790b0ffbfb9d29d7f2f602b7ab0b2b1990949

Observation 439a5904-bd81-4259-9ab6-974e0356269a · outbound

This paper cites Omni-rgpt: Unifying image and video region-level understanding via token marks.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Omni-rgpt: Unifying image and video region-level understanding via token marks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.927175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:c1c29e5bef15940570c3f23a81fe7bb51c57c36221c122dada795b9b37f90f82

Observation 67edd1eb-495a-4cd3-8238-ec895844a8ee · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:45:34.779591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:b19f5150b71c1e39187b3936371288fb6807b1d8e3a6b467a103b5580c690e73

Observation 26f41eed-5359-4dcb-a9f4-e9fd5ac1600b · outbound

This paper cites Sa2va- i: Improving sa2va results with consistent training and inference.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Sa2va- i: Improving sa2va results with consistent training and inference

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.749187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:5f1da0a4fa1b67cf70da35d6086299811ebba3a4e8f086a93868f8587360b4ab

Observation 1e33ad75-ace6-4194-9d3c-7c060947b04c · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual grounding in videos.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Videoglamm: A large multimodal model for pixel-level visual grounding in videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.862249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:b1cfd0df6f702e9fa3680594f51c83f9824b44caedec0594f8243c248a3b4ed8

Observation 7bbc80e7-0eeb-43aa-aea5-a9a56bb47d97 · outbound

This paper cites Videorefer suite: Advancing spatial- temporal object understanding with video llm.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Videorefer suite: Advancing spatial- temporal object understanding with video llm

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.871711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:f56b555d505fdb65e2472a5e313d8c732a54e6e951b437910881b4e2a50ab1c0

Observation 0d8e9173-957b-45ba-a782-b4599bbda1dc · outbound

This paper cites Vip-llava: Making large multimodal models understand arbitrary visual prompts.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Vip-llava: Making large multimodal models understand arbitrary visual prompts

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.866093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:118c151ab37c945789fcd9830d241d06cef36b734dd91772c64b9098d30b5e5e

Observation 4c7421cf-7e64-4f08-9a93-2e0f1954ba18 · outbound

This paper cites Generalized decoding for pixel, image, and language.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Generalized decoding for pixel, image, and language

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.867884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:2f83fb1922d4fa1973620af5d231ec52efc6870f3d31163ad9a7bd2a5c42e169

Observation 53a9a5f1-e08c-4ec3-a4d5-e6e3ce76512e · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Gsva: Generalized segmentation via multimodal large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.905750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:568decaebc8d51d61cfea66a7af3fe6dd72bdbdc483e136ff0caf8f4903490f6

Observation 23edf9ce-eec1-4996-b474-7cdcb2dfd904 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Glamm: Pixel grounding large multimodal model

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.911046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:4b1eccacfd9f094c2cf0a3d40b84e6f2e605ef6f996b07a7f3267d375ee5f46d

Observation 63c51d81-258c-414b-812d-20b1ea2b92f4 · outbound

This paper cites Lisa: Rea- soning segmentation via large language model.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Lisa: Rea- soning segmentation via large language model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.873349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:83893e7e398b85e5378fcadc0c64e45ecc3696ab335ddf1882a7288359bee561

Observation 575d8c05-c7f7-42c1-a20a-3df06e7a08c6 · outbound

This paper cites CoLLaVO: Crayon Large Language and Vision mOdel.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding CoLLaVO: Crayon Large Language and Vision mOdel

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.776035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:6b8ebb81858a57753ae34cf5e51f42f3ab5bde2b53be1cc6afe750e8ec54b50a

Observation f9e505eb-d405-4ced-8f5e-c157b04b840b · outbound

This paper cites Segment anything.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Segment anything

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.876630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:d39516626caac5f50caf006c560166da28f2d2f18eedbf285b7acdae6c0795cc

Observation 1dfbf54c-991e-4217-83b0-d42f4c33c253 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding SAM 2: Segment Anything in Images and Videos

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:45:34.779470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:602d10fd7ac63dc3b95177473faf57dafe07ed5745729159f2200ede1700d7c5

Observation 8ae83314-bd9e-4409-bae3-9e23baddcca8 · outbound

This paper cites GroundingGPT: Language enhanced multi-modal grounding model.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding GroundingGPT: Language enhanced multi-modal grounding model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.884170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:0e9491652f9330680f8065ad68462c4a58fc0a1c2809758267c7041766d77bb0

Observation 5b05204d-7ba7-46d1-8390-b6a87e3d9a5e · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding LITA: Language Instructed Temporal-Localization Assistant

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.782507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:30cce56438f2daee5d54ca6455f0cc328ca5ae3baaddf2a4ea4030c1ab9594ee

Observation 4ce1d25a-5e4d-49f4-b27a-7c775f745683 · outbound

This paper cites VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.786182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:de1c7b5d1e029cc811e20bfeb7789d5a0df4b67cfd3489d2c0a4c4bf1e150f37

Observation 53319488-39a1-4251-bfa8-9a8cf9e6612a · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.885931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:cb36da17973e6e3335e4ad3d33803d546f08375ba600fee6191fedeb57c82d86

Observation 7ba36f90-8d5b-4224-860b-09a8bac66510 · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:45:34.755963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:9c5343c7c90b928919bbe442be06e07aa42e79bf897092773b3007aa09c33c9f

Observation b8b842b8-d96b-4015-8c38-0b713131a6bb · outbound

This paper cites VTimeLLM: Empower LLM to Grasp Video Moments.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding VTimeLLM: Empower LLM to Grasp Video Moments

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.768754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:5efe43c1676b6bf007af0a3d2fc3439584f29ef2825f05d94083f7ed14bf5a43

Observation 39369fc2-1877-4b2d-8361-27fbb69b7a60 · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.782509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:c7c0121ce5c5ecdaf4db276da405168f175a0709f50c7b71ecdae9c669a7ef7c

Observation 34063787-d5e0-4c4b-a715-e96d5de010e5 · outbound

This paper cites Hawkeye: Training video-text llms for grounding text in videos.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Hawkeye: Training video-text llms for grounding text in videos

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.895432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:2ad2e3f210e8609f9a378fbcbecb967390b7cd5340c5ca2ecc0877baa0e2de72

Observation d46e3d6c-f44f-496b-9a3d-863043f454d6 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.762438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:270663e0a13a6a3f0d2a881915981aa2e69dcacae60398ef439dfddd5afebd61

Observation b7feabb6-0a7b-4090-9f69-22bb55cfd53c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:45:34.768697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:7c83e501931f1320dd07c5674ef6a7f36525482b3e935b9606f0f8a2983d4379

Observation 94957e16-60aa-46c2-8b6f-105f9b7a6672 · outbound

This paper cites Tall: Temporal activity local- ization via language query.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Tall: Temporal activity local- ization via language query

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.902357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:53f32b170dc70538713f1a7f10ff948f67a02a3512113403c7a335850ef85cdb

Observation 2453830b-a0cd-45fa-af24-12106c294f30 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity under- standing.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Activitynet: A large-scale video benchmark for human activity under- standing

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.923364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:53d404ebf041d4f5ab971ed9c881853d8aedad54a2f2792702b450d942f732ed

Observation 07afdd43-58a4-41f8-bc00-833bafa767f4 · outbound

This paper cites Detecting moments and highlights in videos via natural language queries.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Detecting moments and highlights in videos via natural language queries

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.856383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:9fa99a4d7cb5d5df08efeecb3d7810ea8837f244c5d43c8276ad6aa8b02f0d61

Observation 441a83b3-bcfa-43fe-9b7c-3c71bebbe125 · outbound

This paper cites Long Context Transfer from Language to Vision.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Long Context Transfer from Language to Vision

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:45:34.772580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:8ab6b9f15e8d4111a4f8ba6ca96f372879f4edfaaeac5a5d3115ee4ea2e3c14d

Observation a101b8b2-1b9b-4349-a689-88ed2a3ae3b1 · outbound

This paper cites Yoloe: Real-time seeing anything.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Yoloe: Real-time seeing anything

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:34.789794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:3979565d618164f57a240ca8a643a812a6cc1b96d7c3f817ba6efaac75cb65ab

Observation 256f25d8-18cb-442b-8ccc-d94f234e76c4 · outbound

This paper cites Two men both dressed in athletic gear are standing and talking in an indoor weight lifting gym filled with other equipment.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Two men both dressed in athletic gear are standing and talking in an indoor weight lifting gym filled with other equipment

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.858395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:52d410014aef24b28fed6e6a32b412cf7a604942387eec2a4be086942f619c66

Observation 169c417a-584e-4927-91a6-1a0b8ec978e7 · outbound

This paper cites Overall trends are consistent with those observed on ActivityNet.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Overall trends are consistent with those observed on ActivityNet

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T12:52:29.864028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:9527952635c82c43dd8f302dc298b1c4f66d8c51750145320753a5708d6c2496

Observation 226ef46f-8d68-447e-b4b3-2ec24e30e022 · outbound

This paper cites Table 6 reports the effect of different color parameteriza- tions on ActivityNet.

MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Table 6 reports the effect of different color parameteriza- tions on ActivityNet

Reference 64

Resolution
malformed identifier
raw_fallback, observed 2026-07-06T12:52:29.869944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:41:31.700367Z digest=sha256:78800bf4fe80960b87289962a724496b2da0af88ff08c78147e8e42b9eb465a6

Pith citing papers

No inbound Pith citation observations are available.