Pith. sign in

Paper Citation Record · LEDGER

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding

As of 5 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 0 inbound Pith citation observations for arXiv:2512.06673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.06673 v2

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T00:54:53.789523Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

86 of 86 outbound references displayed

  • verified exact28
  • verified fuzzy57
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d39bed59-f4bd-46df-b1b2-9bcfe84c731f · outbound

This paper cites Localizing mo- ments in video with natural language.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Localizing mo- ments in video with natural language

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.320293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:ce5f915f9050c18a0184024fb6afe619cd322c01286adc7a8355af395454e7ae

Observation 14512fa4-c8ed-4176-b7b8-789f16e95ff5 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Lawrence Zitnick, and Devi Parikh

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.413456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:60b145da882fd23f33edaeffb572936f775d9bc87d97891fec0faab74b4c4df4

Observation ae986ee4-75aa-4e6f-ab89-be88cdb5fd34 · outbound

This paper cites Qwen2.5-VL Technical Report.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.564125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:9e2f5e9bdd0f2d7b3eb8eb29d596f3618e2b89f1ca02bddb8de18ff00b48817e

Observation 3c8621c7-08a1-48c5-8f1a-20f15b7302f8 · outbound

This paper cites One token to seg them all: Language instructed reasoning seg- mentation in videos.Advances in Neural Information Pro- cessing Systems (NeurIPS), 37:6833–6859.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding One token to seg them all: Language instructed reasoning seg- mentation in videos.Advances in Neural Information Pro- cessing Systems (NeurIPS), 37:6833–6859

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.352926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:918e30c19ddfad041fb83bdfeb94d0e218ad8f4777f7dbe511f3b5d152138252

Observation 81260634-4bc2-4b18-819a-4d363323e5ca · outbound

This paper cites End-to- end object detection with transformers.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding End-to- end object detection with transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.383759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:cde7f64d62a27bfe4816b4afb210e658cecc4d14815ccba41eacdfbf2a135b06

Observation 17b636cd-f3b3-4751-870c-f9954cd11183 · outbound

This paper cites Chen and William B.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Chen and William B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.368742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:136dbadf411275cf6bac39fac8087946a13fa925b380cfae988188ed96212157

Observation 91bdf40b-93f9-4bc6-8bdb-ad0bc3fe7364 · outbound

This paper cites TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.578489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:efb5dd9d81e0ec6f132265fbdae41ff588ad936badd94dd757cdaeeaf5500322

Observation 559cd5a6-674d-4b9a-85c5-f993e80f5adc · outbound

This paper cites Sutrack: Towards simple and unified single object tracking.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Sutrack: Towards simple and unified single object tracking

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.325412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:2947fe93b732478f3604378187baf541fcb85856ecff8734d792f4aad52f31b7

Observation 1b18494c-ce0e-40c8-a5a4-ed71a1834bc6 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.571484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:091e070d6ba7d30ed1b6f83220ad5edf0eee0ed55ac2136f7894144a3f0614ac

Observation 478d5077-302c-4dfc-bc15-5b34516c64e8 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.567418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:7ff4365059943de85c92e4d5a315c07f475ee0351ec5b5effef3b3c8e7d406db

Observation 05787ab9-47af-4202-a225-2e5012a2b0f3 · outbound

This paper cites V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.542264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:7443ff77a491d2b012d4a2b3d021c6277c7ef18805f5ad6ef8b7ea92900de1f3

Observation ea1b4a1f-8dee-4cb9-af8a-f60352e59c7a · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.529425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:ee3999643796dea54e179c93a9b8230874da64438f61722ee58f53f619fbae4e

Observation 605e6d0c-d22d-473b-81fc-5adda7a1f7b7 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.589332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:77282093b84a2e90a854c7b774eebaa91fe12f29dc28104667d161f7f1bc9803

Observation 04a5bd31-04a1-4392-ad73-e48d61960717 · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.551714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:5b9016e665e04c36f02abaa4ea62680af36909bcae99b9ee093d945b7991a48d

Observation c70e0ede-58dc-471e-b6dc-af7f568b3217 · outbound

This paper cites Tall: Temporal activity localization via language query.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Tall: Temporal activity localization via language query

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.407793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:68057789e120fc5ca2e65309fea8f44cc0c6a473566d2802aaf254d7c9da2511

Observation d92ec788-1196-4346-9dd3-47634f82f031 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.383486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:be1f202195b0bc2f97b38917fe653b075ff5047d8fa43c1a555e29289b737332

Observation dda6359b-8f55-455f-bebd-98c639fb5b2a · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Ego4d: Around the world in 3,000 hours of egocentric video

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.397915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:68bfd57adcb8870da221caf6e1cfdecdad4d5f6da239b662da5da110107719cd

Observation b2deb3ce-b3e1-4338-b375-6b3edea312e9 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.376375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:0254eef9a71617081e1de56a2d575b305e99f85ad32e439fbf8e73dd672ae959

Observation b8ba7102-3c03-4912-bf30-b8486ea8f632 · outbound

This paper cites Context-guided spatio-temporal video grounding.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Context-guided spatio-temporal video grounding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.401093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:34d26caad5547b983f64058ef357da615f4282fe180613826ff22a7d0f9ff128

Observation 62678143-30fe-4a46-942c-f04e3065bf57 · outbound

This paper cites Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.560844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:3a287c0a097a39ea47214a3bbf67dbd50c789c46b6bd3f5b7c9955296a48cfbc

Observation 9f7a8c12-6772-493b-afab-5c4d15ca4c88 · outbound

This paper cites TRACE: Temporal Grounding Video LLM via Causal Event Modeling.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.582114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:c2b348d6ad2f8c3433f7ee621ca4c12fc7d34fb27748e3e37df5e9b75c204a95

Observation 9cfe936b-f2db-4cb4-9de6-098390cdaf1e · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.372130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:073e3b8c82237727c3d7dba503ed5ca9b6d03ae4093a44af203d331f60392b81

Observation fa5dab23-0308-4fde-b04b-3d7f4835fb54 · outbound

This paper cites Creating summaries from user videos.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Creating summaries from user videos

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.410687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:bd7ad2401979b80e6ba2300517b779df49f94d2f2344a61bb1c4da3e3ae1d3fb

Observation f63be08b-25b4-4032-b6ef-d97e7fbbaa0c · outbound

This paper cites Lora: Low-rank adaptation of large language models.Inter- national Conference on Learning Representations (ICLR), 1 (2):3.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Lora: Low-rank adaptation of large language models.Inter- national Conference on Learning Representations (ICLR), 1 (2):3

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.342059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:a3abb348e0165f93cd1fc32548b3aed76205076389c6a616517ad0fe35105d1f

Observation cb000905-2dd9-4bda-ac22-47530f9a5e60 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Vtimellm: Empower llm to grasp video moments

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.338148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:e3debc628e3863438a9817f2d9fa6ca6c620383a5b09c38204b7c612e1dae757

Observation e31a8702-54d6-4968-be8d-f0a0d2829b17 · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Lita: Language instructed temporal-localization assistant

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.330180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:b3f3ce4ccc458910f952b59e22620a270d003aefe3268e4ec85ec73f41a57ce2

Observation 289f36d7-8a6d-40d6-9866-7730ef4ab9a7 · outbound

This paper cites GPT-4o System Card.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding GPT-4o System Card

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.585131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:915130572a41155ec0ddac9cbd980a1cdac46bc98476c6b86f784cf03538dfae

Observation 9c37f211-c121-497d-be4f-965d2a61ffbb · outbound

This paper cites Embracing con- sistency: A one-stage approach for spatio-temporal video grounding.Advances in Neural Information Processing Sys- tems (NeurIPS), 35.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Embracing con- sistency: A one-stage approach for spatio-temporal video grounding.Advances in Neural Information Processing Sys- tems (NeurIPS), 35

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.416663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:19efa111806dcdc4fd7df66a8acceb4b6fb5b990e36576e069353b645ec7ffb0

Observation ad30844a-01db-414c-9e27-a4ba1cf84b14 · outbound

This paper cites Language repository for long video understanding.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Language repository for long video understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.376136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:e62a125b6b6446a45c14f1a837e2d85656c3f279bd85dc52d041e1936d228c1f

Observation abdc1bd4-909d-4223-b571-bcd3f492cd85 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.386474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:7f150ffb74fb8b2cb98afac72526938ce32964f024753612b232916044679af2

Observation c7763273-4139-44aa-95a3-3aef49417473 · outbound

This paper cites Dense-captioning events in videos.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Dense-captioning events in videos

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.322625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:4bdc2ccc133d3275f2025b820bba42c96a19a56f54d528671ea1ac48b843ba0e

Observation d3550782-3e28-4ce4-9811-8939a0bc7ba8 · outbound

This paper cites Berg, and Mohit Bansal.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Berg, and Mohit Bansal

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.391122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:202688f2e24d4a761fc21a2fda4097ee45662edeb3302a5b49fe3b534fcf30b9

Observation f1ed475d-ccd7-4a0d-8357-fa598ae586fe · outbound

This paper cites Detecting moments and highlights in videos via natural language queries.Advances in Neural Information Processing Systems (NeurIPS), 34:11846–11858.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Detecting moments and highlights in videos via natural language queries.Advances in Neural Information Processing Systems (NeurIPS), 34:11846–11858

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.526519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:5f1eca1d4c3500f293fa7cf879c86e964992d3b7da5a6c6b37f6a61bf1dfe096

Observation 3c75648d-0ef4-49e2-b309-a9a3f18588ce · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.618430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:9da3ac7b228a2af8c1e8da8fff2d07683ba72fe44e2c0fa55409fbbdf0c7b66a

Observation 44b6eeb1-b677-4c82-a566-fdfe4a14c356 · outbound

This paper cites Llava-ST: A multimodal large language model for fine-grained spatial- temporal understanding.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Llava-ST: A multimodal large language model for fine-grained spatial- temporal understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.387524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:634f9be2e559057121e9a017a1ef879310e7f071b1299ab085b4f7aaa348978d

Observation 3c26515f-7644-4eaa-8667-6f0379f61cb5 · outbound

This paper cites Llava-st: A multimodal large language model for fine-grained spatial- temporal understanding.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Llava-st: A multimodal large language model for fine-grained spatial- temporal understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.585448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:ab3f57667dfbeb4dfee26407e3028973fa55ca228906b386311c2f899fa762e4

Observation 4926f75f-0dfc-422e-8ca1-4d41c72e0e5a · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VideoChat: Chat-Centric Video Understanding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.588348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:ff041bf5cfe97d3a8aab5b56c66664424c04d4f5a200c1858096ff0b34574ed5

Observation dab7d110-72cf-45aa-bde7-ede966d4b50a · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.362837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:e3a9b87e272cc95f9a58a599bc08c575bd3f7861862921677234cb3fdfe2d39a

Observation 98d8840d-f323-4d65-b85c-e426373fbeed · outbound

This paper cites Referdino: Referring video object segmentation with visual grounding founda- tions.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Referdino: Referring video object segmentation with visual grounding founda- tions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.542121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:f82c39022131a0d0ebf19372424682367cb821145bf8e2c4bb151f96aa618db9

Observation 055628cf-ae76-45e4-b59b-f1189c9bdaf5 · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.342329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:fc0a0f93c9c5c46fd8bf84246e3187ca20af02a959a9f2a82090ed5b87ae7a9b

Observation 18a6556e-e704-41d1-8a9f-e41e1cc07152 · outbound

This paper cites Glus: Global-local reasoning unified into a single large lan- guage model for video segmentation.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Glus: Global-local reasoning unified into a single large lan- guage model for video segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.359537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:f8abb8cfca4378b82f54a52f66c04bc07b4a04dd1f8e0ba2b53b735218def8b2

Observation 73071e62-ca17-42ba-9f8e-93748517271f · outbound

This paper cites Collaborative static and dynamic vision-language streams for spatio-temporal video ground- ing.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Collaborative static and dynamic vision-language streams for spatio-temporal video ground- ing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.596988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:242f0128e835ec83e38a0030a736375dcece1101f3464c09628ce9342a4c98cb

Observation e70e5bf2-40dd-446c-b33f-21eb56bef62c · outbound

This paper cites Grounding DINO: Marrying dino with grounded pre-training for open-set object detection.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Grounding DINO: Marrying dino with grounded pre-training for open-set object detection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.392999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:95c921e7f5f8c8888da0ee345902c13a56eebcd2d9669581da50e285e483303c

Observation 5989b7f1-3842-456a-84c9-24deb37c597d · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding TempCompass: Do Video LLMs Really Understand Videos?

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:17.144047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:acd2a8257b2f004c1563b1cda44c8a089a7ac1dc51508832d79104e8b4b849f4

Observation 854c8ae5-24f8-4e62-983c-a1376105faee · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Swin transformer: Hierarchical vision transformer using shifted windows

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.512763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:972e65c6f42d64cbc5f6586c92d86f3b479cd814a9deaf2fa4b921b33a44d467

Observation 470f744a-a41c-44e2-9102-190dfa7b92dd · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.612461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:32b07f097c97acd1a562bafac8fcdb921c605f11f1c410b236289b9a485bfd88

Observation 3780dad6-ab2a-414b-8f9d-9e81fe75ccdc · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.556812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:83433b58a501d716ba2b67531f4e7a5ad4ffb834c1b892238d21cc79b5a5b7fd

Observation ad8f6c2f-e4f5-4540-b780-e9d6848fd243 · outbound

This paper cites Video-chatgpt: Towards detailed video un- derstanding via large vision and language models.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Video-chatgpt: Towards detailed video un- derstanding via large vision and language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.404626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:5d4680f0a5dfa649852d08bedf89b1bfd26ba2b908106a322b8d8b1007c30bb9

Observation 91e44e5a-74f7-4fa7-8e38-f613afa1264f · outbound

This paper cites Genera- tion and comprehension of unambiguous object descriptions.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Genera- tion and comprehension of unambiguous object descriptions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.578060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:e7387eac3bee137d0e9ba3c820da4518b476d65efffdee60e87a5449caace4e8

Observation dc73d32e-0f37-46d9-acae-47524f284c9b · outbound

This paper cites Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.538814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:3ca78f98e782a1cd012c67df3e8c7db0c4ed5e9ce6276212f79f2dc1994be340

Observation 6a904aa3-5c33-48a0-a3d9-cf921f28d043 · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.574958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:aacebe5203209ab88e26c152370d0b7fad8ab9da563f8dd1ba6824a293a8229c

Observation f98fbac6-d048-4c40-aa82-6c942f6f8cea · outbound

This paper cites Streaming long video understanding with large language models.Advances in Neu- ral Information Processing Systems (NeurIPS), 37:119336– 119360.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Streaming long video understanding with large language models.Advances in Neu- ral Information Processing Systems (NeurIPS), 37:119336– 119360

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.593181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:7f25e1ef478a7f17c2d9d5cd76e67c09a187479fdb49634f0c4ce5078d77315e

Observation 24f6e6ea-3ca1-41f3-ac95-1482a9a73e57 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding SAM 2: Segment Anything in Images and Videos

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.598362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:8f085ad90b240ef632d9f86f18929b26e7e3bce4290d37043f01c18e33a2395c

Observation 3e07ad01-53fd-4060-9769-0c15e0ec069f · outbound

This paper cites Ground- ing action descriptions in videos.Transactions of the Associ- ation for Computational Linguistics (TACL), 1:25–36.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Ground- ing action descriptions in videos.Transactions of the Associ- ation for Computational Linguistics (TACL), 1:25–36

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.563954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:7843c5eade32b783c1f9a7eaf0f868d1bc67fbf91c389a89bb2be395a907b8b6

Observation 0bbdfc4e-8885-4c6e-bfd0-55933f105876 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.380110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:16c397dbfe562046391c23cc34f1a42d2c4b091661e3c5181e9304d74823ade4

Observation 9e5eac96-d3a7-4977-b03f-6de7d9092c6a · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.523998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:21dab0e0c3c90d697fb458e2dc5b2a47137cb7ca947f08e8b2655369c55254f7

Observation 7ec0f3a3-2e73-4831-8b4c-eec4f2f10c1e · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Moviechat: From dense token to sparse memory for long video understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.560268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:0aacdb1b200cf826c7f91268165df2e2da1771cb0b4facc1c8c479df65f26e31

Observation e85449fd-1b50-4556-9d36-f763e20685f9 · outbound

This paper cites Tvsum: Summarizing web videos using titles.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Tvsum: Summarizing web videos using titles

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.567376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:dcf3bc3c43a87f07bcf3a93428741784249512c5c13e0531261f4e20ab999d79

Observation 60eb5f8a-2ab5-4393-8d2f-8521301de291 · outbound

This paper cites STVGBERT: A visual- linguistic transformer based framework for spatio-temporal video grounding.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding STVGBERT: A visual- linguistic transformer based framework for spatio-temporal video grounding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.574581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:2d48c30f8d8bad6a8b32630467b967f040d4b935789b5ac978c2bce46c340965

Observation dc50b21c-4e74-489d-9e1e-b41505fdee92 · outbound

This paper cites Human-centric spatio-temporal video grounding with visual transformers.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Human-centric spatio-temporal video grounding with visual transformers

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.554120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:67607cd71b571b2d22f631f98d4cf01a8afe2836b934c2e5b9e9e6f172a7e4aa

Observation 19a5d37e-6486-4a8a-b153-7e6e0d905da5 · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.609124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:367e282995e4ac71698f6e3234a104aa02fc6724ff239e8cced6e65f7e4a14c3

Observation 994323ee-fe0b-49b7-8ad5-1f43387f8019 · outbound

This paper cites SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.621861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:af8ba1d0e3fdc817a1b99823bf4e217b11247dace8c8f7c35a677e55d61b30bf

Observation 040a74d2-2dee-4802-b188-0c456b63ffee · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.547114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:81e96921d94e5da213fc713b3e40d118ae30dbb179223a430d51942463beadd5

Observation 39b8fe18-aaaf-4c37-9a40-2fc6a16bc31c · outbound

This paper cites Can i trust your answer? video question answering.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Can i trust your answer? video question answering

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.366458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:250eb4a04bbea8382791a369878a2922909562380a5aafe534b4805582834fe6

Observation 26411682-4ebb-4120-b129-3efad12db353 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Msr-vtt: A large video description dataset for bridging video and language

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.548266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:f35cd2b1ce87d84455cada14c5fc5edb86ff618a1b5d97e746b976d792c1acf2

Observation 5ce41bf4-68fd-4d56-847f-814f054843f0 · outbound

This paper cites Visa: Reasoning video object segmentation via large lan- guage models.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Visa: Reasoning video object segmentation via large lan- guage models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.551056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:52480caa68e451ced4c74f896951cb59424adbb41aa578df77c270df640a33a6

Observation 2cd3a306-6a8e-474e-96b0-bb55670525d1 · outbound

This paper cites Task preference optimization: Improving mul- timodal large language models with vision task alignment.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Task preference optimization: Improving mul- timodal large language models with vision task alignment

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.557132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:337071de3f3933304ff7497747193909238b28c627e878311eabb0a48f8dfa11

Observation 93b6b99a-001c-48a6-b010-fd366be95e5f · outbound

This paper cites Tubedetr: Spatio-temporal video ground- ing with transformers.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Tubedetr: Spatio-temporal video ground- ing with transformers

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.570612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:e850a4c1c4c3b71cd929609b947e13b0d699bb8b09c45be754f3a6fd5f0c45f9

Observation 2192f487-8071-427c-b643-8aa91420d2df · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models.Advances in Neural Information Processing Systems (NeurIPS).

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Zero-shot video question answering via frozen bidirectional language models.Advances in Neural Information Processing Systems (NeurIPS)

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.581802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:b4c40e13f93e10aa6e74324093e61bc1c539dc9193067a696c89b17f5c95fa26

Observation 79f23c7e-a294-4381-830c-ae8e0993b51a · outbound

This paper cites Qwen3 Technical Report.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Qwen3 Technical Report

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.534178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:aa2d40dec4b80e6867aef9920616bcf976dd21bb4720901279b668fbfec33467

Observation 3487de46-2ea9-4107-9fa9-66e23515b09e · outbound

This paper cites Tenenbaum.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Tenenbaum

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.545045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:418e68761e320da3a9ae8a66322449b3122ce9115a7a4eb373bb78417f7c9ece

Observation a8fb2e04-5606-460b-a560-f9ee8b8dc6b6 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Self-chained image-language model for video localization and question answering

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.394885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:4e6a074107bcc346898f751d04c54bfdc3b1894bc11bf1a24076698a3b223669

Observation 0bf8e50c-39f6-4e84-860c-d0531d8dfc4e · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.519397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:128f940cb6f8e508eeba3b30f2e7662b586bf32956e87916fa900549c30f9b2c

Observation d3f6849d-288b-4d9f-86fd-6668e103d25e · outbound

This paper cites Sigmoid loss for language image pre-training.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Sigmoid loss for language image pre-training

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T01:01:25.345422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:62b5ab9accb26a41153e6cad3433530a8362818b3d62058a522956fd6e60161c

Observation 645d8179-2cb5-4179-9228-8250a797ca45 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.605355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:d015fdaa2e78907ed2eab54a96e8acf704eb01060e8e4fc4aaecf1aec88af165

Observation 0b9a1aa6-493e-42f0-baad-440c81e1a141 · outbound

This paper cites A simple llm framework for long-range video question-answering.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding A simple llm framework for long-range video question-answering

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.536188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:983fe57a156575585784f50cda29506f25c18c84a17fb825cd680c3776b47c0c

Observation f4f65383-30d6-4e85-a504-81211bb2062d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.615391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:8d4300f03fd22cae3017bccede274f255c01d6f6e7f96b73baba842c2cf6404f

Observation 129f71a8-9fbb-4def-97d4-53c46bd94599 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:58:46.601891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:a7dbbcb074e4cc4519a59fe406fc8da79d9fe181a901e3122ecfe61bb84b1c3d

Observation 890a4901-07c9-41c8-bfd3-773d0ae1204e · outbound

This paper cites Where does it exist: Spatio-temporal video grounding for multi-form sentences.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Where does it exist: Spatio-temporal video grounding for multi-form sentences

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.539066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:e3718a115195882f594abccfcd2f8ddc4ab741e4e83d2d451be89972a77d65fc

Observation 5c190997-ec5f-41ba-b72f-d5d5304883e6 · outbound

This paper cites An Open and Comprehensive Pipeline for Unified Object Grounding and Detection.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.594864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:b410d2f78793e78bab60246fab78c57cb47700a6972627014e94af251204e55d

Observation 733014c9-ed2f-488a-b7dd-faff8d6fd368 · outbound

This paper cites Given textual featuresQand video featuresV, LLMs employ a textualized box decoder to predict a sequence of tokensy 1:L = (y 1.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Given textual featuresQand video featuresV, LLMs employ a textualized box decoder to predict a sequence of tokensy 1:L = (y 1

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.518611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:9b3901a2b018f9621345c325051caa9c8ade427b66aa8ee6f53fb8416e6d9c91

Observation b75207c1-56e7-42a1-9fba-a15237e2014c · outbound

This paper cites <image> Locate the visual content described by the query <query></query> in the image.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding <image> Locate the visual content described by the query <query></query> in the image

Reference 82

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T00:58:47.515615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:2529c9269bceeeb9510ddba5248d8b95b5e64c95a0445ece4bc9582530f5392b

Observation 3dffa583-cbad-45d7-b506-2eaafeef2ccf · outbound

This paper cites What‘sundertheskier’sfeet?.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding What‘sundertheskier’sfeet?

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.522541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:2bd2ffd936e8c95f098511becda0e12a7503997c3b9b5b32b6dbb76e120b4e26

Observation 53cd6288-2f35-4907-8d4b-8d88dd258ac2 · outbound

This paper cites Rows marked with*are zero- shot settings.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Rows marked with*are zero- shot settings

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.533023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:4a485c6581c657d86ada9ca0439e949e530d35eefb67ee69ddefe9310459aae8

Observation 791fcf43-268b-41ec-97f0-4e99d341ef3a · outbound

This paper cites 6–11 present additional qualitative results across im- ages and videos.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding 6–11 present additional qualitative results across im- ages and videos

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.509958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:24022fdc8e5b000584be6a20f9d7d9e95593e5a6506bdfc696e02234fa6a923b

Observation 2f563edb-f2f8-474d-a3e0-9305b23cfa48 · outbound

This paper cites Consequently, DEViL is trained to emit a single RST and retrieve one tube per query, without explicitly modeling multiple entities or their roles.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Consequently, DEViL is trained to emit a single RST and retrieve one tube per query, without explicitly modeling multiple entities or their roles

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T00:58:47.529950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:1d5d787cfee9e5954dd9d80fc5ee3a2baf849b0ab0ed6f61f65d12f3039f5abd

Pith citing papers

No inbound Pith citation observations are available.