Pith. sign in

Paper Citation Record · LEDGER

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data

As of 8 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 1 inbound Pith citation observation for arXiv:2509.03501.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03501 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:58:20.690107Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T09:20:32.920925Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T09:21:20.658722Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy44
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f32d129a-eba2-490c-9936-8d5edfae3144 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.439297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.439297Z digest=sha256:8b5c12cac07812cc365b3bff483655b187d1cef172f172117712fbc73e6343a8

Observation b3ccf81e-1b56-427c-9b08-7e20054f5d45 · outbound

This paper cites Vicas: A dataset for combining holistic and pixel-level video un- derstanding using captions with grounded segmentation.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Vicas: A dataset for combining holistic and pixel-level video un- derstanding using captions with grounded segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.478984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.478984Z digest=sha256:cc52e9fbdacc1463ff5b6b97334cb6e32bd76e62ba329dc64dcd0ea64025f4ad

Observation 5393e087-47aa-431c-a067-46362548aeee · outbound

This paper cites Qwen2.5-VL Technical Report.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.511934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.511934Z digest=sha256:eb7b071998cc25513fd2e7bc397e6017c421b03a3157e613851fcd3adbae5db8

Observation 0acf6b67-da61-46a4-8d88-8b9b91104798 · outbound

This paper cites PySceneDetect: Video Scene Cut Detection.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PySceneDetect: Video Scene Cut Detection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.546854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.546854Z digest=sha256:655c9e465d5f86f3922e4337542a18287c052fa659d0c8712e618ebeaeff450f

Observation 0942c527-958a-49d8-a7a2-3e6248af82cb · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Sharegpt4video: Improving video understanding and generation with better captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.594910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.594910Z digest=sha256:2798f3d946387b7afa0fbac015c3c0a3eaf7162dfcc1ac09fda768d4776d46c6

Observation 8a097da6-4d02-45a7-9027-91034c18167b · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.625667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.625667Z digest=sha256:81bc23d16c3b694e3b29826bf365d472ee2c57403564a4b063645403ed002e0c

Observation 5b0f2dbe-11db-4cda-bb90-e1f1021855e8 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.702318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.702318Z digest=sha256:d1e60e6475a50f46988162e744869c464452473d3e1025a49dc1c2c1557c1c1d

Observation 2d03d88a-5f72-4a3d-bf4a-d778e64fa9fc · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.748440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.748440Z digest=sha256:5de87a277f587f99c2fb42b569d4921f62f924766af07d08068d21ce5db38450

Observation 048c6c3c-bdd1-4fc5-86b3-6ebce32f8450 · outbound

This paper cites PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.783270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.783270Z digest=sha256:5a17b546d03675aa1d890cb8a37e1b5de6829efe687dab54fcbf8a33ef48bad2

Observation bd4597e5-0732-488b-8a8c-dc1b7b90ad62 · outbound

This paper cites Unifying Specialized Visual Encoders for Video Language Models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Unifying Specialized Visual Encoders for Video Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:58:21.479672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:15.827629Z digest=sha256:8eac6f22a58695ae1c3a02dbf785f69e3b4757b823d316fcfd23f452c485a733

Observation 39097890-3252-438c-801d-864c4e3436f5 · outbound

This paper cites Videorefer benchmark evaluation for general mllms.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Videorefer benchmark evaluation for general mllms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:15.862688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:15.862688Z digest=sha256:1e8d292846714e2835f366b6a1e9e35cbec22d97d7872242c0f893c4ffb53b27

Observation ea89e910-4024-4328-91d2-40ecbdec82b8 · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:36.378923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:15.886354Z digest=sha256:a1acbb6bd119a081557aeba6f4e46e81fd414c0407ac493eda0e2f27572775db

Observation 20a1c4c3-7e96-4250-857e-91bb626e0f3c · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:36.140120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:15.959077Z digest=sha256:df1aeff8e4c708f32a1a20e5b2378c5b0e53bc2c699dc6288ff0c9c36fe45d5b

Observation c4716635-c7ba-48e6-befe-dd37db008796 · outbound

This paper cites an unresolved cited work.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:58:35.911323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:15.964475Z digest=sha256:3ff8168a2be663e64c3f4e3af245fd0490192fe78b820bf6567c558b9e4ead1c

Observation dc4f1d6e-e92b-4609-96e1-011c9c7cfa23 · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:35.677128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:15.997832Z digest=sha256:d34c4b67f8dd9874cede587efeabb32089291f1c2a0476c188b587a55fa16f99

Observation 239fdc4d-74dd-4aac-8b13-50ca7b2927b2 · outbound

This paper cites Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.021381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.021381Z digest=sha256:df6f63f0cf34de7080e434126e45b6b0e09a0ad78497f6714f63ae087436b3d4

Observation 6b30dc52-6e55-4609-94cd-8a3165b86594 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data CogVLM2: Visual Language Models for Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.100830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.100830Z digest=sha256:3fe87586b007f31545b0646956ccdc678d5f39106a653e3db8101e32c6a8c0ef

Observation 8d6e30f3-90d9-4439-b413-0fe2d25dcc04 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Vtimellm: Empower llm to grasp video moments

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:35.448603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:16.153565Z digest=sha256:66bd9fbb5c4ed77468d260d8018fff1191701f7a5653f93bc66def64a73812c1

Observation ca09e8d2-9d88-44e3-bebf-71617c12196e · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:35.167898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:16.212317Z digest=sha256:eff7c918c457018f8f6e5c0c23d34f187192dae5527561cce172c3d46f410fed

Observation c7f1e1c6-471c-4efa-b6ad-ef27fa137087 · outbound

This paper cites Referring to any person, 2025.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Referring to any person, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:34.871035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:16.255854Z digest=sha256:a2868ee1231726356e0286cce8020027193cdf9334b7ad7cfa576aa929b845c6

Observation 8d707d43-44e9-44ae-bad0-275f8ebeb5f5 · outbound

This paper cites Miradata: A large-scale video dataset with long durations and structured captions.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Miradata: A large-scale video dataset with long durations and structured captions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:34.590640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:16.276413Z digest=sha256:e11da9c7df7e6e4e94b2039a1d79a3e351a2298ee97c2c94772dc57dba0797c0

Observation 64164d42-fc27-40eb-8b7d-6a7dd2404fe1 · outbound

This paper cites Large-scale Pre-training for Grounded Video Caption Generation.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Large-scale Pre-training for Grounded Video Caption Generation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:58:21.238931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:16.322852Z digest=sha256:1d8847314462883cfe280aba1173c6bc5f69ff5c1769e71d85c646aac2b1a84c

Observation 490269bf-c118-43c1-9d77-91433c9dc0bd · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Detecting mo- ments and highlights in videos via natural language queries

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:34.262660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:16.360030Z digest=sha256:8c171d4e6226c3325f29f3d3190f33d42422957519e5ca709bd54e16c3a001fd

Observation bdbad2d6-08e4-4336-8bb6-36cb5609d366 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.394371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.394371Z digest=sha256:f2f640475f73546db175fd56d10a90b6f9a2bb784cd8b5309f5a9f917c0df109

Observation 565fbcb7-361f-4c0f-a99e-47a1c40b960c · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data VideoChat: Chat-Centric Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.444568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.444568Z digest=sha256:ea3b82ad99d4c03b15a991a0d056cab3e0d4a91da446a7812f833f7a6a38fd35

Observation 4e2dcb1b-3744-4395-a5b1-e1c1d46511d2 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:33.986752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:16.510722Z digest=sha256:f24300843b024d85575c214c966506faaf420767838fcafc2a96461c457f35dc

Observation 5f901906-4b2c-4118-aa92-62de57bd8c5d · outbound

This paper cites Temporal reasoning transfer from text to video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Temporal reasoning transfer from text to video

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:33.713478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:16.553789Z digest=sha256:ffbc11c2f1013e217913db4fc815f5354ebe4abe0753898aa2626b8b4170a35e

Observation b7e4c9cd-d50d-4af7-8df4-24e422889d97 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Llama-vid: An image is worth 2 tokens in large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:33.367721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:16.591382Z digest=sha256:f1b54020c5c2c54584830c7b318cad7bee94b1a7034d65201c5bdf673ad79f65

Observation 17a543bd-35a1-449b-abf9-58aa9b4c13ea · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Describe Anything: Detailed Localized Image and Video Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.670856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.670856Z digest=sha256:5548f987ca1ad527ef83979ff961dc7c62e4d1259b1ef50687599d69bdd465a4

Observation 51407507-def0-4822-9ef0-dcab963e35dc · outbound

This paper cites Unleashing hour-scale video train- ing for long video-language understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Unleashing hour-scale video train- ing for long video-language understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.708257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.708257Z digest=sha256:cc8c984fedf042c6020045f77a4d1f05e98e1c740c4f4cdfccff5cc7c99f5a48

Observation 2b0f3896-8ba2-43c2-ac5a-aa020adf157d · outbound

This paper cites Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.722899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.722899Z digest=sha256:6dfa806711003fb33298ffc6ab8207bf0af0b4b5e19db30adbed1aa01bba6360

Observation 7d10a3ff-980e-41d7-a53f-13c215e27703 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.779731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.779731Z digest=sha256:42fe1830bc290342509c0379b6a82104ebd1dd6554fc6e1e995218d635ca81d3

Observation 10c054de-5eed-4652-8f0e-67c4dddb1291 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data TempCompass: Do Video LLMs Really Understand Videos?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.815963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.815963Z digest=sha256:606922c30a45b45d376ab70de228055414eafd42d0a03cbf3ed0249e2a7e5ee7

Observation 8baa8687-9803-408e-b636-62050abd02b7 · outbound

This paper cites Groma: Localized visual tokenization for grounding multimodal large language models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Groma: Localized visual tokenization for grounding multimodal large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:33.010944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:16.858999Z digest=sha256:545324ae90c73f3d6635f8d5a59fcc8079d49baf4e2218c86bd3b01f39d81648

Observation a6ad5f79-7adc-411a-9357-50dd0a011a34 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.903200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.903200Z digest=sha256:dcdbc344ce13bc7cb347c5ae9f7d793c8f432e7f20d66208ebcdc71d424387b1

Observation bd380dbb-488a-4b83-8b6c-16e9af6165b4 · outbound

This paper cites Point and Ask: Incorporating Pointing into Visual Question Answering.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Point and Ask: Incorporating Pointing into Visual Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.959800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.959800Z digest=sha256:62ca330302331bc44484f80b2747a6f1593ec4c38eae16ec5db9e4bd1eb90a66

Observation 94876515-aa3e-476b-a6bb-d9e3c9eeeb79 · outbound

This paper cites PG-Video-LLaVA: Pixel Grounding Large Video-Language Models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PG-Video-LLaVA: Pixel Grounding Large Video-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.017701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.017701Z digest=sha256:618587a62a62681a2f137c15dfcb4f0675183464d023a4de37532e5a7c574da7

Observation 716af1ec-e4e1-4635-9627-c571d526b777 · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level vi- sual grounding in videos.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Videoglamm: A large multimodal model for pixel-level vi- sual grounding in videos

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:32.712204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:17.050270Z digest=sha256:b035d813ab786788a9a871da451b4d09171375a09dcbab8fc87adbde0fee5950

Observation 5ff4291b-5505-4f60-b17f-e45a306f10fb · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.121792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.121792Z digest=sha256:4bd958ed1f002e582fb69378281bb83cc9690eeb20b782327eea4b7ce1374442

Observation 0b2eee42-add2-4e91-92c4-61e8bce2a2ad · outbound

This paper cites Artemis: Towards referential understanding in com- plex videos.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Artemis: Towards referential understanding in com- plex videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:32.345406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:17.184642Z digest=sha256:06352a82f798f05d5364290c724b9de8581a8b821117db4ed99bd91d65f389f5

Observation f1ab99bb-d21e-4920-99d6-f9af787c780e · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Glamm: Pixel grounding large multimodal model

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:31.969203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:17.240520Z digest=sha256:2b87c12a1b315d629ea9390ab965d99979f6835e8bfc001714b873fa2d6b93d2

Observation cd859a89-9070-41aa-871c-aff310a6b94a · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data SAM 2: Segment Anything in Images and Videos

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.383157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.383157Z digest=sha256:f44c6b5967a9e92bb6156d36f943dcde418dc480c8baec4c4a849d8ee1f183c3

Observation a4b7911d-4e68-4380-aa44-c84057af99cf · outbound

This paper cites Grounded sam 2: Ground and track anything in videos with grounding dino, florence-2 and sam 2.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Grounded sam 2: Ground and track anything in videos with grounding dino, florence-2 and sam 2

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:31.616302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:17.432596Z digest=sha256:de409f8152c8a4e05eab5dc2abfdf7a3d3132db6d9e0b954d26f8fe6b30db285

Observation 2f9abcff-b849-49cd-b4bc-18e75325fbd6 · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.487286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.487286Z digest=sha256:32a4a6a89e8194ec61336d364195c67ea4240fd6fe9853d612b3a2aaa45a3e97

Observation 99bcb17b-27a2-48d3-93ce-16d3844b925f · outbound

This paper cites NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.549752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.549752Z digest=sha256:905eba17209d6aff0edfe2fa841fb2f433e3677e1fced4a31b1353455b9d05b7

Observation 1bbee3d4-1b02-4f24-acb0-d7539bf57205 · outbound

This paper cites Sama: Towards multi-turn referen- tial grounded video chat with large language models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Sama: Towards multi-turn referen- tial grounded video chat with large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.583118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.583118Z digest=sha256:9c08fc373c9a7419170aa139a1543868fc030a5f7c68110d95eb47bbda4dad59

Observation 0c54bd9b-9a9c-4f0e-9a94-cc1157014dd7 · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Qwen2.5: A party of foundation models, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:31.332165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:17.612753Z digest=sha256:2a42966c5f47bcdbddf55bf32094d321bd52376eebc31bfd636a0deb455be27a

Observation 7214d2b9-9d9e-41f2-8fbf-0bf0ee11b373 · outbound

This paper cites Natural language processing with Python and spaCy: A practical introduction.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Natural language processing with Python and spaCy: A practical introduction

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:31.040447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:17.667045Z digest=sha256:9b4d0a7949f77ceb511c5bcf19b82b486105e1ca23850fcec6c14828c98596ef

Observation 41ca1c2b-a3bf-4f21-bef5-9b4434fbc12a · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:17.747639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:17.747639Z digest=sha256:375429927acad786c41def9d9bf3c37be10d45fc7efea877645a65d0e31329ab

Observation 765d890c-9a71-4e79-97ce-21c300b094b2 · outbound

This paper cites Elysium: Exploring object-level perception in videos via mllm.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Elysium: Exploring object-level perception in videos via mllm

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:30.665992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:17.831748Z digest=sha256:d640855c770ccfc9693a5a2e4442d78d260085eb6be42eeaeef53ee1c590ea60

Observation ff965332-63fc-4e5b-87ca-42b14e74a0f2 · outbound

This paper cites Tarsier: Recipes for training and evaluating large video description models, 2024.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Tarsier: Recipes for training and evaluating large video description models, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:30.357181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:17.857705Z digest=sha256:38194e53e3b98cfe6d8dc2222ef44f193a2bbdb1a7926e94130fba273ba426cf

Observation b15deeb9-3579-48ff-a139-a73346e85ea6 · outbound

This paper cites Language as queries for referring video object segmen- tation.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Language as queries for referring video object segmen- tation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:30.042271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:17.889827Z digest=sha256:6f83881cadedf7611bac8f903eb53fa4354c26c8c4bdbb94afddc1c7de5a3373

Observation 88627f00-7b73-4104-9607-25201e2aaa50 · outbound

This paper cites LongViTU: Instruction Tuning for Long-Form Video Understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.003545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.003545Z digest=sha256:6ceba393e52c614bf9e8180d940e2064fe296fb856afa277340f03b930d38582

Observation 21cda32f-45a3-4ae8-a22a-8e9a1467a8ae · outbound

This paper cites Number it: Temporal grounding videos like flipping manga.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Number it: Temporal grounding videos like flipping manga

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:29.734566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:18.090304Z digest=sha256:1d19d971070d00dd00c76694ab6198b10a0c2683d5f307572a269d668d32fcd3

Observation db8b1beb-fc4e-4632-af2d-dd2f0be45627 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Next-qa: Next phase of question-answering to explaining temporal actions

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:29.392629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:18.182650Z digest=sha256:449cd8c7359551095617ecd8eb50786178bed181aa193b43731dcebfbfb0c632

Observation fca9e09a-1ba7-4a9c-b092-17d5b39bb4d6 · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:29.082214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:18.254568Z digest=sha256:94c872eecc05771ceda334b6361f20f8e1526d90fd9ef439fdef05a45c1e55e6

Observation 29bf663a-d044-4637-a009-e82cb6143f55 · outbound

This paper cites Pixel- aligned language model.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Pixel- aligned language model

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:28.697397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:18.284425Z digest=sha256:7a5221c0eb3c5540d85df38c6788fa2b43e5e18bc5604e82ae64d54f07902a1f

Observation fec93ee3-44fe-4d4c-95c9-34e70626e67a · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.338413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.338413Z digest=sha256:c1227a99c070fea5cb223c782cd7cd903bd1fdfa5ea4656ed1780d15e1c75ecf

Observation e49916f7-c529-4d14-8a85-dbf10b8ef374 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.397448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.397448Z digest=sha256:91a212e76b8350b4cb51bb92d1089d984433fa4757029f83152777316f103a78

Observation 823bc005-5621-42bd-9c92-70213b7cdcbd · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data xgen-mm (blip-3): A family of open large multimodal models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.453965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.453965Z digest=sha256:9f9f42445010bc9d92940c2c2cadb9e98e25faa647a3ed82ab182dcd7f721115

Observation 77f0dbfa-cd10-4f3b-99b6-b97953a77307 · outbound

This paper cites List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.485504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.485504Z digest=sha256:9c2d398a829081171360e14b035225ac1bc71e0706b4b73c56429842077aab7c

Observation 98767322-0937-462c-b96f-9cabf0103a4f · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.599712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.599712Z digest=sha256:f95ff310c98002520f4b39a7d7a38d3cf177a098a4c33e2a1b7a7e489e7193c6

Observation 93f24843-fe72-45a5-b580-731de430261e · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:18.651851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:18.651851Z digest=sha256:302ffaaa2dd2203e3bb4493d641e39cadd8756b16dbb8cd142857b7a008e49d1

Observation ce8cfc46-6dae-45ea-98e2-9f2899656ef4 · outbound

This paper cites Merlin: Empowering multimodal llms with foresight minds.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Merlin: Empowering multimodal llms with foresight minds

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:28.238086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:18.742480Z digest=sha256:b26988998444a11488bd07ec8e41cb5cbcfaa4a82ba4bae0b593c6646d603804

Observation 2ebddabe-83ff-4b38-bf40-e465dcc7d3d2 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:27.920430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:18.800425Z digest=sha256:43ac5cb9d1d14947670091af09d2fe0c3978940a739077bc164a1bc73adf6071

Observation 9aece6be-3583-454e-8cbd-a93e978b42d7 · outbound

This paper cites Osprey: Pixel un- derstanding with visual instruction tuning.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Osprey: Pixel un- derstanding with visual instruction tuning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:27.591043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:18.873688Z digest=sha256:9a9beb0d94b15491a325bd18945212bf0a38e8e8a05c92cab51a7bade431306b

Observation 3a73b83d-ca66-4222-8e4d-56efb146d259 · outbound

This paper cites Videorefer suite: Advancing spatial- temporal object understanding with video llm.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Videorefer suite: Advancing spatial- temporal object understanding with video llm

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:27.131183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:18.929032Z digest=sha256:c4e225898156dac7067af5aee26ccc4e420c20efc3c98a1f8a6de75cd26c6504

Observation d5d9aa8a-38c1-4266-bf08-5ad6c73145e3 · outbound

This paper cites Sigmoid loss for language image pre-training.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Sigmoid loss for language image pre-training

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:26.785238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:19.027799Z digest=sha256:e575c580286755c0755f1af148419b7044e721a4900671da5e952c0d618b42f0

Observation 9d81e749-85c4-4ba8-ae83-d53265e10d7d · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:19.096136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:19.096136Z digest=sha256:20d4d839ec63e7057858623463338c7b0ea3470794facd54bfead2e55d0938e0

Observation 8918d51d-37d9-45ad-be60-d56aad392477 · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Llava-grounding: Grounded visual chat with large multimodal models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:26.325015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:19.153012Z digest=sha256:ab48ad432b7eb70b6d6da42c0e37d78f93a06f4361bcc4f7ea9d4e62d8aeec70

Observation 65a71161-201a-4948-826e-c30a6127c282 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:19.257522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:19.257522Z digest=sha256:508eaa0f958cc0ef7dc772b6e6d9f6b9821e35d6dfc7d4df5ed4d4803c0281e3

Observation 103af01e-08d4-495a-a86f-d49e1f4b426e · outbound

This paper cites Gpt4roi: Instruction tuning large language model on region- of-interest.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Gpt4roi: Instruction tuning large language model on region- of-interest

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:25.935731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:19.342543Z digest=sha256:d9f79efb3c7791326e59436325b530819be21425ff963e77c3d4d0955e15df71

Observation 0db6e86d-0d55-44b4-8dbc-7acf2127e4cf · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Video instruction tuning with synthetic data, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:25.520312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:19.436000Z digest=sha256:ddd22dac10dcf5f589a1aceea1548cede68d32bad5ccc0caaef9376912f9da83

Observation 59a300d9-b8e6-4ef6-a34b-318b14e219a4 · outbound

This paper cites VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:19.515582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:19.515582Z digest=sha256:a360d96f20fd1b4761db4b985a69acfc899f0cb6d8b11c55eb15312f72654f32

Observation 29fb73ab-44d5-4544-b861-1cfecb24fc66 · outbound

This paper cites Frames are extracted only from the seg- ment of the video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted only from the seg- ment of the video

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:25.072468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:19.612113Z digest=sha256:1f78fd285780e03b89c6fa55221575bfb1efb0817f63c8e0ea54d57f4bc0bf13

Observation 796b7f92-13a3-4030-89cb-11fb6d2ec278 · outbound

This paper cites Sorry, I’m not sure.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Sorry, I’m not sure

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:24.720638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:19.706926Z digest=sha256:223e1a800be63de1da1b4932df47a49f3c8248290013ff0ff8055825006fbb04

Observation 278c6d5f-1929-483f-8e31-c7a64564ef6a · outbound

This paper cites Frames are extracted only from the seg- ment of the video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted only from the seg- ment of the video

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:24.365748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:19.711705Z digest=sha256:c1fd4bbcda36af9032c3f46664014d4bd6367a04b7ce975337d881cb6a96fb53

Observation 7ad677b9-c848-4f99-9c64-189edb0320df · outbound

This paper cites Yes” or “No.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Yes” or “No

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:23.998339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:19.893683Z digest=sha256:43ceab05d287e3e28e026fa80baaf013017806262aee1bf699cedd367b4b04e8

Observation 7d300569-30cc-4e3d-842d-d40c885d5598 · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:23.669493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:20.028670Z digest=sha256:658c33aa6b569849968a1bc38984948f670d3330b41df5ee670df217a718328c

Observation 6bbfbe7e-2d92-4072-90ff-90d6db7ad08a · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:23.298483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:20.152919Z digest=sha256:246e7663ac3da9938df8009274094c68e91fc5d18a2219d562105d0f895a496e

Observation 2f61eb7b-21ac-44fd-b0ec-b3351f8cb620 · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:22.941919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:20.210594Z digest=sha256:9226133bd4132e107a283fe87b5d91100dbade4da2cdc3588ea428c3d5878089

Observation 5fa46932-f1db-4106-b288-55198e7607f5 · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:22.613750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:20.340722Z digest=sha256:7005e8a728b587ada67e9d58785652cef5aef8a5f67e0dd75de99ff88c764cdb

Observation c01a5e71-d7c6-47ad-811f-86fc083866b2 · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:22.332818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:20.543748Z digest=sha256:79da00bafea81114bcf69689457d36f528a9d0957057619916f62e9b1c86f67c

Observation 0c0d1e74-60de-41c2-bae6-e8f3673ee118 · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:22.033684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:20.629580Z digest=sha256:7a29b60215c3ea9aae62de726a5b6c6bddd3092cb1170b1ec056422c27b787da

Observation ac81c5fd-500b-46bf-a0ae-bd3438dac2a7 · outbound

This paper cites Frames are extracted from the full video.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Frames are extracted from the full video

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:58:21.791400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T10:58:20.690107Z digest=sha256:ec311741849d8e2b7c6014424aa72273b40eb2b99650a4643d440475d835408c

Pith citing papers

Observation 7d5ed8f9-de26-4309-b4df-0653128618b2 · inbound

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly cites this paper.

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:20.660900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T09:20:32.920925Z digest=sha256:18d7996692b840e39af78fbbfa5228339c506e2ca0e91d30c491423722b3d414