Pith. sign in

Paper Citation Record · LEDGER

Object-centric Video Question Answering with Visual Grounding and Referring

As of 19 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2507.19599.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19599 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:19:40.237501Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved35
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f4a2b7bb-2e0c-4a06-b9fb-9f8b22e30d21 · outbound

This paper cites GPT-4 Technical Report.

Object-centric Video Question Answering with Visual Grounding and Referring GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:30.864771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:30.864771Z digest=sha256:82be92e1ae7f422bef6fda9ca02e9e56f5324b242af64f2f23eb73bf786aef01

Observation 15a59f5c-3d8f-4f09-ab56-e0507cb122af · outbound

This paper cites Burst: A benchmark for unifying ob- ject recognition, segmentation and tracking in video.

Object-centric Video Question Answering with Visual Grounding and Referring Burst: A benchmark for unifying ob- ject recognition, segmentation and tracking in video

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:52.299469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:30.941003Z digest=sha256:317a0449b2f8cefeb6ce8e1d1ed0b3055e02f4768ea26ff085d488bfa679e89e

Observation 72cd616b-060b-4990-93c0-ea6b3f62f79b · outbound

This paper cites Qwen2.5-VL Technical Report.

Object-centric Video Question Answering with Visual Grounding and Referring Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:31.060322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:31.060322Z digest=sha256:a1f64612dac1ef0a3975a62d5aff273b01b5b1b4103f683456ebdebae73c6d32

Observation 64102306-6b7f-4f91-80c3-f32aa1652cfa · outbound

This paper cites One token to seg them all: Lan- guage instructed reasoning segmentation in videos.

Object-centric Video Question Answering with Visual Grounding and Referring One token to seg them all: Lan- guage instructed reasoning segmentation in videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:52.160107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:31.203054Z digest=sha256:a9d51803254bdcaba310952b9970780d77830011cc6ad87e14ac3ab08a282ec1

Observation de2d2fdc-fc69-4e84-a048-286ac09946bf · outbound

This paper cites Xmem++: Production-level video segmentation from few annotated frames.

Object-centric Video Question Answering with Visual Grounding and Referring Xmem++: Production-level video segmentation from few annotated frames

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:52.024268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:31.309544Z digest=sha256:51e8f687d604f4392ef4eb0f1542e82621ae7ab13536e75bc02047e35d680c5a

Observation 3897ae52-368e-4cf8-a23e-9833ac40268c · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

Object-centric Video Question Answering with Visual Grounding and Referring Coco-stuff: Thing and stuff classes in context

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:51.766255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:31.414465Z digest=sha256:938568d3e3a4c046b34cd52bf8931e00c3a0d5c86e08ff325c94d8d6c4cf0f05

Observation 6fca08f5-55d2-42a4-8ad2-c38d9365bb64 · outbound

This paper cites Vip-llava: Making large multi- modal models understand arbitrary visual prompts.

Object-centric Video Question Answering with Visual Grounding and Referring Vip-llava: Making large multi- modal models understand arbitrary visual prompts

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:51.512367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:31.529633Z digest=sha256:be86c2601de3b362defa12eb7af52c429f5c6b176cc4753cf9c3121d19b4b5a2

Observation 5d06ad8b-31d9-4502-86fc-bb077eb6f79d · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Object-centric Video Question Answering with Visual Grounding and Referring Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:31.659606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:31.659606Z digest=sha256:6f076934c9b1b9f9f3e5ac0964d27177c42a95f17388cff4ff41c5881beef3f9

Observation 1ee91e18-9a87-407e-9fd5-9213426db69f · outbound

This paper cites Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos.

Object-centric Video Question Answering with Visual Grounding and Referring Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:31.776850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:31.776850Z digest=sha256:6d7bb1113f48ccfc66cded77bda9e3917ef223b409c0f72a9485b62f13a7ae1b

Observation c8d0e2f8-5dd9-4fac-8c69-b65e8e9c55ea · outbound

This paper cites Detect what you can: Detecting and representing objects using holistic models and body parts.

Object-centric Video Question Answering with Visual Grounding and Referring Detect what you can: Detecting and representing objects using holistic models and body parts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:51.309218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:31.883436Z digest=sha256:3b071caeaedcb3eae47eca2e54a78974435ca5a18167fa87756c0a0a18edae35

Observation 87cfd5da-20c0-4dd2-be86-c82accb6bea4 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Object-centric Video Question Answering with Visual Grounding and Referring How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:31.985474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:31.985474Z digest=sha256:ddc238930067fd9e9437204292418eab8ae96ef7a3fd005195464281155fcd51

Observation 08867776-fa48-4a4a-afc0-7032580eddf9 · outbound

This paper cites Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model.

Object-centric Video Question Answering with Visual Grounding and Referring Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:51.085303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:32.101781Z digest=sha256:858bfc00b34ac32515b16c444f94904ec9fa518f2b25c63d539224384defb947

Observation 467cceab-2790-4e27-8aab-b4c40e40fda3 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Object-centric Video Question Answering with Visual Grounding and Referring VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:32.183705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:32.183705Z digest=sha256:a9d2cb45617d40aa2a0cf58270ec51bbac702ae61814ee0809f34493bfffc236

Observation 6e4f8a82-6789-47d6-ab6f-b8fe37747739 · outbound

This paper cites Grounded question- answering in long egocentric videos.

Object-centric Video Question Answering with Visual Grounding and Referring Grounded question- answering in long egocentric videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:50.848944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:32.316400Z digest=sha256:e69b5c504ce4b40dc3cbcf98e6e885dd990c4f0c7bdd0bf6c046de93c8913868

Observation b9e99d39-f6dc-4276-9c58-823895421b22 · outbound

This paper cites Mevis: A large-scale bench- mark for video segmentation with motion expressions.

Object-centric Video Question Answering with Visual Grounding and Referring Mevis: A large-scale bench- mark for video segmentation with motion expressions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:50.598223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:32.471391Z digest=sha256:44bb9397c9d6bb6fabd3ff1e7227fc9b5db7fc5765aaf5be3b675d3837047f8d

Observation d3ab6386-277c-40ad-a059-8b69694a63cc · outbound

This paper cites Mose: A new dataset for video object segmentation in complex scenes.

Object-centric Video Question Answering with Visual Grounding and Referring Mose: A new dataset for video object segmentation in complex scenes

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:50.279506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:32.592931Z digest=sha256:b064c29989c83c61c71c0e7d4531f96cbb814c562607507158c6a036951076b1

Observation a69ad915-4348-4f40-b026-8ab69739055a · outbound

This paper cites LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:32.702316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:32.702316Z digest=sha256:5e6d9f8a761ebe36beb0581da2e3db238c389c0d863fe7c1efa3480b05332a29

Observation 615ca1a8-b839-481e-86be-2a949d326ed0 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Object-centric Video Question Answering with Visual Grounding and Referring LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:32.811828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:32.811828Z digest=sha256:58fa44b107010c8168697028a692033ec41b190484be182f0ed7aaf39471a386

Observation 2a6d0135-259a-4498-af7e-5c2be488fbd3 · outbound

This paper cites GPT-4o System Card.

Object-centric Video Question Answering with Visual Grounding and Referring GPT-4o System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:32.918892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:32.918892Z digest=sha256:145f7d2f2c23699efd8196031866bea961eac6f575a23f0e3d74ce5ab353c1f0

Observation a4d82382-ce65-4f2f-90f6-875055bd318e · outbound

This paper cites Cotracker3: Simpler and better point tracking by pseudo-labelling real videos.

Object-centric Video Question Answering with Visual Grounding and Referring Cotracker3: Simpler and better point tracking by pseudo-labelling real videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:49.949105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:33.036045Z digest=sha256:107e3a8a1fb89702a745a0bbd55a8182a981255b18aedc564df1148bca19a2f4

Observation 8e024cfc-dbb8-4abf-99fb-9249ed304db2 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Object-centric Video Question Answering with Visual Grounding and Referring Referitgame: Referring to objects in photographs of natural scenes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:49.720019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:33.140267Z digest=sha256:98e33fabb84ac7769dfb9508548c8c1f851dace6ea6dbb8a89050b28dc64f4c9

Observation 29c9ab82-2837-4706-add4-46a147faccd3 · outbound

This paper cites Video object segmentation with language referring ex- pressions.

Object-centric Video Question Answering with Visual Grounding and Referring Video object segmentation with language referring ex- pressions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:49.479956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:33.238590Z digest=sha256:3c0f80ab8e671e6cbe1eea6e19006fdece16986007dd14c0dd1804a5d2e0cb62

Observation a306dbb0-bdd4-4aa2-8e53-6b96dab7dc20 · outbound

This paper cites Segment anything.

Object-centric Video Question Answering with Visual Grounding and Referring Segment anything

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:49.210901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:33.380141Z digest=sha256:447e53950e89759911337332a22f87ce895bd981f546b257a02786ed8b03ba81

Observation 79e37907-115d-4a89-8f14-eeb865c5259e · outbound

This paper cites Grounding language models to images for multi- modal inputs and outputs.

Object-centric Video Question Answering with Visual Grounding and Referring Grounding language models to images for multi- modal inputs and outputs

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:48.929539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:33.502374Z digest=sha256:9d1df5f53cd18c522d283906f6d26f0235954126f0d8597486e5376e5b01cb69

Observation 58447ff6-88b7-4e0d-ab6b-135f1d9fc163 · outbound

This paper cites Generating images with multimodal language models.

Object-centric Video Question Answering with Visual Grounding and Referring Generating images with multimodal language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:48.636508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:33.605622Z digest=sha256:bc5a735aa7b2205ca9d996c894db245699492e94f90b4aba4478f363c96be4d6

Observation d3f1ad95-2d6d-41cb-b41c-26b35e07aed9 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Object-centric Video Question Answering with Visual Grounding and Referring Lisa: Reasoning segmentation via large language model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:48.380741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:33.708532Z digest=sha256:2ab36bfa94505b8b15ffd84bc373266d64c4d325f9bb28233496acccf0ee0ae2

Observation f50cf220-2210-43f9-852c-2d9ad08ed121 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

Object-centric Video Question Answering with Visual Grounding and Referring MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:33.841194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:33.841194Z digest=sha256:233422484885affeb93410a6907791ca74fee5ce50706b736d9fd4f5ea3c54bb

Observation 95ea52ed-4eec-4bce-990b-893c9990c20e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.009655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.009655Z digest=sha256:e8fa3b4d424e8961afc0fbe3226d716a5b4e75ef12418e92528894872ce82e0f

Observation f7cad20f-21af-4df1-92e8-4210147ffecf · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.124422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.124422Z digest=sha256:0c234f78ff9b0ef141e3e688a90281638973334e96d0ecc43f2fcfd9223ec2c3

Observation 172277c2-04cd-4578-9645-535d88e70850 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Object-centric Video Question Answering with Visual Grounding and Referring Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:48.097986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:34.242376Z digest=sha256:e43636f45393a419f401e283ff536b61ba683001c21f65c59678cfba1ecb67fd

Observation 3a772343-879c-4ae3-8d14-0f5d729134dc · outbound

This paper cites Towards Robust Referring Video Object Segmentation with Cyclic Relational Consensus.

Object-centric Video Question Answering with Visual Grounding and Referring Towards Robust Referring Video Object Segmentation with Cyclic Relational Consensus

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.362625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.362625Z digest=sha256:3d57bb8f81dcadac9469438edd53dcfeabe983df3f6f2aa05ee123f8d443eef8

Observation 0b453265-6047-4a43-867c-f4127e8ee3d2 · outbound

This paper cites Describe anything: Detailed localized image and video captioning.

Object-centric Video Question Answering with Visual Grounding and Referring Describe anything: Detailed localized image and video captioning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:47.846987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:34.543370Z digest=sha256:0bda8d1464a6d0f4237592048e2d3957e975942e2f46db65abdff95a789b0fe0

Observation b81e087f-251f-40a7-ae09-ceb99b331e7a · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Object-centric Video Question Answering with Visual Grounding and Referring Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.701074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.701074Z digest=sha256:daa534f6c16cbecd365931bae9db7281060d0c3f7c5d89bdab15c0f8473628ef

Observation 1237dd00-ff9e-45c2-adab-a977ca078d9e · outbound

This paper cites Rouge: A package for automatic eval- uation of summaries.

Object-centric Video Question Answering with Visual Grounding and Referring Rouge: A package for automatic eval- uation of summaries

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:47.568836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:34.792454Z digest=sha256:1884007aa42670c9ff298fe76015102e954139afffe919f8261e4041fb3f7f65

Observation a155d143-04a0-4678-ac25-5bd7cbdcefed · outbound

This paper cites Glus: Global-local reasoning unified into a single large language model for video segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring Glus: Global-local reasoning unified into a single large language model for video segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:47.292410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:34.885264Z digest=sha256:d83ae324b00673fa7e1baf51f11e93ee27f1c40d7fe9956c9322102031b2ec52

Observation a14696b6-30a5-4cce-8ba6-08e350b0cc02 · outbound

This paper cites Visual instruction tuning.

Object-centric Video Question Answering with Visual Grounding and Referring Visual instruction tuning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:47.019287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:34.986936Z digest=sha256:c15e054c4c95e134de71eee63c2f6515b906507ee629be880cbc229b5493c109

Observation 47c75236-174c-43bb-82be-0cc419d7d744 · outbound

This paper cites Lamra: Large multimodal model as your ad- vanced retrieval assistant.

Object-centric Video Question Answering with Visual Grounding and Referring Lamra: Large multimodal model as your ad- vanced retrieval assistant

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:46.770684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:35.099805Z digest=sha256:949fa3ad874e0ea3adb52b63c88eef5e577653ba1b335828603eb4db5a5e2970

Observation 2dc44f75-013a-403d-be4b-9560b923ed50 · outbound

This paper cites Decoupled Weight Decay Regularization.

Object-centric Video Question Answering with Visual Grounding and Referring Decoupled Weight Decay Regularization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:35.242260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:35.242260Z digest=sha256:e26eff827a03f5e665edc0a59e32af2b43cbc11e2a5b221d32c8e548b0bbf6f0

Observation aa765c0b-ad33-427e-93f2-d3503289bb97 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Object-centric Video Question Answering with Visual Grounding and Referring Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:35.344758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:35.344758Z digest=sha256:d17149da8a04d715273881321f2616c58d3dc4e560403ddff863c9392b643dce

Observation 6cbe70e5-59c8-4e31-98c6-eee78a3e78ec · outbound

This paper cites Generation and comprehension of unambiguous ob- ject descriptions.

Object-centric Video Question Answering with Visual Grounding and Referring Generation and comprehension of unambiguous ob- ject descriptions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:46.495180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:35.462523Z digest=sha256:8e23c6922922983c908e5f6d1467034ae74f769e6d773b7b943ac9db7f869428

Observation 1ea753bb-c75a-4a3c-af41-dcc1cd2756be · outbound

This paper cites Large-scale video panoptic segmentation in the wild: A benchmark.

Object-centric Video Question Answering with Visual Grounding and Referring Large-scale video panoptic segmentation in the wild: A benchmark

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:46.303905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:35.567110Z digest=sha256:aed380531a9e1c69dd8d34678011458955c91c730ceeddd76dd854f7dc7d6c4b

Observation 18aacfb3-e0e4-4215-bf55-5ba9bff3d63f · outbound

This paper cites V-net: Fully convolutional neural networks for volumetric medical image segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring V-net: Fully convolutional neural networks for volumetric medical image segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:46.043914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:35.730518Z digest=sha256:52f301c8db16361db28489d5a73f8c408e1f9de81557a3b3776667b84bf0ccea

Observation 0ed6dda5-284f-482d-af20-571d707c01f7 · outbound

This paper cites Bleu: A method for automatic eval- uation of machine translation.

Object-centric Video Question Answering with Visual Grounding and Referring Bleu: A method for automatic eval- uation of machine translation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:45.806001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:35.851977Z digest=sha256:89d6fa90a20290bd1c3f22b44a53d605329fc38c68cf964ed1a2535154f3fd65

Observation 4fc8125a-1305-44cd-931f-26d4ddd572c4 · outbound

This paper cites Perception test: A diagnostic bench- mark for multimodal video models.

Object-centric Video Question Answering with Visual Grounding and Referring Perception test: A diagnostic bench- mark for multimodal video models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:45.588351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:35.974409Z digest=sha256:c1f56fef61a187d5066da6c40fe2f9ab2d3337165e29574dde02578f35d7455d

Observation 07e3f2f3-ce0a-4fc7-b345-467de8bfbfab · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Object-centric Video Question Answering with Visual Grounding and Referring Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:36.065603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:36.065603Z digest=sha256:d492385a90d175d1f99003e42a661020ffb90c87b6945a8e7ac3016f0b96727a

Observation fa54ea9a-a496-4f7e-aace-b97d167440a3 · outbound

This paper cites Occluded video in- stance segmentation: A benchmark.

Object-centric Video Question Answering with Visual Grounding and Referring Occluded video in- stance segmentation: A benchmark

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:45.342672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:36.172270Z digest=sha256:34bebcf98752fcf4e8f26f34489d8f56074ab1d8fc3fb53bd2d1e887d6539a2c

Observation c5968d3d-6c7a-4dee-8860-bda3c817aa7e · outbound

This paper cites Artemis: Towards referential un- derstanding in complex videos.

Object-centric Video Question Answering with Visual Grounding and Referring Artemis: Towards referential un- derstanding in complex videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:45.082859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:36.327772Z digest=sha256:e9f8b0765a058fd58902f3a337aee3bfaa046933a0c69ccb7d083d96af22a928

Observation fc7d881a-c0f5-4e22-b453-0570dc9067a9 · outbound

This paper cites Paco: Parts and attributes of common objects.

Object-centric Video Question Answering with Visual Grounding and Referring Paco: Parts and attributes of common objects

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:44.828041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:36.435318Z digest=sha256:9d64c2c374f979b8f4162fdc6ae5afcbe81a89988be4dbe9e3917309724bb110

Observation ca092a3f-da64-4fc2-aee2-c549055601d9 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Object-centric Video Question Answering with Visual Grounding and Referring SAM 2: Segment Anything in Images and Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:36.552673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:36.552673Z digest=sha256:c0ae0a38810e6e13ebba76986c7fdd28af91735d0b7ae797a134baaf78d8f89d

Observation 5c03be5e-3b00-4980-88c4-78066345f1ff · outbound

This paper cites Hiera: A hierarchical vision transformer without the bells-and-whistles.

Object-centric Video Question Answering with Visual Grounding and Referring Hiera: A hierarchical vision transformer without the bells-and-whistles

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:44.575409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:36.644218Z digest=sha256:f6300645dacbfa8ee99a5896a6fbd81ede6ccae303fe2af64e9e28e6adc57352

Observation 297c4461-aa4a-413f-8cf5-41104126d5c6 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

Object-centric Video Question Answering with Visual Grounding and Referring Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:44.344740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:36.749964Z digest=sha256:c91c88a48c9419301849421fa2daaec9f5294b3125e7f01d107b5043a89a54f6

Observation 624e5f94-9f0c-470f-acef-f2167f960c44 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Object-centric Video Question Answering with Visual Grounding and Referring Emu: Generative Pretraining in Multimodality

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:36.979072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:36.979072Z digest=sha256:8258bf038cd5bef5d1889297a4838b9a96a30281869e635456c43e77ac65bf05

Observation 1b464f74-7deb-44c6-8ad2-d543e84718e8 · outbound

This paper cites Cider: Consensus-based image descrip- tion evaluation.

Object-centric Video Question Answering with Visual Grounding and Referring Cider: Consensus-based image descrip- tion evaluation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:43.856383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:37.123600Z digest=sha256:9ffbc77d9395ccbd49d56b94b95fff91bd308b29af3f5950e8605431659d4ef6

Observation bc9be667-cf9b-423a-94e6-5f3465b07219 · outbound

This paper cites Ov-vis: Open-vocabulary video instance segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring Ov-vis: Open-vocabulary video instance segmentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:43.599672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:37.236690Z digest=sha256:842ef091d0324c7ffe0dcef8c7f309f52b409bebaee0d96b911c74a1f870c844

Observation a4c34e4a-92bc-43f1-8a2c-09516a9bef1c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Object-centric Video Question Answering with Visual Grounding and Referring Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:37.356424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:37.356424Z digest=sha256:ffbf5e84e999db0da3d7ed9654c6663b3f72e3bb75ce66140be15a8229aa60a0

Observation ff284c6b-ee5d-44f0-b956-a21b58c65482 · outbound

This paper cites Unidentified video objects: A benchmark for dense, open-world segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring Unidentified video objects: A benchmark for dense, open-world segmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:43.351712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:37.497702Z digest=sha256:4faf4fb47a61b307c09a1de5514def821d731c169fded1fd33fd5f17b4cd6aac

Observation 6a9aaa68-a589-4f06-9676-baa83d4ef10e · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Object-centric Video Question Answering with Visual Grounding and Referring InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:37.661534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:37.661534Z digest=sha256:6550f65f735241114f52600e4cc4a65d6eedd6904851146ef6b444799ad42578

Observation 177d8f52-5c82-4d46-86ca-5d8315692384 · outbound

This paper cites Internvideo2: Scaling foun- dation models for multimodal video understanding.

Object-centric Video Question Answering with Visual Grounding and Referring Internvideo2: Scaling foun- dation models for multimodal video understanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:43.127670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:37.795167Z digest=sha256:af00e659359c9eea214854dd011dac50dda411b012fd6b93f18dad00c73e6cad

Observation ff67cfcd-c3e3-4eb0-bb9c-114846a0b438 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Object-centric Video Question Answering with Visual Grounding and Referring Next-qa: Next phase of question-answering to explaining temporal actions

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.919131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:37.926811Z digest=sha256:c0bad536c905e7313ff42e237ea0281649709e1cf338f485fcef3c1c07017489

Observation 872524ac-7cfc-420a-a0ea-c0d3541c37e5 · outbound

This paper cites Visa: Reasoning video object segmen- tation via large language models.

Object-centric Video Question Answering with Visual Grounding and Referring Visa: Reasoning video object segmen- tation via large language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.729703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:38.024130Z digest=sha256:c4aa3b8a863f5336918e5e185e0b5e91afa0403059954035e84812e421566b3e

Observation 455e99f2-4c90-42d3-be2f-6e87c41904f0 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Object-centric Video Question Answering with Visual Grounding and Referring VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.109831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.109831Z digest=sha256:c293eadd1fd200163d90a858460539abd6c1b2fc8f56cead93f88205ac1a9979

Observation 44a2257f-8dea-4cd1-8131-6d0e80f5ee68 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Object-centric Video Question Answering with Visual Grounding and Referring Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.229151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.229151Z digest=sha256:b1abcf69100eb5e50079a1e3ca05fc21f9784c905b9c64902eb2bb3498b7036b

Observation 827c5a9c-da2c-4dd7-80b6-f910464b5d2e · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

Object-centric Video Question Answering with Visual Grounding and Referring LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.352762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.352762Z digest=sha256:7e41aa17d943959dda058b30d4dc89f5ad00b9cbb7b2cc29f880f85b3fbc7ec6

Observation cb2f858e-5f53-4f4b-b6dc-fac94a69787c · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Object-centric Video Question Answering with Visual Grounding and Referring mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.498533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.498533Z digest=sha256:a67e0cede97817d9488075d1946af6d34182fab689cb1b090fd80f1c40b226cb

Observation dd125f56-671f-4185-abee-9e1900f90ef0 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Object-centric Video Question Answering with Visual Grounding and Referring Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.603030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.603030Z digest=sha256:4db8f62a53caccc2e8cd7e5b1ed0592d1fb9de0d3771b14fe94b28cfe95ff1ff

Observation 450421f4-78e9-495d-b65b-092441bda764 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Object-centric Video Question Answering with Visual Grounding and Referring Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.745855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.745855Z digest=sha256:ab15f46672ad7a95f4c3703063b432d7ef7e0a0e886224f2f7c1c76a4548626a

Observation 972d9ea1-6637-4f61-8713-1686d8528fd1 · outbound

This paper cites Sa2va: Marrying sam2 with llava for dense grounded understanding of im- ages and videos.

Object-centric Video Question Answering with Visual Grounding and Referring Sa2va: Marrying sam2 with llava for dense grounded understanding of im- ages and videos

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.467851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:38.813448Z digest=sha256:cb13a2775978fd91a60eee6e41c1496449c447e592a8b57a621b7d3963814a00

Observation a545401d-1949-49a9-8bc6-1d7869e20b88 · outbound

This paper cites Os- prey: Pixel understanding with visual instruction tun- ing.

Object-centric Video Question Answering with Visual Grounding and Referring Os- prey: Pixel understanding with visual instruction tun- ing

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.256487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:38.918347Z digest=sha256:5c5d603ab0a7f6d79db298f0cf1222cf526ef11321d5ad11c42eb45d8e532193

Observation ccf18085-1bfe-4023-b3b1-542b44531262 · outbound

This paper cites Videorefer suite: Advancing spatial-temporal object understanding with video llm.

Object-centric Video Question Answering with Visual Grounding and Referring Videorefer suite: Advancing spatial-temporal object understanding with video llm

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.021520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:39.024526Z digest=sha256:236d498ca7d0c0b551dd5c7127c40243ce693d33f190491c761afa8ed5f1f5b2

Observation 025328a4-de54-48dd-bb8d-b7f8679de765 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Object-centric Video Question Answering with Visual Grounding and Referring VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.140269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.140269Z digest=sha256:4cb2114c8e688f63fcbed98a2437e3def23ead5f84850216786ec7bbc7566e4e

Observation b7f6f91c-70a2-46e0-adda-b72c587b9531 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Object-centric Video Question Answering with Visual Grounding and Referring Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.258349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.258349Z digest=sha256:5a3f3b634e0de876c43e73a4ed32c8769223457482e8b624a309b55426a93973

Observation 99f4e4f6-d6ea-447e-b0c2-7f198124978f · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Object-centric Video Question Answering with Visual Grounding and Referring LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.360265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.360265Z digest=sha256:90e592ea63a52d635a19fc15accf7ec8cebffd57971055f7de81100c1ba42d11

Observation b5ee245e-d564-4974-a412-7de86e962e88 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

Object-centric Video Question Answering with Visual Grounding and Referring GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.474732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.474732Z digest=sha256:c7c63de788085594379dd802617efbc08a655a214e56cae534432079569c7888

Observation 67cb2719-8b81-48f7-adeb-ca73373abd9c · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.606964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.606964Z digest=sha256:993e012a244ae30b97b3ba227a2d4236c286a53f9568b01b3ce7cc02b86b7425

Observation 0d770155-febd-4a35-a8c6-8d9ae90f717e · outbound

This paper cites Scene parsing through ade20k dataset.

Object-centric Video Question Answering with Visual Grounding and Referring Scene parsing through ade20k dataset

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:41.775584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:39.713362Z digest=sha256:f818f07949c8666127afae84e013767f1d36005432f23c5bb9da32f50042a86d

Observation f836cefc-0b8a-427a-9a92-f9483ff2bffc · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Object-centric Video Question Answering with Visual Grounding and Referring MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.834178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.834178Z digest=sha256:982e877922dfc91ad0878c65fc097cefc4afaa1c44ddea8963f4ccfa54816e17

Observation b64863f9-c6e2-441b-8d25-9f17e8265ac1 · outbound

This paper cites The complete list of used datasets in training is presented in Tab.

Object-centric Video Question Answering with Visual Grounding and Referring The complete list of used datasets in training is presented in Tab

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:41.512014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:39.976573Z digest=sha256:3187722bdac271bb95326672bb5755ec517a3bab4ba9e40d2b97a0c633dd9bef

Observation 129b29ba-459e-434b-8336-2757de218c6a · outbound

This paper cites an unresolved cited work.

Object-centric Video Question Answering with Visual Grounding and Referring Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:19:41.268991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:40.126377Z digest=sha256:fdb9e98ee85fff082e3daa287fcde3e8c094d51f5579e2a89a7ed743983b9c95

Observation b5b5aa3e-f768-44d4-b4c0-9b6077d8db28 · outbound

This paper cites an unresolved cited work.

Object-centric Video Question Answering with Visual Grounding and Referring Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:19:40.965085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:40.237501Z digest=sha256:644dfd07a80bfd1cb7d31684178ec30f2da9e3f0e93aeaa025a438f9599d451c

Observation c21f2be1-ac58-4b45-9031-bf31ee57b6f1 · outbound

This paper cites an unresolved cited work.

Object-centric Video Question Answering with Visual Grounding and Referring Unresolved cited work

Reference 223

Resolution
parse uncertain
raw_fallback, observed 2026-08-06T14:19:44.107775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T14:19:36.874703Z digest=sha256:fde67548a0c95fae42904161bcdb77f91907a1cb1bee0b64b01324e46f8794af

Pith citing papers

No inbound Pith citation observations are available.