Pith. sign in

Paper Citation Record · LEDGER

Object-centric Video Question Answering with Visual Grounding and Referring

As of 8 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2507.19599.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19599 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:19:40.237501Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved35
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f4a2b7bb-2e0c-4a06-b9fb-9f8b22e30d21 · outbound

This paper cites GPT-4 Technical Report.

Object-centric Video Question Answering with Visual Grounding and Referring GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:30.864771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:30.864771Z digest=sha256:46911b1080517fa44d40af54237f6eb145aab89f56dd1481c0d220f90dfbea2d

Observation 15a59f5c-3d8f-4f09-ab56-e0507cb122af · outbound

This paper cites Burst: A benchmark for unifying ob- ject recognition, segmentation and tracking in video.

Object-centric Video Question Answering with Visual Grounding and Referring Burst: A benchmark for unifying ob- ject recognition, segmentation and tracking in video

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:52.299469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:30.941003Z digest=sha256:4a6183be4430909a86a435d75047a6e2fa9c64b028df00448fbe2698062e7e9f

Observation 72cd616b-060b-4990-93c0-ea6b3f62f79b · outbound

This paper cites Qwen2.5-VL Technical Report.

Object-centric Video Question Answering with Visual Grounding and Referring Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:31.060322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:31.060322Z digest=sha256:af7e0a876d3f0edf29097bec9b65e174ca301376c416c3b4cb435aea19d01c8f

Observation 64102306-6b7f-4f91-80c3-f32aa1652cfa · outbound

This paper cites One token to seg them all: Lan- guage instructed reasoning segmentation in videos.

Object-centric Video Question Answering with Visual Grounding and Referring One token to seg them all: Lan- guage instructed reasoning segmentation in videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:52.160107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:31.203054Z digest=sha256:d5e6111cea7fa7d7c428ba27ade2ff5ebfbb8890b5b8ec0b5de9e2c02f97de54

Observation de2d2fdc-fc69-4e84-a048-286ac09946bf · outbound

This paper cites Xmem++: Production-level video segmentation from few annotated frames.

Object-centric Video Question Answering with Visual Grounding and Referring Xmem++: Production-level video segmentation from few annotated frames

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:52.024268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:31.309544Z digest=sha256:feaf014878b9d3222c211310ffab71870c5ae617005eb90862c7af03ad727976

Observation 3897ae52-368e-4cf8-a23e-9833ac40268c · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

Object-centric Video Question Answering with Visual Grounding and Referring Coco-stuff: Thing and stuff classes in context

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:51.766255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:31.414465Z digest=sha256:f2a445060d2de4bfe20c48eb5fd60d9ad8a3bdf32d286ebf2ac82142521c3875

Observation 6fca08f5-55d2-42a4-8ad2-c38d9365bb64 · outbound

This paper cites Vip-llava: Making large multi- modal models understand arbitrary visual prompts.

Object-centric Video Question Answering with Visual Grounding and Referring Vip-llava: Making large multi- modal models understand arbitrary visual prompts

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:51.512367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:31.529633Z digest=sha256:8d7db93a0ec96b86f3a2abd9e9a17591c388fb3ab4d30633631cefae84de5745

Observation 5d06ad8b-31d9-4502-86fc-bb077eb6f79d · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Object-centric Video Question Answering with Visual Grounding and Referring Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:31.659606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:31.659606Z digest=sha256:3175ed3e8e9fc6c4f86ab4cc686fe318b518e2e804349d862f811f02613db97d

Observation 1ee91e18-9a87-407e-9fd5-9213426db69f · outbound

This paper cites Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos.

Object-centric Video Question Answering with Visual Grounding and Referring Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:31.776850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:31.776850Z digest=sha256:a1e3ad87ae7ca6818642de93176c3aa02f37561090e239aa5faabc7b62864400

Observation c8d0e2f8-5dd9-4fac-8c69-b65e8e9c55ea · outbound

This paper cites Detect what you can: Detecting and representing objects using holistic models and body parts.

Object-centric Video Question Answering with Visual Grounding and Referring Detect what you can: Detecting and representing objects using holistic models and body parts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:51.309218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:31.883436Z digest=sha256:392014bba52e406187e0e34af3ad000a0ba08bfa5cf7cf59c8ea2a95df35e227

Observation 87cfd5da-20c0-4dd2-be86-c82accb6bea4 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Object-centric Video Question Answering with Visual Grounding and Referring How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:31.985474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:31.985474Z digest=sha256:803c5dbf164941dfeb07cf84e562bfbc6a9cf2308a3edb66b452a8025ec3b064

Observation 08867776-fa48-4a4a-afc0-7032580eddf9 · outbound

This paper cites Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model.

Object-centric Video Question Answering with Visual Grounding and Referring Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:51.085303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:32.101781Z digest=sha256:a4a4c7113789a0cc74b737a32536411e7192b8bb586feacc02aa0c1775d2adfa

Observation 467cceab-2790-4e27-8aab-b4c40e40fda3 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Object-centric Video Question Answering with Visual Grounding and Referring VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:32.183705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:32.183705Z digest=sha256:9a2e860a2dc52c1caf88b2ddb55998360504b79462523290a50c4ec9cb35f03f

Observation 6e4f8a82-6789-47d6-ab6f-b8fe37747739 · outbound

This paper cites Grounded question- answering in long egocentric videos.

Object-centric Video Question Answering with Visual Grounding and Referring Grounded question- answering in long egocentric videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:50.848944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:32.316400Z digest=sha256:481c347d29f575bee8fed2dfc2a87e3a6ea60786ff7dab2609ec6b5e14871d4f

Observation b9e99d39-f6dc-4276-9c58-823895421b22 · outbound

This paper cites Mevis: A large-scale bench- mark for video segmentation with motion expressions.

Object-centric Video Question Answering with Visual Grounding and Referring Mevis: A large-scale bench- mark for video segmentation with motion expressions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:50.598223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:32.471391Z digest=sha256:483383c89238bdfd51ee8e4a9535ce124817caddd6ee87de0536b029b5d501eb

Observation d3ab6386-277c-40ad-a059-8b69694a63cc · outbound

This paper cites Mose: A new dataset for video object segmentation in complex scenes.

Object-centric Video Question Answering with Visual Grounding and Referring Mose: A new dataset for video object segmentation in complex scenes

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:50.279506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:32.592931Z digest=sha256:8f681066ce209fbd0e746993b9d82dea1627b063c7ed3c849ede8313d4f1dbb7

Observation a69ad915-4348-4f40-b026-8ab69739055a · outbound

This paper cites LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:32.702316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:32.702316Z digest=sha256:4f0c91d16d76c54615d4eff44180697e132a3760386d9f5cd64e47b156a28488

Observation 615ca1a8-b839-481e-86be-2a949d326ed0 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Object-centric Video Question Answering with Visual Grounding and Referring LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:32.811828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:32.811828Z digest=sha256:3ed3653a686f94ecc8b3f8aa8018c9d061102cc39c36af752f6c2fd9435111eb

Observation 2a6d0135-259a-4498-af7e-5c2be488fbd3 · outbound

This paper cites GPT-4o System Card.

Object-centric Video Question Answering with Visual Grounding and Referring GPT-4o System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:32.918892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:32.918892Z digest=sha256:1c53d5a6379bf308580e5b63bfc5cc5bc489724930a398d02d34c623fe2ffb53

Observation a4d82382-ce65-4f2f-90f6-875055bd318e · outbound

This paper cites Cotracker3: Simpler and better point tracking by pseudo-labelling real videos.

Object-centric Video Question Answering with Visual Grounding and Referring Cotracker3: Simpler and better point tracking by pseudo-labelling real videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:49.949105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:33.036045Z digest=sha256:529dda589e3cc988742936e75de7da842204e894b4c85a5277a2e8d3cd749ad5

Observation 8e024cfc-dbb8-4abf-99fb-9249ed304db2 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Object-centric Video Question Answering with Visual Grounding and Referring Referitgame: Referring to objects in photographs of natural scenes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:49.720019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:33.140267Z digest=sha256:b92a199705ec0fe0aeb03997b47826609387805b0c735d4fb717768d77a539f5

Observation 29c9ab82-2837-4706-add4-46a147faccd3 · outbound

This paper cites Video object segmentation with language referring ex- pressions.

Object-centric Video Question Answering with Visual Grounding and Referring Video object segmentation with language referring ex- pressions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:49.479956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:33.238590Z digest=sha256:90f1d143cc8aa4656fe12ba4bd05a63559ea778e1b6f6fcc43c097bafe15b43b

Observation a306dbb0-bdd4-4aa2-8e53-6b96dab7dc20 · outbound

This paper cites Segment anything.

Object-centric Video Question Answering with Visual Grounding and Referring Segment anything

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:49.210901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:33.380141Z digest=sha256:20073bb5d4644d1985a27015598004324bdbd7dc12096a5f50ad39501890db0c

Observation 79e37907-115d-4a89-8f14-eeb865c5259e · outbound

This paper cites Grounding language models to images for multi- modal inputs and outputs.

Object-centric Video Question Answering with Visual Grounding and Referring Grounding language models to images for multi- modal inputs and outputs

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:48.929539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:33.502374Z digest=sha256:93b7e1a92da9ff0f0249168a616f0b8378ad5a4e1019bb3f7100bb8c32358db3

Observation 58447ff6-88b7-4e0d-ab6b-135f1d9fc163 · outbound

This paper cites Generating images with multimodal language models.

Object-centric Video Question Answering with Visual Grounding and Referring Generating images with multimodal language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:48.636508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:33.605622Z digest=sha256:732768a3229aa92720007c50efd2d081986c9743040ee8e4f918c92daf930e9c

Observation d3f1ad95-2d6d-41cb-b41c-26b35e07aed9 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Object-centric Video Question Answering with Visual Grounding and Referring Lisa: Reasoning segmentation via large language model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:48.380741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:33.708532Z digest=sha256:b025c80a75cc357c51dea26c6387e91afad2a02f3d6e9afbf0707f87b00e8531

Observation f50cf220-2210-43f9-852c-2d9ad08ed121 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

Object-centric Video Question Answering with Visual Grounding and Referring MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:33.841194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:33.841194Z digest=sha256:1517a978bb384bb4a08e0db6ace9c1ae7c1abd9a2f5eac89eb6a52626182cd8d

Observation 95ea52ed-4eec-4bce-990b-893c9990c20e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.009655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.009655Z digest=sha256:60e71d019f836694955b2a4d0e05f56e83c0f5bfe8ab4ea4a3e2dc2d3e87e04a

Observation f7cad20f-21af-4df1-92e8-4210147ffecf · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.124422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.124422Z digest=sha256:1b690aa675409f5995713ab0ee5d30a8ccb7aef3e6390a6bd42469e2e9f0b5e3

Observation 172277c2-04cd-4578-9645-535d88e70850 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Object-centric Video Question Answering with Visual Grounding and Referring Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:48.097986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:34.242376Z digest=sha256:32d4ae5a5c2e40e0749773b0972e2e0eec9bf1e77c5f3a37a15d50822a90aab1

Observation 3a772343-879c-4ae3-8d14-0f5d729134dc · outbound

This paper cites Towards Robust Referring Video Object Segmentation with Cyclic Relational Consensus.

Object-centric Video Question Answering with Visual Grounding and Referring Towards Robust Referring Video Object Segmentation with Cyclic Relational Consensus

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.362625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.362625Z digest=sha256:930835d58839794515468dc0ab28f405d185a527698ef36a4f61344b70283675

Observation 0b453265-6047-4a43-867c-f4127e8ee3d2 · outbound

This paper cites Describe anything: Detailed localized image and video captioning.

Object-centric Video Question Answering with Visual Grounding and Referring Describe anything: Detailed localized image and video captioning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:47.846987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:34.543370Z digest=sha256:f52efdb27ef226ed41af36fefe3a1ada326b48d6a93c66de7b6b9c4e043833e9

Observation b81e087f-251f-40a7-ae09-ceb99b331e7a · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Object-centric Video Question Answering with Visual Grounding and Referring Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.701074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.701074Z digest=sha256:3f959a6a02e7076fe410c3dcc823f17fc88524a104b24181337267668fdfa2de

Observation 1237dd00-ff9e-45c2-adab-a977ca078d9e · outbound

This paper cites Rouge: A package for automatic eval- uation of summaries.

Object-centric Video Question Answering with Visual Grounding and Referring Rouge: A package for automatic eval- uation of summaries

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:47.568836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:34.792454Z digest=sha256:56a5414684008e7a51eb138fb52422c4c1833faea688bc3a907ecc68f316ef78

Observation a155d143-04a0-4678-ac25-5bd7cbdcefed · outbound

This paper cites Glus: Global-local reasoning unified into a single large language model for video segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring Glus: Global-local reasoning unified into a single large language model for video segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:47.292410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:34.885264Z digest=sha256:046c9b08c70008af40265b84eb7a5831da3bc0005f6f6e6801abcb336abc7569

Observation a14696b6-30a5-4cce-8ba6-08e350b0cc02 · outbound

This paper cites Visual instruction tuning.

Object-centric Video Question Answering with Visual Grounding and Referring Visual instruction tuning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:47.019287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:34.986936Z digest=sha256:ee064ba958f0426a19b6e7cb565d29688c5ce5e7c9f729ef258a1bca6156995b

Observation 47c75236-174c-43bb-82be-0cc419d7d744 · outbound

This paper cites Lamra: Large multimodal model as your ad- vanced retrieval assistant.

Object-centric Video Question Answering with Visual Grounding and Referring Lamra: Large multimodal model as your ad- vanced retrieval assistant

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:46.770684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:35.099805Z digest=sha256:ee946e69738a5aff6bd13fb31bb397f1dc2eb2fbebd0bc84f7d6f3ea43e21565

Observation 2dc44f75-013a-403d-be4b-9560b923ed50 · outbound

This paper cites Decoupled Weight Decay Regularization.

Object-centric Video Question Answering with Visual Grounding and Referring Decoupled Weight Decay Regularization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:35.242260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:35.242260Z digest=sha256:a57d185844a9f6eb7e67e173844ebd91000d0d51b8ae8928f2b0e3dbab8dd526

Observation aa765c0b-ad33-427e-93f2-d3503289bb97 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Object-centric Video Question Answering with Visual Grounding and Referring Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:35.344758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:35.344758Z digest=sha256:c008d655fad5b9b50c1eb4d0de8f2758a2c8d37166587c9946b6c5abc96fe4ea

Observation 6cbe70e5-59c8-4e31-98c6-eee78a3e78ec · outbound

This paper cites Generation and comprehension of unambiguous ob- ject descriptions.

Object-centric Video Question Answering with Visual Grounding and Referring Generation and comprehension of unambiguous ob- ject descriptions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:46.495180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:35.462523Z digest=sha256:d2fdd5ddc33e910f522f17cfc7badf7e85b9d2a084c4a6b096f1e56a85750fa6

Observation 1ea753bb-c75a-4a3c-af41-dcc1cd2756be · outbound

This paper cites Large-scale video panoptic segmentation in the wild: A benchmark.

Object-centric Video Question Answering with Visual Grounding and Referring Large-scale video panoptic segmentation in the wild: A benchmark

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:46.303905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:35.567110Z digest=sha256:5d0dd601336dfcad676de8a94d51177d03cac8baed9b74ab333ba6c11cdd6c84

Observation 18aacfb3-e0e4-4215-bf55-5ba9bff3d63f · outbound

This paper cites V-net: Fully convolutional neural networks for volumetric medical image segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring V-net: Fully convolutional neural networks for volumetric medical image segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:46.043914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:35.730518Z digest=sha256:63138f49c09d0dbdf5710b5a0d5e88495b77ff628ea89c7de05b95bcc351dd6a

Observation 0ed6dda5-284f-482d-af20-571d707c01f7 · outbound

This paper cites Bleu: A method for automatic eval- uation of machine translation.

Object-centric Video Question Answering with Visual Grounding and Referring Bleu: A method for automatic eval- uation of machine translation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:45.806001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:35.851977Z digest=sha256:330fe398928c3ffc17e6abee98e765864d709f7d8f10982865650f2c8c0e7f2d

Observation 4fc8125a-1305-44cd-931f-26d4ddd572c4 · outbound

This paper cites Perception test: A diagnostic bench- mark for multimodal video models.

Object-centric Video Question Answering with Visual Grounding and Referring Perception test: A diagnostic bench- mark for multimodal video models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:45.588351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:35.974409Z digest=sha256:ce471d67fd3df6fd61329c79685ac3f2d4fdda092915ed930dbc51d93fb1a0f3

Observation 07e3f2f3-ce0a-4fc7-b345-467de8bfbfab · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Object-centric Video Question Answering with Visual Grounding and Referring Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:36.065603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:36.065603Z digest=sha256:ef36f01a7c67f3a53b9ba5aad3f89bf9a719ee4c380633fe745ef51a1398c1a0

Observation fa54ea9a-a496-4f7e-aace-b97d167440a3 · outbound

This paper cites Occluded video in- stance segmentation: A benchmark.

Object-centric Video Question Answering with Visual Grounding and Referring Occluded video in- stance segmentation: A benchmark

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:45.342672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:36.172270Z digest=sha256:9fa1702f2eb668a0719c91ec4638aa0c819d02ebb3afa38becd4fc1b460f6164

Observation c5968d3d-6c7a-4dee-8860-bda3c817aa7e · outbound

This paper cites Artemis: Towards referential un- derstanding in complex videos.

Object-centric Video Question Answering with Visual Grounding and Referring Artemis: Towards referential un- derstanding in complex videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:45.082859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:36.327772Z digest=sha256:c6286ab6ef752c10b02d652613bd8814cce1ee9b789f1a7c1cb825e1e6348e1a

Observation fc7d881a-c0f5-4e22-b453-0570dc9067a9 · outbound

This paper cites Paco: Parts and attributes of common objects.

Object-centric Video Question Answering with Visual Grounding and Referring Paco: Parts and attributes of common objects

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:44.828041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:36.435318Z digest=sha256:85eee6f65c0013dcec0133720e9cba1083946c2b584c2c75c343eb8381f86984

Observation ca092a3f-da64-4fc2-aee2-c549055601d9 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Object-centric Video Question Answering with Visual Grounding and Referring SAM 2: Segment Anything in Images and Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:36.552673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:36.552673Z digest=sha256:24a17597c933145b3dea31efb48dc06c4a831d06a6866754dd1d7393dcc5be4d

Observation 5c03be5e-3b00-4980-88c4-78066345f1ff · outbound

This paper cites Hiera: A hierarchical vision transformer without the bells-and-whistles.

Object-centric Video Question Answering with Visual Grounding and Referring Hiera: A hierarchical vision transformer without the bells-and-whistles

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:44.575409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:36.644218Z digest=sha256:a353b4262815cf857875d1cf732f7bd3b4e6ae78c5a9153ed4e72673b8af60ae

Observation 297c4461-aa4a-413f-8cf5-41104126d5c6 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

Object-centric Video Question Answering with Visual Grounding and Referring Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:44.344740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:36.749964Z digest=sha256:1097704d24d2147e9bc101e097bf7ba5a2ae40783335c728e64de7fb1ee666db

Observation 624e5f94-9f0c-470f-acef-f2167f960c44 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Object-centric Video Question Answering with Visual Grounding and Referring Emu: Generative Pretraining in Multimodality

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:36.979072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:36.979072Z digest=sha256:143c0d0c2a9fa5e602fdbc45721b674023a1270818e4605f94729caf21d288d6

Observation 1b464f74-7deb-44c6-8ad2-d543e84718e8 · outbound

This paper cites Cider: Consensus-based image descrip- tion evaluation.

Object-centric Video Question Answering with Visual Grounding and Referring Cider: Consensus-based image descrip- tion evaluation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:43.856383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:37.123600Z digest=sha256:0b1aecdeea9c9862d2d8531a83399ecf70848b46fbc5661672be8baf25dad309

Observation bc9be667-cf9b-423a-94e6-5f3465b07219 · outbound

This paper cites Ov-vis: Open-vocabulary video instance segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring Ov-vis: Open-vocabulary video instance segmentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:43.599672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:37.236690Z digest=sha256:6a059da8f6bd5c2bb345a7e6aee5ddead0a15b0db96c5efdae6d296ce6ffaf86

Observation a4c34e4a-92bc-43f1-8a2c-09516a9bef1c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Object-centric Video Question Answering with Visual Grounding and Referring Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:37.356424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:37.356424Z digest=sha256:a98ccf045f295f96baaa1625331739daf4316d9ea2606005f2cae0207a4c23b4

Observation ff284c6b-ee5d-44f0-b956-a21b58c65482 · outbound

This paper cites Unidentified video objects: A benchmark for dense, open-world segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring Unidentified video objects: A benchmark for dense, open-world segmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:43.351712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:37.497702Z digest=sha256:0a54dd24cd984764efb5ecd442dcede5ca7fe18b3947d61766e01e2605a0ef70

Observation 6a9aaa68-a589-4f06-9676-baa83d4ef10e · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Object-centric Video Question Answering with Visual Grounding and Referring InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:37.661534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:37.661534Z digest=sha256:1a26e861ea6d0d622a12a03a47859503b1309c5ea93a827a2ddeff6f2487a892

Observation 177d8f52-5c82-4d46-86ca-5d8315692384 · outbound

This paper cites Internvideo2: Scaling foun- dation models for multimodal video understanding.

Object-centric Video Question Answering with Visual Grounding and Referring Internvideo2: Scaling foun- dation models for multimodal video understanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:43.127670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:37.795167Z digest=sha256:5da8b2d7c1c9e0e93fb182bd079fd494bd4c0021de3fe50358aa6c8f31e8c6b8

Observation ff67cfcd-c3e3-4eb0-bb9c-114846a0b438 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Object-centric Video Question Answering with Visual Grounding and Referring Next-qa: Next phase of question-answering to explaining temporal actions

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.919131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:37.926811Z digest=sha256:1634bc323bb9390aaa27fa430a20056c50d96da1342c0464fc7d620375ed108f

Observation 872524ac-7cfc-420a-a0ea-c0d3541c37e5 · outbound

This paper cites Visa: Reasoning video object segmen- tation via large language models.

Object-centric Video Question Answering with Visual Grounding and Referring Visa: Reasoning video object segmen- tation via large language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.729703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:38.024130Z digest=sha256:eab8abe53d2e4a39663cdf781b8afad36419e6339c21ed9ba003ad4d3d517bd6

Observation 455e99f2-4c90-42d3-be2f-6e87c41904f0 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Object-centric Video Question Answering with Visual Grounding and Referring VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.109831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.109831Z digest=sha256:ef250ef1efb54bcca6a96d5cbe7ef61c7d4b964b63c6ab253dd6333db4f0df47

Observation 44a2257f-8dea-4cd1-8131-6d0e80f5ee68 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Object-centric Video Question Answering with Visual Grounding and Referring Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.229151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.229151Z digest=sha256:935b90dbe8f48742f604881f6bd90141226b51b17b4a3e5324d2354c5666917b

Observation 827c5a9c-da2c-4dd7-80b6-f910464b5d2e · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

Object-centric Video Question Answering with Visual Grounding and Referring LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.352762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.352762Z digest=sha256:bb341526b8838f41af6d44c29b00c6acd82c33c7f53a3573f0fb3065a59d7bbc

Observation cb2f858e-5f53-4f4b-b6dc-fac94a69787c · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Object-centric Video Question Answering with Visual Grounding and Referring mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.498533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.498533Z digest=sha256:b81563a96b4b7e60679f024878a3f4310cca6bf825c115dcb21f7d8816ac3245

Observation dd125f56-671f-4185-abee-9e1900f90ef0 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Object-centric Video Question Answering with Visual Grounding and Referring Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.603030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.603030Z digest=sha256:09d287e5c0dcb169855234e70a4225dc216093548a8c9651b3d312e353557063

Observation 450421f4-78e9-495d-b65b-092441bda764 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Object-centric Video Question Answering with Visual Grounding and Referring Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.745855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.745855Z digest=sha256:ee7d03ed08f2acae597d0df4f474f05e4c2c3811578c8f2a6f73f1f296bba0a4

Observation 972d9ea1-6637-4f61-8713-1686d8528fd1 · outbound

This paper cites Sa2va: Marrying sam2 with llava for dense grounded understanding of im- ages and videos.

Object-centric Video Question Answering with Visual Grounding and Referring Sa2va: Marrying sam2 with llava for dense grounded understanding of im- ages and videos

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.467851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:38.813448Z digest=sha256:a48f0af58d56f20cf6d64db12020ff16a2d9817f3c85798982429eed1ba30f5b

Observation a545401d-1949-49a9-8bc6-1d7869e20b88 · outbound

This paper cites Os- prey: Pixel understanding with visual instruction tun- ing.

Object-centric Video Question Answering with Visual Grounding and Referring Os- prey: Pixel understanding with visual instruction tun- ing

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.256487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:38.918347Z digest=sha256:f6e5d800f393dbbbc7c87d250319e97ee4aadacc0130fa632454dafb75070ec8

Observation ccf18085-1bfe-4023-b3b1-542b44531262 · outbound

This paper cites Videorefer suite: Advancing spatial-temporal object understanding with video llm.

Object-centric Video Question Answering with Visual Grounding and Referring Videorefer suite: Advancing spatial-temporal object understanding with video llm

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.021520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:39.024526Z digest=sha256:e4fe67f26ded09c857c5da5f72e13cf1402f89f0e35b52096f0ce3d3f5faefd5

Observation 025328a4-de54-48dd-bb8d-b7f8679de765 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Object-centric Video Question Answering with Visual Grounding and Referring VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.140269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.140269Z digest=sha256:8e59077a3c255f9e829b07f46a1f7ef800372a8b0afffcf2261f2b0f9a9db56f

Observation b7f6f91c-70a2-46e0-adda-b72c587b9531 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Object-centric Video Question Answering with Visual Grounding and Referring Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.258349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.258349Z digest=sha256:421136ba366b1bf69330670b94c77007a41a577369ebfe21d335df9c902e223f

Observation 99f4e4f6-d6ea-447e-b0c2-7f198124978f · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Object-centric Video Question Answering with Visual Grounding and Referring LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.360265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.360265Z digest=sha256:55db6fcd0dc04c77c01add3bb6a32c7d02a23416bef6490ff48e9e660a5993fa

Observation b5ee245e-d564-4974-a412-7de86e962e88 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

Object-centric Video Question Answering with Visual Grounding and Referring GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.474732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.474732Z digest=sha256:f2424774d6bc5a7db39d76be03180c0d248e4e095805f29941ee9e6c0441a29a

Observation 67cb2719-8b81-48f7-adeb-ca73373abd9c · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.606964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.606964Z digest=sha256:d158ca01fe4b213ab3b585943cd8dc544f0887f56728e79a7fed26577bbfe241

Observation 0d770155-febd-4a35-a8c6-8d9ae90f717e · outbound

This paper cites Scene parsing through ade20k dataset.

Object-centric Video Question Answering with Visual Grounding and Referring Scene parsing through ade20k dataset

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:41.775584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:39.713362Z digest=sha256:2caa6440720a8f5dfb577465ed50d4a883b69b0f28acc64b128d6f7dcbbd3f35

Observation f836cefc-0b8a-427a-9a92-f9483ff2bffc · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Object-centric Video Question Answering with Visual Grounding and Referring MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.834178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.834178Z digest=sha256:7d9a15a029e5b446674588c6f91315caae313d842024c828920c21d15da5be47

Observation b64863f9-c6e2-441b-8d25-9f17e8265ac1 · outbound

This paper cites The complete list of used datasets in training is presented in Tab.

Object-centric Video Question Answering with Visual Grounding and Referring The complete list of used datasets in training is presented in Tab

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:41.512014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:39.976573Z digest=sha256:6cc80133327fafa55a88b524e558a3708a83a83389f5da75e13886f0418dc80f

Observation 129b29ba-459e-434b-8336-2757de218c6a · outbound

This paper cites an unresolved cited work.

Object-centric Video Question Answering with Visual Grounding and Referring Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:19:41.268991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:40.126377Z digest=sha256:59459fe26c9abd13c9913643401f246cdf61ee787ea39fb97468585e0bb2d760

Observation b5b5aa3e-f768-44d4-b4c0-9b6077d8db28 · outbound

This paper cites an unresolved cited work.

Object-centric Video Question Answering with Visual Grounding and Referring Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:19:40.965085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:40.237501Z digest=sha256:29d0e475956afb72f6ee3ec36ab1aaf1bdfd188aee53f079dc6285780763ce2f

Observation c21f2be1-ac58-4b45-9031-bf31ee57b6f1 · outbound

This paper cites an unresolved cited work.

Object-centric Video Question Answering with Visual Grounding and Referring Unresolved cited work

Reference 223

Resolution
parse uncertain
raw_fallback, observed 2026-08-06T14:19:44.107775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:19:36.874703Z digest=sha256:b1feafe11c074acaf0bbe95acf085bbf5c81f799e17ba001b0544d1397ad276b

Pith citing papers

No inbound Pith citation observations are available.