Pith. sign in

Paper Citation Record · LEDGER

VoCap: Video Object Captioning and Segmentation from Any Prompt

As of 20 August 2026, this Paper Citation Record lists 100 of 110 outbound references and 2 inbound Pith citation observations for arXiv:2508.21809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21809 v1

Coverage vector

measured 100 of 110 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:01:18.047778Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T00:56:18.867382Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.698319Z

Reference resolution

100 of 110 outbound references displayed

  • verified exact3
  • verified fuzzy44
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9177577e-76bf-47c2-b879-04cb54ee956f · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

VoCap: Video Object Captioning and Segmentation from Any Prompt Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:15.691477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:15.691477Z digest=sha256:92f165ab2d95feead7e16eff62a08b4e5c30f4d4edc4e1e61747797aaf5bf7df

Observation 4b28fbf4-cae5-49e4-b130-0ecb18ca97cd · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

VoCap: Video Object Captioning and Segmentation from Any Prompt Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:15.746526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:15.746526Z digest=sha256:454d632ebd3a3151c8a4e747a52c255e9c3c7d9777f85d7adde169432ea5ca41

Observation 0e08dbdd-2edc-4f41-bce3-00a182780fd5 · outbound

This paper cites Context r-cnn: Long term temporal context for per-camera object detection.

VoCap: Video Object Captioning and Segmentation from Any Prompt Context r-cnn: Long term temporal context for per-camera object detection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:15.824077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:15.824077Z digest=sha256:a36184e98241cc535caf7ed9e04e742bba443557e4bbb5d85425751bc7b71eef

Observation 2874a56a-9b25-4054-b70e-13bcec7da245 · outbound

This paper cites JAX: composable transformations of Python+NumPy programs, 2018.

VoCap: Video Object Captioning and Segmentation from Any Prompt JAX: composable transformations of Python+NumPy programs, 2018

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:15.880769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:15.880769Z digest=sha256:4c87a98e5bf36ae263e43cd4208bc625a3bf73dc75523f19051333416a568dc4

Observation 1f439b7b-b054-4b52-8a43-5bb14697f2f5 · outbound

This paper cites The 2019 DAVIS Challenge on VOS: Unsupervised Multi-Object Segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt The 2019 DAVIS Challenge on VOS: Unsupervised Multi-Object Segmentation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:15.973996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:15.973996Z digest=sha256:d679158351a49a668e5f35855173ffb865370fe6296ef5f512b5050bed4788de

Observation 1d125a6c-7184-413c-8e2a-93cdfe3cb89e · outbound

This paper cites nuscenes: A multimodal dataset for autonomous driving.

VoCap: Video Object Captioning and Segmentation from Any Prompt nuscenes: A multimodal dataset for autonomous driving

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.063128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.063128Z digest=sha256:c65239e5cbc099c0b783b360d4c66f26b4e7dd8c00f454de6d7abd42ab7ddede

Observation df8eea38-b958-46b0-88d4-aedcbe49966f · outbound

This paper cites Stablevideo: Text-driven consistency-aware diffusion video editing.

VoCap: Video Object Captioning and Segmentation from Any Prompt Stablevideo: Text-driven consistency-aware diffusion video editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.133523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.133523Z digest=sha256:cefe855404e177481592066ef956d8114dda244799b18052433f49b7f7f89116

Observation 7b95b0c0-f4fe-4c59-9daf-220cedfc8e0e · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

VoCap: Video Object Captioning and Segmentation from Any Prompt PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.184646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.184646Z digest=sha256:a9aae018b7192e8bf6de2e0e5219d20c2930c8e14cade723dbb53dddcabd3646

Observation b58a7584-aef2-4cc0-8aa1-b51c705e54b3 · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Per-pixel classification is not all you need for semantic segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.263063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.263063Z digest=sha256:2f99cc7b9deabd41498862557c3631378a7721b63232113301832ad4a59e5151

Observation c1e9961d-3d4b-4c70-bf9a-e5c47ce592f4 · outbound

This paper cites Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model.

VoCap: Video Object Captioning and Segmentation from Any Prompt Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.348008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.348008Z digest=sha256:af0acb66aebc84092416a1b9e036472bbc2e64b8dbab6de1c831e69dc1faac6c

Observation fd00fb5a-7009-4953-9119-15a1db801fea · outbound

This paper cites Putting the object back into video object segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Putting the object back into video object segmentation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.434138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.434138Z digest=sha256:1a8e0d4232367c258a9e1b8088225a32305a3c52c9742595b54d67fbb81749e9

Observation d3d34e96-2a6a-4207-a531-14fd3f85fee9 · outbound

This paper cites Segment and Track Anything.

VoCap: Video Object Captioning and Segmentation from Any Prompt Segment and Track Anything

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.517409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.517409Z digest=sha256:34d8b58f41e2fe9e0de224bb3442fbb16a31813260ceca2cc0ae8fbcaa7c017a

Observation 59ef5bd6-536b-4f9c-a433-c3cbf1968cd2 · outbound

This paper cites Find first, track next: Decoupling identification and propagation in referring video object segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Find first, track next: Decoupling identification and propagation in referring video object segmentation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.597953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.597953Z digest=sha256:9eee4bfbfbe5b0208b0a5018c0bee27a55c38ce5b7a0d033bc1265f42459ba20

Observation b6eb5932-be1c-44b8-927a-9f6c338c64f2 · outbound

This paper cites Ow-viscaptor: Abstractors for open-world video instance segmentation and captioning.

VoCap: Video Object Captioning and Segmentation from Any Prompt Ow-viscaptor: Abstractors for open-world video instance segmentation and captioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.689430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.689430Z digest=sha256:da041fee0f237c103a1d9ca45360f776008a3fdb6225376b985e7234d97dfaeb

Observation 721f38f1-5d3f-41d8-88f2-94a18394aa17 · outbound

This paper cites Scenic: A jax library for computer vision research and beyond.

VoCap: Video Object Captioning and Segmentation from Any Prompt Scenic: A jax library for computer vision research and beyond

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.748592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.748592Z digest=sha256:bdf9ea6bb1c19fbdb980a94ef10f05871c594fe86d9c8878624542d86857a1a7

Observation 16a7ea64-fec1-4e92-9025-369d3f0ecf66 · outbound

This paper cites Memsam: Taming segment anything model for echocardiography video segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Memsam: Taming segment anything model for echocardiography video segmentation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.827030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.827030Z digest=sha256:01dc7aa8f842663908f5b26d61969ca00bbe395b9f4b586dd39632afb8dfa715

Observation 508e5e38-71f8-4802-a7f4-7bda04a4eca4 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

VoCap: Video Object Captioning and Segmentation from Any Prompt Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.886981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.886981Z digest=sha256:7a3555eeeaac0aa8bd27ccdc1645d0b6da171e2ac53068537b0b2ca07d3ac493

Observation 45a4fb90-4416-4353-bf32-56a713abf84c · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

VoCap: Video Object Captioning and Segmentation from Any Prompt Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.961237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.961237Z digest=sha256:872527e9d3d3ff0c4215be064654efc840f875ac2d4e1c32a75f92330911da41

Observation 00dd6e78-bbe7-4fe8-9136-235ab55110a5 · outbound

This paper cites MOSE: A new dataset for video object segmentation in complex scenes.

VoCap: Video Object Captioning and Segmentation from Any Prompt MOSE: A new dataset for video object segmentation in complex scenes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.022115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.022115Z digest=sha256:81d837b754ff39be18f144692df5ffbb3935ea53e7eaddf8251097134fadcc87

Observation f67cbf05-cbd4-482f-920d-67c153d78ab6 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

VoCap: Video Object Captioning and Segmentation from Any Prompt An image is worth 16x16 words: Transformers for image recognition at scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.098094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.098094Z digest=sha256:389929136944e09bf4e6e6cb6908ecc80666a373763832f556907907513c55c9

Observation 6747a0f9-37f7-4fcf-9395-2915513161f3 · outbound

This paper cites EVA-02: A Visual Representation for Neon Genesis.

VoCap: Video Object Captioning and Segmentation from Any Prompt EVA-02: A Visual Representation for Neon Genesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.183470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.183470Z digest=sha256:858ae847c0ab37bdfc1fae1582e76cc6bb5d819babf5128155f8492a655861ab

Observation 5b629c41-94ce-44c9-8d93-a6c4abe62451 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VoCap: Video Object Captioning and Segmentation from Any Prompt Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.239917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.239917Z digest=sha256:3259aa0a90369652cf3351d0855ff4f8076a41330f9c94474053964d8b9dd5b7

Observation 0b8176bb-ec41-460f-a59b-d8711ef0de9e · outbound

This paper cites VideoSAM: Open-World Video Segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt VideoSAM: Open-World Video Segmentation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:01:18.703058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.298703Z digest=sha256:136e57804fd8f50b35e937e874563a8f3f4af94b3842d0c1a56d60df28af7e0f

Observation 6d7c230e-35c5-4568-906a-76a94457ba27 · outbound

This paper cites Masked autoencoders are scalable vision learners.

VoCap: Video Object Captioning and Segmentation from Any Prompt Masked autoencoders are scalable vision learners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.366945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.366945Z digest=sha256:a128d812632eb9c968b23f0306c51d1e134068e39515d492fb0f9a4fd3379fc9

Observation 6702ecee-8c0c-42ba-a0f7-6d2634eee613 · outbound

This paper cites Decoupling static and hierarchical motion perception for referring video segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Decoupling static and hierarchical motion perception for referring video segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.448996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.448996Z digest=sha256:9bab0a261413e2dce31c86ff427ff0f9081e5d6e9b2659ed0f224ae46de81e46

Observation 1c6717fd-d5db-4a7e-a5e3-b13941ee6882 · outbound

This paper cites Instruct-imagen: Image generation with multi-modal instruction.

VoCap: Video Object Captioning and Segmentation from Any Prompt Instruct-imagen: Image generation with multi-modal instruction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.502682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.502682Z digest=sha256:8bbd62c81965596cdffc15cf9d5b23463c9dd3b82c40236128320512f0a3a4c4

Observation 259459b1-63b5-4526-a464-f708dd6b1677 · outbound

This paper cites Segment and caption anything.

VoCap: Video Object Captioning and Segmentation from Any Prompt Segment and caption anything

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.553355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.553355Z digest=sha256:9b7b98322628cf4a271fc19d9f3ad0128ed21b8cea315ade3e1f2d150fa16414

Observation 33d9da2d-f6c1-4ba2-b1bf-fd554fb63942 · outbound

This paper cites A better use of audio-visual cues: Dense video captioning with bi-modal transformer.

VoCap: Video Object Captioning and Segmentation from Any Prompt A better use of audio-visual cues: Dense video captioning with bi-modal transformer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.587039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.587039Z digest=sha256:32fb6dd9d930239dd95baef4cd161604f9d889b9dc7fbec43a92dce51c568b8d

Observation 81b991e6-60d9-40a5-ba43-cdb91bcaa6b7 · outbound

This paper cites Perceiver: General perception with iterative attention.

VoCap: Video Object Captioning and Segmentation from Any Prompt Perceiver: General perception with iterative attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.595806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.595806Z digest=sha256:3115d9b2da175cfd3eb5b46df6164f9fe2fe7ae1ed673ebd6443660850803873

Observation 075a4d78-04a8-40e2-b57e-23de063ff4f5 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

VoCap: Video Object Captioning and Segmentation from Any Prompt Scaling up visual and vision-language representation learning with noisy text supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.602222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.602222Z digest=sha256:f391cfc5e6883034c50f5a3e4afd30249ebda895e5e63b207ecf7cc383151c2f

Observation 6c3529bb-aa1e-468e-8f57-7a1d95d60664 · outbound

This paper cites Densecap: Fully convolutional localization networks for dense captioning.

VoCap: Video Object Captioning and Segmentation from Any Prompt Densecap: Fully convolutional localization networks for dense captioning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.610712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.610712Z digest=sha256:b10accb1008af417150338027aae05bf37a34b593c0ef9d815897b3a2f8e6d1d

Observation c77ef26a-3202-4983-88b7-8ebf93db0f22 · outbound

This paper cites Kanani, Sriparna Saha, and Pushpak Bhattacharyya.

VoCap: Video Object Captioning and Segmentation from Any Prompt Kanani, Sriparna Saha, and Pushpak Bhattacharyya

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.621165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.621165Z digest=sha256:de50e310e17d7bece4706e71bae0573c2b0b80259a09cabab4eba60528571f35

Observation 2069cbfc-56f4-47ec-9832-55555d20f88a · outbound

This paper cites Video object segmentation with language referring expressions.

VoCap: Video Object Captioning and Segmentation from Any Prompt Video object segmentation with language referring expressions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.627534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.627534Z digest=sha256:8bad308ce0b192d0a0e9adf393b9adbf5e7f40d4d6bdfa5cd3b5f008dbd10bbb

Observation a6c63d63-bd18-4096-8acf-48f1bf76e627 · outbound

This paper cites Video panoptic segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Video panoptic segmentation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.633229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.633229Z digest=sha256:40f53da257c35af17d97d73457a453d0ff8c5ad06e183c8e84bc713b540774f9

Observation c0e447a3-09f0-47d3-88bc-0a9aaaa470be · outbound

This paper cites Segment anything.

VoCap: Video Object Captioning and Segmentation from Any Prompt Segment anything

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.639259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.639259Z digest=sha256:990313e307a8c19d586513fac9c23c145c1bd79e755c18d922ec87dfa01b16cc

Observation f4ff2a8b-c79d-417b-8f09-65ed0a8c227c · outbound

This paper cites Dense-captioning events in videos.

VoCap: Video Object Captioning and Segmentation from Any Prompt Dense-captioning events in videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.644666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.644666Z digest=sha256:d2415c016991fc9aefc48bc075c4e4b9ff5798f1d45cbf23f82c44f65b501cf6

Observation eac9522b-363c-4045-8bff-59058700cd60 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

VoCap: Video Object Captioning and Segmentation from Any Prompt Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.541463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.650657Z digest=sha256:9bb6abb5ec84cc9e37f8feea85132e9d61e0ec01ee5d07fb56fd0aee32e4473b

Observation 02709fa6-6806-4d60-91e0-406cdbac60d9 · outbound

This paper cites Bidirectional Correlation-Driven Inter-Frame Interaction Transformer for Referring Video Object Segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Bidirectional Correlation-Driven Inter-Frame Interaction Transformer for Referring Video Object Segmentation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:01:18.655013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.656301Z digest=sha256:72b9aaa7e04d27fefd54c0379b4b38bd8f2fc6d9a2f20c23fee37920ab2daa00

Observation e5c2ef69-ed6d-4bd2-a648-46efa7d3c36b · outbound

This paper cites Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks.

VoCap: Video Object Captioning and Segmentation from Any Prompt Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.521097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.662923Z digest=sha256:be75a1e727b9b75c0a7ce26c13458afbf0f89d930b9fb9ef088acfa1f32450db

Observation 88920aa7-0cda-42af-9c7d-dc9954f24fbd · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

VoCap: Video Object Captioning and Segmentation from Any Prompt Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.489542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.667573Z digest=sha256:dd7eec50494332e55fced0537aec3e89b9fb966a315efa5a9b67c17d6e7aff51

Observation 0aad7dbe-5ff1-4662-80ab-fe44c90c0184 · outbound

This paper cites Learning object context for dense captioning.

VoCap: Video Object Captioning and Segmentation from Any Prompt Learning object context for dense captioning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.461200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.673872Z digest=sha256:ba4d5d7c355d95b5dd18bdca5005c819e5ff2179f2da504ef87d71acc758d33e

Observation d712224b-a72a-4e4a-8f3e-a0bfed0eaa41 · outbound

This paper cites Exploring plain vision transformer backbones for object detection.

VoCap: Video Object Captioning and Segmentation from Any Prompt Exploring plain vision transformer backbones for object detection

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.438111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.678736Z digest=sha256:a42fab30e11c4040a05eae4eccf8fbdfd70c994b06cc41af39cd86c136dc5908

Observation 0a7bcacd-eeb5-4fb0-93e9-c08967baec2c · outbound

This paper cites Beyond mot: Semantic multi-object tracking.

VoCap: Video Object Captioning and Segmentation from Any Prompt Beyond mot: Semantic multi-object tracking

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.401508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.684439Z digest=sha256:e73d66ac1487a187f42021807d73d99c683420083468970a08d9f03598cc47b0

Observation b819030b-abb2-474b-a43c-1fdb5303f7b7 · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

VoCap: Video Object Captioning and Segmentation from Any Prompt Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.690922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.690922Z digest=sha256:f8560947ae5e847764b2933db0c589690301c200b528766472c5ebc046460b82

Observation c5e05588-bf31-435b-ac20-27d2e85eb79f · outbound

This paper cites Visual instruction tuning.

VoCap: Video Object Captioning and Segmentation from Any Prompt Visual instruction tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.696125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.696125Z digest=sha256:c949e6678c6ee9389470833498e463dc82dca5fce94d33a70374e40acadc0d3c

Observation 209ca39c-956b-4d91-a462-2f5cd2b9b926 · outbound

This paper cites Image segmentation using text and image prompts.

VoCap: Video Object Captioning and Segmentation from Any Prompt Image segmentation using text and image prompts

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.337607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.702937Z digest=sha256:c2f8399deaa008b57c62fd450962b18634c6570c099cddb8ab5c47b8422c9e92

Observation ac956e77-cd96-456a-bf9d-d075ed517381 · outbound

This paper cites Soc: Semantic-assisted object cluster for referring video object segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Soc: Semantic-assisted object cluster for referring video object segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.311031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.708656Z digest=sha256:c891536df1959e1c6f7881fd5b953451dac7271315fba56b560d3d0d87727d95

Observation 1b214c00-92d9-4c56-820f-076d44c8e8f8 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

VoCap: Video Object Captioning and Segmentation from Any Prompt Generation and comprehension of unambiguous object descriptions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.285378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.714036Z digest=sha256:e5796cd095c5a9dcb78a6cbbbd0d616492bc3620196b1338092a84ef14506921

Observation 97fb6b6c-8501-423f-9c11-619ec0520100 · outbound

This paper cites Scaling open-vocabulary object detection.

VoCap: Video Object Captioning and Segmentation from Any Prompt Scaling open-vocabulary object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.257724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.719388Z digest=sha256:232bd2ae95da56b28bce1fba1dbe57c71e2d1cb287c222c9f33557c61ccd396d

Observation c0e9d8b2-f9da-4f7d-a238-3d5f21ac534c · outbound

This paper cites Pivot: Iterative visual prompting elicits actionable knowledge for vlms.

VoCap: Video Object Captioning and Segmentation from Any Prompt Pivot: Iterative visual prompting elicits actionable knowledge for vlms

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.226009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.726585Z digest=sha256:bf0835a5c720dd19c27dbf0b65f429a5662cf3f7780a0585153712a3c7909c72

Observation e8613517-5df2-446e-b494-aa657ac9f9a4 · outbound

This paper cites Gpt-4v(ision) technical work and authors.

VoCap: Video Object Captioning and Segmentation from Any Prompt Gpt-4v(ision) technical work and authors

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.186781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.732795Z digest=sha256:ac37025c2b34189858b50d58d6db40b513ad64c83d8f200c993bf58cafc3a781

Observation c99daecb-c971-4c8e-b2cf-4124212af177 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

VoCap: Video Object Captioning and Segmentation from Any Prompt Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.740311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.740311Z digest=sha256:b8a0703bbbe339c6e1ec03ae785242d2a64ef1309f08d6666f0c36ac7438e0e2

Observation 2a349ecd-bd05-47ae-9458-a62320139450 · outbound

This paper cites Perazzi, J.

VoCap: Video Object Captioning and Segmentation from Any Prompt Perazzi, J

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.151456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.750541Z digest=sha256:9e8558eabf79c8ef61268b7de17a901ec48a4f9db40b84588dd8b4d3fc23390a

Observation 047b2d75-2e1b-4597-8ec3-1352b02d854f · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

VoCap: Video Object Captioning and Segmentation from Any Prompt Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.128513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.762144Z digest=sha256:32273bdc73c10ae2ba4828e43cb4b99b4bfe3678247752befeab21de14a70844

Observation daf5929b-7733-4dd7-b7d2-2f9781d1fcf8 · outbound

This paper cites Connecting vision and language with localized narratives.

VoCap: Video Object Captioning and Segmentation from Any Prompt Connecting vision and language with localized narratives

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.769767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.769767Z digest=sha256:158db0caebc402c37ca175d7bdc5ea09081c3b683ad897018409b82d6d19392c

Observation 2f433843-de19-4feb-9c59-54571dcbd76e · outbound

This paper cites an unresolved cited work.

VoCap: Video Object Captioning and Segmentation from Any Prompt Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:01:20.087763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.778521Z digest=sha256:a0fd6713edde1cd9265f9274bdb109a595b0a863a09ad27e81de4d0bb549dfe2

Observation 00934094-f590-4efc-8065-b9b4418ece4c · outbound

This paper cites Learning transferable visual models from natural language supervision.

VoCap: Video Object Captioning and Segmentation from Any Prompt Learning transferable visual models from natural language supervision

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.059637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.785371Z digest=sha256:b2fdbbde9f95dbfae1f801619f3eed8af4af8faea6791bc61e9d256a1bae08c3

Observation 67b91da0-785f-4eaf-a0f7-61732255d9dd · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

VoCap: Video Object Captioning and Segmentation from Any Prompt Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.790907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.790907Z digest=sha256:e74adc5cb87f0a8cef539d6a48c8fae79910db1785f59ec721ca23cbe02a05db

Observation 015ba149-0ab4-47dd-a96f-453410770f6f · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

VoCap: Video Object Captioning and Segmentation from Any Prompt SAM 2: Segment Anything in Images and Videos

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.799150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.799150Z digest=sha256:2b655e9ddf9ed668e74e389a2c096854ca147530fe79e6aeab80a70fd5e89573

Observation 298b0a5c-55b7-446e-9eec-1bfdfe473786 · outbound

This paper cites Youtube- boundingboxes: A large high-precision human-annotated data set for object detection in video.

VoCap: Video Object Captioning and Segmentation from Any Prompt Youtube- boundingboxes: A large high-precision human-annotated data set for object detection in video

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:20.007307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.807605Z digest=sha256:711e4ec30489e7b9618453c81e9ea79e1ba1bd44b0e676dc374ce248438bf95d

Observation d1f04666-ef6f-4fd3-b188-1f2075ce5503 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt U-net: Convolutional networks for biomedical image segmentation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.812043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.812043Z digest=sha256:6f774ffba7800b6cb4b5fefa576e5699347cce6f960acbed838bce4bd75f46de

Observation 2c89d1ec-60f0-4ec5-8062-292d5eefe26a · outbound

This paper cites Berg, and Li Fei-Fei.

VoCap: Video Object Captioning and Segmentation from Any Prompt Berg, and Li Fei-Fei

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.968197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.817510Z digest=sha256:c218652a6602411bdcd21aee8908c8cb2e59d3db5b709cc2efe212bef876b0b6

Observation 9c714a6c-cc32-4d0d-8fb1-c6e0d7e68ec9 · outbound

This paper cites Hiera: A hierarchical vision transformer without the bells-and-whistles.

VoCap: Video Object Captioning and Segmentation from Any Prompt Hiera: A hierarchical vision transformer without the bells-and-whistles

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.934662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.822085Z digest=sha256:b9cf4a7316233da3b4aa32016bd4aa16a57d25ef33c1cd8b6587c5559e3d4ef5

Observation c36ae280-65e0-471a-a2e6-738d8c47fe3d · outbound

This paper cites Tokenlearner: Adaptive space-time tokenization for videos.

VoCap: Video Object Captioning and Segmentation from Any Prompt Tokenlearner: Adaptive space-time tokenization for videos

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.890989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.827140Z digest=sha256:d3d9c3d16b4a7288f8d446234d64af13df5f598b7f5c207ad850735c1ac8adaf

Observation de29efa0-952a-47d5-94d1-8310d0897d9c · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

VoCap: Video Object Captioning and Segmentation from Any Prompt Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.862210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.833973Z digest=sha256:b0b9e764ce122b25f3fa4739e80473f8c909c81fdceb36205cd56e0468a3ff1e

Observation eb82c4d2-4bf1-49f2-9550-83c922099e37 · outbound

This paper cites Annotating objects and relations in user-generated videos.

VoCap: Video Object Captioning and Segmentation from Any Prompt Annotating objects and relations in user-generated videos

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.814186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.839406Z digest=sha256:4ccc8a30ea189f92e709d3abd643ef662cd6a7cb218108e48df58fc62f663d90

Observation 317edd14-e39c-4dbb-984f-57f9c1f2ccab · outbound

This paper cites Region-object relation-aware dense captioning via transformer.

VoCap: Video Object Captioning and Segmentation from Any Prompt Region-object relation-aware dense captioning via transformer

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.771189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.845053Z digest=sha256:26be1ce66a133ab08b5260921cdd6f363649d9a4da715ea0742377a8dead93a5

Observation 3b4551fa-fdd1-4b54-a5a1-245f1020df62 · outbound

This paper cites What does clip know about a red circle? visual prompt engineering for vlms.

VoCap: Video Object Captioning and Segmentation from Any Prompt What does clip know about a red circle? visual prompt engineering for vlms

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.740260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.853275Z digest=sha256:c363c4dcc664f8f4a713876b279e298d05f27023af2760a26f0691efbbe46fb6

Observation 41ad1629-a574-43b7-84d9-88bfbd4b6591 · outbound

This paper cites Video foundation models for animal behavior analysis.

VoCap: Video Object Captioning and Segmentation from Any Prompt Video foundation models for animal behavior analysis

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.702233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.860408Z digest=sha256:cb7185807cc5f6532c3e9ecfcf0b3e2e4b23190401aab8ff8f8a096a54a0a8e1

Observation 8e79880d-d477-4a52-9edc-08113a32497b · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset.

VoCap: Video Object Captioning and Segmentation from Any Prompt Scalability in perception for autonomous driving: Waymo open dataset

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.865273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.865273Z digest=sha256:d6e06faa289a86df9ce3aa93f585b96d1f679be8d78c25ae19d8b37c9580b1a4

Observation 66638174-134c-448a-8f47-8dd4c160250f · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

VoCap: Video Object Captioning and Segmentation from Any Prompt Gemma: Open Models Based on Gemini Research and Technology

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.871477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.871477Z digest=sha256:1d165a17824fed9a2cbc27c9018025c8f87090a5f6fe5a1b67b71949b8028955

Observation 494a617f-2afe-4654-8bf5-8521bb3e8b92 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

VoCap: Video Object Captioning and Segmentation from Any Prompt Gemma 2: Improving Open Language Models at a Practical Size

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.880826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.880826Z digest=sha256:b4e3a6e60f346c80a6461cfe03f95cfa7d4c3b9d15aa1fdb385f054cdb524bc5

Observation 3b3fd1a2-1124-4050-8d20-b132447d8328 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VoCap: Video Object Captioning and Segmentation from Any Prompt LLaMA: Open and Efficient Foundation Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.886308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.886308Z digest=sha256:bca62294d5f0c9fb1bb3f293a1879fa7ef24256ca2440fb7d2852bd2597640ce

Observation ee97b73b-be38-47f1-8c9a-70004761ad0f · outbound

This paper cites Attention is all you need.

VoCap: Video Object Captioning and Segmentation from Any Prompt Attention is all you need

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.891687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.891687Z digest=sha256:2edac2974a40207133104b512080825f0768b54e97a8d4c509a812c8233368ea

Observation 6fd80b34-72bd-4cc3-9c36-b1359b36f6c2 · outbound

This paper cites Cider: Consensus-based image description evaluation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Cider: Consensus-based image description evaluation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.621641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.896550Z digest=sha256:46374f9e77fa8926e2cd2f3de4e038db6d98c264c2ae0f73681c7d8fdf18109a

Observation 04161a33-0846-43ff-a739-6bf328814047 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

VoCap: Video Object Captioning and Segmentation from Any Prompt Phenaki: Variable length video generation from open domain textual descriptions

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.598780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.905108Z digest=sha256:1416ce6817b9d3c624a28798f0826f0d7d7efa3fd2123e3196b2333b91f5c3f3

Observation 93bb1291-efe9-4fea-a86c-4c3458880499 · outbound

This paper cites Connecting vision and language with video localized narratives.

VoCap: Video Object Captioning and Segmentation from Any Prompt Connecting vision and language with video localized narratives

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.574760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.910810Z digest=sha256:99f881fb825f8caaba1ec938511bbd76c38843d405e23795d38153d95ffe6d45

Observation 76f7b4b0-3460-45d5-a143-cdc5a9dbe52a · outbound

This paper cites Git: A generative image-to-text transformer for vision and language.

VoCap: Video Object Captioning and Segmentation from Any Prompt Git: A generative image-to-text transformer for vision and language

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.554837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.915952Z digest=sha256:2cda5770aed4791deae5b206ae46bcdc0111ca15572ac410f5080b91c00373e3

Observation ddd0f84e-71ef-41e2-b5b5-b85091610457 · outbound

This paper cites End-to-end dense video captioning with parallel decoding.

VoCap: Video Object Captioning and Segmentation from Any Prompt End-to-end dense video captioning with parallel decoding

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.518822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.920936Z digest=sha256:8a63e5f363f1332f52a72cd2755ccf100217feaf9c50f4e405ccdd0a2b5cac34

Observation d78b6e5f-978c-4106-b90a-5f503598f5ac · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

VoCap: Video Object Captioning and Segmentation from Any Prompt Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.925533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.925533Z digest=sha256:2fbe38353e7e9e7ff1425e639ff758a72dd7d08bbc294c2c85eccd90b5abd777

Observation edfe104c-d0a3-4856-b530-fe9e64d72b31 · outbound

This paper cites Unidentified video objects: A benchmark for dense, open-world segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Unidentified video objects: A benchmark for dense, open-world segmentation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.485562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.930332Z digest=sha256:33b2327b5524fa49e9ba81d7a104377d59636690f04f0020d2df2fef0c403a8d

Observation 5dccc15b-8b05-4d5b-9b86-998459bd7ad9 · outbound

This paper cites The all-seeing project: Towards panoptic visual recognition and understanding of the open world.

VoCap: Video Object Captioning and Segmentation from Any Prompt The all-seeing project: Towards panoptic visual recognition and understanding of the open world

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.451196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.935974Z digest=sha256:c1e85329d3d3847776f05c9a7bc9466e708e5c24bcf6c3147a5d4a3ea5f666f1

Observation 1daf37a9-7708-467f-9ccb-7631b853dd2b · outbound

This paper cites The All-Seeing Project V2: Towards General Relation Comprehension of the Open World.

VoCap: Video Object Captioning and Segmentation from Any Prompt The All-Seeing Project V2: Towards General Relation Comprehension of the Open World

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.942135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.942135Z digest=sha256:cb1b1e401ddff78267083bb9b647cb116f4c1abde355efc4c747e13c36c24889

Observation a59b876c-4e5d-41bb-aa62-b96cdfca49c2 · outbound

This paper cites Instancediffusion: Instance-level control for image generation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Instancediffusion: Instance-level control for image generation

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.426563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.948334Z digest=sha256:bec3c1cb23119a004701b3b3faba1998dccf88390431291e96e669ef0de83415

Observation d24b68d0-16ec-4dfa-bc4b-bb4a07d1beb5 · outbound

This paper cites Language as queries for referring video object segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Language as queries for referring video object segmentation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.397909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.954568Z digest=sha256:3ad5d1d5dc0b137378ec677d950fdb37174823e6b6d32c086a28cd0bfe787e95

Observation 2e6c14ce-75b9-49c9-96ef-bcefd03f0a57 · outbound

This paper cites UniRef++: Segment Every Reference Object in Spatial and Temporal Spaces.

VoCap: Video Object Captioning and Segmentation from Any Prompt UniRef++: Segment Every Reference Object in Spatial and Temporal Spaces

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:01:18.397693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.960362Z digest=sha256:3ba4d682f1ed1c7d6f3fec651202fb0f2c0c377fa0207374a06966017408ddbc

Observation b41afe1f-22a7-4ab7-9a9e-29c1d14a7a23 · outbound

This paper cites General object foundation model for images and videos at scale.

VoCap: Video Object Captioning and Segmentation from Any Prompt General object foundation model for images and videos at scale

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.377670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.966174Z digest=sha256:9215352f19636b4f5b710140c6c2c53c139ec76fdf48d2aa93c53d0845b17eb5

Observation 6ba209f7-85dc-46f5-be63-4b677a35aee7 · outbound

This paper cites Grit: A generative region-to-text transformer for object understanding.

VoCap: Video Object Captioning and Segmentation from Any Prompt Grit: A generative region-to-text transformer for object understanding

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.352240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.975222Z digest=sha256:17596de0e59dc72874fe3a8fa67945babd709cedc4b4c5a4ccd86e0c7bc32355

Observation bd63d38f-194e-4ba3-9319-3ee30bf19a55 · outbound

This paper cites Dettoolchain: A new prompting paradigm to unleash detection ability of mllm.

VoCap: Video Object Captioning and Segmentation from Any Prompt Dettoolchain: A new prompting paradigm to unleash detection ability of mllm

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.328388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.980591Z digest=sha256:ced41a7868917a7634a0a1b4f82f5d9c95b1556a3b29345ffef737e164cdfc0c

Observation a952e8fd-2d3c-4193-abab-4eb57748c530 · outbound

This paper cites Pixel-aligned language model.

VoCap: Video Object Captioning and Segmentation from Any Prompt Pixel-aligned language model

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.282154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.986304Z digest=sha256:b9eeafe9c93e7239ee532a0ffc1cc4f53b8f54ce7ad5570d90a65e04dcd3955d

Observation 10886054-f88b-4cdd-8a12-d32a13b39de9 · outbound

This paper cites Youtube-vos: Sequence-to-sequence video object segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Youtube-vos: Sequence-to-sequence video object segmentation

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.243223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:17.994291Z digest=sha256:eb515632e0708361e68f2cf7066f9c426ccdbd537d3333a4c14a4ffc268a8e9f

Observation 0ea80731-72c7-4620-b19b-ed2f23f557f6 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

VoCap: Video Object Captioning and Segmentation from Any Prompt xgen-mm (blip-3): A family of open large multimodal models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:18.000551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:18.000551Z digest=sha256:795d63226fd2500111d3155b5c49253de9f5304920442ac57901dc893ab6eed8

Observation 524bf4cc-24fa-4109-9c75-fe402ae84dab · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning.

VoCap: Video Object Captioning and Segmentation from Any Prompt Vid2seq: Large-scale pretraining of a visual language model for dense video captioning

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.221648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:18.005261Z digest=sha256:4cacfd6d8110d343894feabf2822ad1da1b54c1f279e39c43c8397546588a7e3

Observation dc07d712-d243-49bb-9167-63d0ab8cde11 · outbound

This paper cites Track Anything: Segment Anything Meets Videos.

VoCap: Video Object Captioning and Segmentation from Any Prompt Track Anything: Segment Anything Meets Videos

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:18.011590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:18.011590Z digest=sha256:5ef06c57632f3fdf51673c8c8f2ec62741fc74eee8a4151177d62f9ca354d1aa

Observation 0deecbdc-995a-4edd-b1b3-a9fed35a2141 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

VoCap: Video Object Captioning and Segmentation from Any Prompt Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:18.017383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:18.017383Z digest=sha256:c13ebcb1a2148068affb3a008ffd2e693013be03281757051509228fec01baae

Observation 1d442d00-cc26-4cd2-ab3a-5d1ef693dfdc · outbound

This paper cites Decoupling features in hierarchical propagation for video object segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Decoupling features in hierarchical propagation for video object segmentation

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.187726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:18.022264Z digest=sha256:0951852ab63f561908cb9f90d60d7823e9320eb02ac080a7bf5f2c3f2bbf569f

Observation 3d09fbf5-b2e4-47fe-925c-160612d99570 · outbound

This paper cites Associating objects with transformers for video object segmentation.

VoCap: Video Object Captioning and Segmentation from Any Prompt Associating objects with transformers for video object segmentation

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.158353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:18.030383Z digest=sha256:70b36d84fa36266f07f60af3cb2303f31878778b8182494ef0d218bbf3bce0a9

Observation a4f0f111-19a8-4746-a4f8-cd040d2366b8 · outbound

This paper cites Scalable video object segmentation with identification mechanism.

VoCap: Video Object Captioning and Segmentation from Any Prompt Scalable video object segmentation with identification mechanism

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.135352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:18.036508Z digest=sha256:b5a405d5083431a4da49bd5a9a114123d6a95ec96ff963be1e980f825563d7a8

Observation 1c1d4855-d106-4023-abbd-8fadeba161e6 · outbound

This paper cites Describing videos by exploiting temporal structure.

VoCap: Video Object Captioning and Segmentation from Any Prompt Describing videos by exploiting temporal structure

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.112159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:18.042153Z digest=sha256:7528cd0abf79ccff3dddc8059dc9a18241f78c3f4339eabf9f5b219d5706a1c7

Observation 37ff9346-5dfa-40b8-b6c7-454f4ad94255 · outbound

This paper cites Modeling context in referring expressions.

VoCap: Video Object Captioning and Segmentation from Any Prompt Modeling context in referring expressions

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:01:19.084617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:01:18.047778Z digest=sha256:018839619ad31478b0b988873a7c0898b709beb15314c75c915de7d50be01011

Pith citing papers

Observation 501080c7-8714-494f-9325-48e7a4702264 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs VoCap: Video Object Captioning and Segmentation from Any Prompt

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.699718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:e00f8520d2b435947c7ed2a012377536cc064c830bc0678e7c5edd86b9a0ce6f

Observation c1f47cb2-a1bc-4129-a261-406d0b614598 · inbound

Video Generation Models are General-Purpose Vision Learners cites this paper.

Video Generation Models are General-Purpose Vision Learners VoCap: Video Object Captioning and Segmentation from Any Prompt

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:98f3ca708ddf27e296f8a84c5b1c8acfca9ff2fa441d25f9889e427478cbe5f6