Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:14.907661Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 15 inbound Pith citation observations for arXiv:2506.05302.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:14.907661Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:51:24.717294Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T17:27:15.782237Z
89 of 89 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5e6c0141-9f53-4a7a-9fb3-8161ec7a3f0d · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Mc-llava: Multi-concept personalized vision-language model.arXiv preprint arXiv:2411.11706, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9caa0b70-e2a3-4d5d-a811-fa791f80c4c2 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b44229d7-89fb-458d-934d-5bdf41176fcd · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78c355f2-42bb-4a4e-886c-ad97de098a34 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Abductive commonsense reasoning, 2020
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53317ec5-f5f6-4ddc-b81d-c892bb967dcc · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Graph cuts in vision and graphics: Theories and applications
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b074da60-bfa4-4d8e-9662-7c15f588d927 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Scene-text oriented referring expression comprehension.IEEE Transactions on Multimedia, 25:7208–7221, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8625921-7c21-4aa7-9a75-856e626b190f · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Activitynet: A large-scale video benchmark for human activity understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d1cb01a-b711-45b9-835b-124c53b2aa96 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Vip-llava: Making large multimodal models understand arbitrary visual prompts
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dac110d-8043-47ee-ae47-2bbdb3a948b2 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Active contours without edges.IEEE Transactions on image processing, 10(2):266–277, 2001
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a0a7662-5adb-4947-9e25-a5fee112ded4 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b362ae58-7c9a-44a4-9249-8f628c09b56a · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Videollm-online: Online video large language model for streaming video
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d57f293-dc7e-400f-8b59-ed0e79f1b019 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 917015bf-08bc-43be-87e3-d8b9bc2d7aa5 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Segment and Track Anything
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 000a4c48-a6f7-46bc-8812-11e99f4309c4 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Total-text: A comprehensive dataset for scene text detection and recognition, 2017
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e559b13-e1bc-4755-b972-952d73d0b449 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos V ocabulary-free image classification, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6bccd78c-1e66-4279-9cda-1f1ba90c107c · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Online action detection
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90bd38f4-2b46-4711-8c71-cd433798ce8c · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Mevis: A large-scale benchmark for video segmentation with motion expressions, 2023
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec455c19-e304-4fd8-a4fd-c922bb6e03b0 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Actor and action video segmentation from a sentence
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23fc4dd0-9b5d-4320-828c-a68112c74745 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Icdar2017 robust reading challenge on coco-text
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93ee6ab1-f238-48e7-bb87-02d288870d6f · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Ego4d: Around the world in 3,000 hours of egocentric video
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fccdb72-31f9-4b8f-8922-eae20a984627 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Regiongpt: Towards region understanding vision language model, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e817e83-8877-43fe-bf03-842cf459c6cb · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos TRACE: Temporal Grounding Video LLM via Causal Event Modeling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a95b82-b009-4ffe-87f4-a71ce956a7b2 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Lvis: A dataset for large vocabulary instance segmentation, 2019
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf17f7ba-2721-403b-921c-fcb3e4933269 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Synthetic data for text localisation in natural images, 2016
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f6b5a53-9742-454a-9744-3d2693faf534 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Omni-rgpt: Unifying image and video region-level understanding via token marks, 2025
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 329c1ee0-6d7c-4793-b219-11a137c499d1 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Segment and caption anything
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f5ee2b8-10b5-4d15-9577-c1b75af42a49 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos GPT-4o System Card
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da1fde41-df61-47a4-b776-fe83df5a39a0 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Visual prompt tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e4fafc4-c40b-4f2c-abba-fae1b630d798 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0320ae6a-736d-466a-9137-320cf8b0beb0 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Icdar 2015 competition on robust reading
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79fca956-3d36-4d13-8c15-fc1fce4ea5a4 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Icdar 2013 robust reading competition
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75fff697-834b-493d-9c0e-ebd32d1da19c · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Referitgame: Referring to objects in photographs of natural scenes
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dd22ac1-8d1d-418a-872d-1c1d175764ca · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Segment anything in high quality.Advances in Neural Information Processing Systems, 36:29914–29934, 2023
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9594d68a-989f-41aa-835f-5303e38fd1d0 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Segment Anything
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02a13c13-2da8-45f8-bc92-16ded190fae2 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Openimages: A public dataset for large-scale multi-label and multi-class image classification.Dataset available from https://github
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbc8b219-ba86-4eeb-a622-83fad787060b · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Shamma, Michael S
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2343811e-b4e8-4161-99c3-e1eed22d929b · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Beyond mot: Semantic multi-object tracking
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3de15e83-db5f-4f6d-8e47-8de78f50c331 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Describe Anything: Detailed Localized Image and Video Captioning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5acc589-816e-4212-9a41-05c03a5b85cc · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Rouge: A package for automatic evaluation of summaries
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c18e82-2491-4fb3-a229-e0f3415129a5 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Lawrence Zitnick, and Piotr Dollár
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 734693f4-cdbe-494d-a851-2c3d8c130a24 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af9be36d-cf66-4a61-ae61-4bb17b033d4f · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos DeepSeek-V3 Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b80ee9d-3b88-48a8-aa56-1bae4df04a53 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Gres: Generalized referring expression segmentation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9626591a-cf6b-4f0e-85b6-962987fec048 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d60893a0-c1ee-47ba-8b40-24254752769b · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Improved baselines with visual instruction tuning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f5a9006-e068-4a88-9059-5b2f4069a2bb · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40857708-b9b1-43ac-8f63-68cc2bbdbcff · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos The 2017 davis challenge on video object segmentation, 2018
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dca6c900-6108-4828-b8be-4fda60a8bde5 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Artemis: Towards referential understanding in complex videos
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0411d034-bbf6-4397-8a7b-2527829345fb · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Artemis: Towards referential understanding in complex videos, 2024
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 615de475-4912-4b52-843a-076e6b94c72f · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Paco: Parts and attributes of common objects, 2023
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b52654c-8e87-4765-9961-b934d46bc038 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80503a7d-ce6a-4a73-abf5-43bb8717d469 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos SAM 2: Segment Anything in Images and Videos
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d3f82b-7679-4289-a451-10f39f150192 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Sam 2: Segment anything in images and videos, 2024
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc7a87ef-9de9-4a0f-a155-21f4c3590af6 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53fdc3f9-85b0-41cd-9010-687b9512139e · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Objects365: A large-scale, high-quality dataset for object detection
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd84cbc5-8476-4f71-9d32-a5ca9a7245e6 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text, 2021
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 003298e2-5e3e-4286-a1da-025dd32bb251 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Icdar 2019 competition on large-scale street view text with partial labeling – rrc-lsvt, 2019
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97ae4dcd-4b1e-4f1b-90fd-60373548ef85 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Human-centric spatio-temporal video grounding with visual transformers, 2021
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f81eebbf-4a2b-4617-958f-4e2a0ea98be1 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Human- centric spatio-temporal video grounding with visual transformers.IEEE Transactions on Circuits and Systems for Video Technology, 32(12):8238–8249, 2021
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6c7bb8a-736d-4403-b5a3-847e4dbdc50a · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Cider: Consensus-based image description evaluation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6538d8ce-016e-4755-9409-05c67de2a623 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Coco-text: Dataset and benchmark for text detection and recognition in natural images, 2016
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2be0f5f2-0406-4073-b936-c8820e72116c · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Elysium: Exploring object-level perception in videos via mllm, 2024
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86242f7d-353d-4fce-9401-a217cb7e31d1 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Elysium: Exploring object-level perception in videos via mllm
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d27a146f-4d23-4b7a-8f8a-50063baddd19 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Towards open-vocabulary video instance segmentation, 2023
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24ea961e-be69-4b03-8336-d84a91f841d6 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Git: A generative image-to-text transformer for vision and language, 2022
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4f51c45-574e-48be-87e2-b80093e76b7c · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos V3det: Vast vocabulary visual detection dataset, 2023
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64243a39-bee3-4aca-b170-7f18f7ea9eea · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Caption Anything: Interactive Image Description with Diverse Multimodal Controls
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b4824b5-1247-46f0-9acb-6b6f537991e6 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce98c32a-3bee-4a17-b336-b03bff860f17 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Grit: A generative region-to-text transformer for object understanding
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation feb50aa7-36ce-4526-9e51-e569d7c557cb · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.Advances in Neural Information Processing Systems, 37:69925–69975, 2024
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e4340d2-d0db-4de9-ba17-cf974d152d9f · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Youtube-vos: A large-scale video object segmentation benchmark, 2018
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6306acd4-37f1-40c5-8523-43cfcc258b31 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Qwen2.5 Technical Report
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bba35fb9-1d5f-43a0-9e47-1b2befb8714e · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Vid2seq: Large-scale pretraining of a visual language model for dense video captioning, 2023
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 038aa238-4d5e-452e-897b-8f9fd415f96a · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Samurai: Adapting segment anything model for zero-shot visual tracking with motion-aware memory, 2024
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f24cf1c-e96e-498b-94ec-f38cf6b5d0a6 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b581093-f66b-4a18-8762-233b527d2abf · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Detecting texts of arbitrary orientations in natural images
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15915057-4c0d-42e2-9ead-e8bb545702d5 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54f1de86-fd8f-4e15-ac4e-e612520875a4 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Merlin: Empowering multimodal llms with foresight minds
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04667787-33fe-4d7a-b681-83be6aea276a · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos, 2025
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3532d1b8-250b-4955-bf83-c62da59b9e9e · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Osprey: Pixel understanding with visual instruction tuning
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60a624db-a6fa-43c6-aedc-ac269ea708c2 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86547f57-af2f-49a4-aa87-4cd8e57a6fa5 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Faster Segment Anything: Towards Lightweight SAM for Mobile Applications
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae642888-c3de-4cbe-aa24-3161505b0301 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e377b8a5-dc62-43a2-b95e-5499da6488b8 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Gpt4roi: Instruction tuning large language model on region-of-interest, 2025
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 614da7e3-d07b-487c-875c-d939bf00e4fa · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Where does it exist: Spatio-temporal video grounding for multi-form sentences
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ccdc906f-94a4-4d76-bf34-91859924d5ba · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 037adc13-83fb-4e76-a27e-08af8718a73b · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Fast Segment Anything
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b5a5567-5bb8-4335-a8ee-4977daff2dd2 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Controlcap: Controllable region-level captioning
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 103c9006-1579-4b31-b1f1-df970f5937e1 · outbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Streaming dense video captioning
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1309521c-a60a-444f-ba0b-13861ed211f6 · inbound
Describe Anything Model for Visual Question Answering on Text-rich Images Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b3d678-10f0-4408-b6bb-06bdb6520ea5 · inbound
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 293
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b0f3896-8ba2-43c2-ac5a-aa020adf157d · inbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bdff3cf-8f46-4932-b4a7-efc5fd55bbdc · inbound
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64af0091-4638-4aaa-8598-57b8cbcfa627 · inbound
MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f78f9f04-29ac-4021-89e9-4f639e9d59ea · inbound
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59ec8268-c8fa-4276-9aee-c8c8bb744edd · inbound
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d313a4e-0c65-4f54-b147-e665b3c3d17c · inbound
Enhancing Foundation VLM Robustness to Missing Modality: Scalable Diffusion for Bi-directional Feature Restoration Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 728775e5-abdc-40d7-84b4-a2b3a515999e · inbound
OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87f48f3a-8a30-4c2e-baec-3acd0e31e9bc · inbound
OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c25c039a-fdd0-4fdc-a0c6-8609bd93d72d · inbound
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fecadecc-bf3c-4c06-850e-e596473f3e66 · inbound
WOW-Seg: A Word-free Open World Segmentation Model Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1ce6cc2-ff2d-43fa-bcdf-2036aa9755d6 · inbound
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 694d2d02-f883-4042-9944-73eace34e136 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f067d0b1-ed72-4548-b86f-4a796ce217ed · inbound
FRFDet: Efficient UAV Small Object Detection with Symmetric Sampling and Scalable Fusion Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.