Pith. sign in

Paper Citation Record · LEDGER

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

As of 7 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 15 inbound Pith citation observations for arXiv:2506.05302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05302 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:14.907661Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:51:24.717294Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.782237Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5e6c0141-9f53-4a7a-9fb3-8161ec7a3f0d · outbound

This paper cites Mc-llava: Multi-concept personalized vision-language model.arXiv preprint arXiv:2411.11706, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Mc-llava: Multi-concept personalized vision-language model.arXiv preprint arXiv:2411.11706, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:07.834487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:07.834487Z digest=sha256:836bba8b0b356dfb7c842477ff3cd0291f363caa29beb87b0649f8ec4e7a458f

Observation 9caa0b70-e2a3-4d5d-a811-fa791f80c4c2 · outbound

This paper cites Qwen2.5-VL Technical Report.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:07.986779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:07.986779Z digest=sha256:b963b0f3ada7f89e12dcf08c8759bcefff8aba118e7ad98cbc36097d470cb5da

Observation b44229d7-89fb-458d-934d-5bdf41176fcd · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.106718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.106718Z digest=sha256:b8346744b93b07fa004c9977b295f25d85312d435405673eec9deb5dbf7f3111

Observation 78c355f2-42bb-4a4e-886c-ad97de098a34 · outbound

This paper cites Abductive commonsense reasoning, 2020.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Abductive commonsense reasoning, 2020

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.211037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.211037Z digest=sha256:0595bdfe2557cbfec4f05ac9af160d1c811048e6c00a1709cec6622e6dcb941a

Observation 53317ec5-f5f6-4ddc-b81d-c892bb967dcc · outbound

This paper cites Graph cuts in vision and graphics: Theories and applications.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Graph cuts in vision and graphics: Theories and applications

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.339094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.339094Z digest=sha256:00d551ea8bb4c07bd9f2d66c574d1ef7cfe12c1a01a6e21774c12fb63f6cfd47

Observation b074da60-bfa4-4d8e-9662-7c15f588d927 · outbound

This paper cites Scene-text oriented referring expression comprehension.IEEE Transactions on Multimedia, 25:7208–7221, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Scene-text oriented referring expression comprehension.IEEE Transactions on Multimedia, 25:7208–7221, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.465002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.465002Z digest=sha256:56250ed4d332436b3dc569d0132a853fdee98ec765acd0418ca32fd63a2c16b5

Observation b8625921-7c21-4aa7-9a75-856e626b190f · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Activitynet: A large-scale video benchmark for human activity understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.647907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.647907Z digest=sha256:96d4b564c3e36ba5945eb79117fa2384ad1db66b382df592747728c4cb7ef807

Observation 6d1cb01a-b711-45b9-835b-124c53b2aa96 · outbound

This paper cites Vip-llava: Making large multimodal models understand arbitrary visual prompts.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Vip-llava: Making large multimodal models understand arbitrary visual prompts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.783473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.783473Z digest=sha256:d7826b0d8677f58fdb3d3e3fcb850fbf22c9b873b06693855a19f0500d994bfe

Observation 9dac110d-8043-47ee-ae47-2bbdb3a948b2 · outbound

This paper cites Active contours without edges.IEEE Transactions on image processing, 10(2):266–277, 2001.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Active contours without edges.IEEE Transactions on image processing, 10(2):266–277, 2001

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.952935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.952935Z digest=sha256:a9c1411d0bbc25cd503481712520222a54982ef64ba068db86891c7d2a109eec

Observation 8a0a7662-5adb-4947-9e25-a5fee112ded4 · outbound

This paper cites an unresolved cited work.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:09.089611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:09.089611Z digest=sha256:0830c3fa4f630a832c10a00055841c689dd37bccaaad3654ec8cf12a2612ab69

Observation b362ae58-7c9a-44a4-9249-8f628c09b56a · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Videollm-online: Online video large language model for streaming video

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:09.157074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:09.157074Z digest=sha256:8b57b58e74f5f9ed3c9552723c9387332f689251a912cb2561657a78da3de4a0

Observation 3d57f293-dc7e-400f-8b59-ed0e79f1b019 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:09.209098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:09.209098Z digest=sha256:13701abf9be769f83e7d066d46204d1ed168997dad082d3ba40de108194c63bb

Observation 917015bf-08bc-43be-87e3-d8b9bc2d7aa5 · outbound

This paper cites Segment and Track Anything.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Segment and Track Anything

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:09.287312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:09.287312Z digest=sha256:1ac1d9455353192b46f4d0087f5162d05d1e152592f0a1b8952cb323429d5baa

Observation 000a4c48-a6f7-46bc-8812-11e99f4309c4 · outbound

This paper cites Total-text: A comprehensive dataset for scene text detection and recognition, 2017.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Total-text: A comprehensive dataset for scene text detection and recognition, 2017

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:22.167513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:09.366044Z digest=sha256:1a51110beed771d47b4c9a1739541bf4eacbbbbfb4cf067d29a4039420cfd2f1

Observation 9e559b13-e1bc-4755-b972-952d73d0b449 · outbound

This paper cites V ocabulary-free image classification, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos V ocabulary-free image classification, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:21.982352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:09.461655Z digest=sha256:d58ff854e0f2c0a5ec4ca885d7dd03bd7a960be25552c8e0c0b0c779520cefa5

Observation 6bccd78c-1e66-4279-9cda-1f1ba90c107c · outbound

This paper cites Online action detection.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Online action detection

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:21.636264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:09.531677Z digest=sha256:e9b1702239502c7a2f0db469310df0fa2f437a9a726d06e4d5d5cdc1df81afb5

Observation 90bd38f4-2b46-4711-8c71-cd433798ce8c · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Mevis: A large-scale benchmark for video segmentation with motion expressions, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:21.160819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:09.622757Z digest=sha256:feefd49a23a722f1284d2b619fd85b09ece3910c94acc37b2d3e47821f8e9005

Observation ec455c19-e304-4fd8-a4fd-c922bb6e03b0 · outbound

This paper cites Actor and action video segmentation from a sentence.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Actor and action video segmentation from a sentence

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:21.019625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:09.710833Z digest=sha256:6f32f758b7e7c0cc131ecb2b1000be13478c150d48655236edc00b77bf45a1e6

Observation 23fc4dd0-9b5d-4320-828c-a68112c74745 · outbound

This paper cites Icdar2017 robust reading challenge on coco-text.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Icdar2017 robust reading challenge on coco-text

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:20.857060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:09.778550Z digest=sha256:ac226c08af29c02d79b7aa798268f1eed5b8773d2ee5a2e170c3fc8d7c141947

Observation 93ee6ab1-f238-48e7-bb87-02d288870d6f · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Ego4d: Around the world in 3,000 hours of egocentric video

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:09.831040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:09.831040Z digest=sha256:32bd3c42593cd34932ec06e091ab5ba2814a996dc5be6f222f79662e21419c7a

Observation 0fccdb72-31f9-4b8f-8922-eae20a984627 · outbound

This paper cites Regiongpt: Towards region understanding vision language model, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Regiongpt: Towards region understanding vision language model, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:20.698363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:09.953283Z digest=sha256:aba7b42c7867228c75d2523f0f6dce41ce0ff90e417c0efaae7470bd82b9a223

Observation 8e817e83-8877-43fe-bf03-842cf459c6cb · outbound

This paper cites TRACE: Temporal Grounding Video LLM via Causal Event Modeling.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.068680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.068680Z digest=sha256:77813e233040d8779c2e7511fdf03bfec5d5eb214ff1fb04915d48f317072573

Observation b1a95b82-b009-4ffe-87f4-a71ce956a7b2 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation, 2019.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Lvis: A dataset for large vocabulary instance segmentation, 2019

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.160396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.160396Z digest=sha256:e4ff72f97b30969fb03e9cf5a1b238321351fb82d9a5aa395aa201e5fc4f1165

Observation bf17f7ba-2721-403b-921c-fcb3e4933269 · outbound

This paper cites Synthetic data for text localisation in natural images, 2016.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Synthetic data for text localisation in natural images, 2016

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:20.526290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:10.245238Z digest=sha256:9f93274f4bddf6277dc4159647b2ce8fc90a05ff50a84a14a70e506ccd994558

Observation 9f6b5a53-9742-454a-9744-3d2693faf534 · outbound

This paper cites Omni-rgpt: Unifying image and video region-level understanding via token marks, 2025.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Omni-rgpt: Unifying image and video region-level understanding via token marks, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:20.368507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:10.341914Z digest=sha256:52c41b1a7686b9ac317cdf9200316623ecc700b47b867c137be24b4eafe46fd2

Observation 329c1ee0-6d7c-4793-b219-11a137c499d1 · outbound

This paper cites Segment and caption anything.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Segment and caption anything

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:20.174249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:10.442313Z digest=sha256:6b6bb035560bc55be5695d3aec2e4bce749cd39bc1b30536b100fc9b11721b26

Observation 8f5ee2b8-10b5-4d15-9577-c1b75af42a49 · outbound

This paper cites GPT-4o System Card.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos GPT-4o System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.544501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.544501Z digest=sha256:3b97720549c667a095e1857f0bdd58e285fc2e6dc737223bfe4e431366d4710e

Observation da1fde41-df61-47a4-b776-fe83df5a39a0 · outbound

This paper cites Visual prompt tuning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Visual prompt tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.660851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.660851Z digest=sha256:26b5bfb23816c677f865d10fc9523e6bc10299483483a9e4c6e0ee455133c7cd

Observation 7e4fafc4-c40b-4f2c-abba-fae1b630d798 · outbound

This paper cites ChatRex: Taming Multimodal LLM for Joint Perception and Understanding.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.735473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.735473Z digest=sha256:2c52775310d82554f10346e94d820b6ff80624681efdb8efe0c7e47176a1e34f

Observation 0320ae6a-736d-466a-9137-320cf8b0beb0 · outbound

This paper cites Icdar 2015 competition on robust reading.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Icdar 2015 competition on robust reading

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.977543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:10.822007Z digest=sha256:2a6554c63178d1767c9e68a0f9c3191416cf2c0c1405092201a7581e02f75391

Observation 79fca956-3d36-4d13-8c15-fc1fce4ea5a4 · outbound

This paper cites Icdar 2013 robust reading competition.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Icdar 2013 robust reading competition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.819618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:10.902744Z digest=sha256:5be35558e5d8a46eb2d7e07faf55c64b73cdffb8ad5cd6948d6507c0dc366ff9

Observation 75fff697-834b-493d-9c0e-ebd32d1da19c · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Referitgame: Referring to objects in photographs of natural scenes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.990032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.990032Z digest=sha256:b2a2ada93da93273eae8aefe566b5f925aeed2d2dd463c2e75076e50ad7e27d7

Observation 7dd22ac1-8d1d-418a-872d-1c1d175764ca · outbound

This paper cites Segment anything in high quality.Advances in Neural Information Processing Systems, 36:29914–29934, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Segment anything in high quality.Advances in Neural Information Processing Systems, 36:29914–29934, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.093403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.093403Z digest=sha256:7e559893a53dc85b2ad2bb4081944005bdcd40cf24ff461d9a97503ef06433e7

Observation 9594d68a-989f-41aa-835f-5303e38fd1d0 · outbound

This paper cites Segment Anything.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Segment Anything

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.178747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.178747Z digest=sha256:f154d0ec156d2e4810c77fb1ff99a1fea3897723325815f2c39816cfe2ca4db3

Observation 02a13c13-2da8-45f8-bc92-16ded190fae2 · outbound

This paper cites Openimages: A public dataset for large-scale multi-label and multi-class image classification.Dataset available from https://github.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Openimages: A public dataset for large-scale multi-label and multi-class image classification.Dataset available from https://github

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.690256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:11.254087Z digest=sha256:14ec2e536397ca13fa4bc84b86db46c8b4f18f1b0dd42e6f76694cfce5303742

Observation dbc8b219-ba86-4eeb-a622-83fad787060b · outbound

This paper cites Shamma, Michael S.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Shamma, Michael S

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.558320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:11.338922Z digest=sha256:8cd9a1ec419e22da79c9249eaa2a58b00e479f3639b81556a58e0012b7ce7503

Observation 2343811e-b4e8-4161-99c3-e1eed22d929b · outbound

This paper cites Beyond mot: Semantic multi-object tracking.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Beyond mot: Semantic multi-object tracking

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.413523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:11.411250Z digest=sha256:011c974cb379986a61ff1d70724a91e34541c115ee225ad413f22de82e4ef9ed

Observation 3de15e83-db5f-4f6d-8e47-8de78f50c331 · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Describe Anything: Detailed Localized Image and Video Captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.485753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.485753Z digest=sha256:e9b938de1c83c8b561bc4679e23e4f6b8e48531cf09acd52c74ba58dcfa0a69b

Observation a5acc589-816e-4212-9a41-05c03a5b85cc · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Rouge: A package for automatic evaluation of summaries

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.573027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.573027Z digest=sha256:3b3677dace00322e88137098ed55c2c200053be1ca741c21cbdc159bd42c0d6e

Observation 77c18e82-2491-4fb3-a229-e0f3415129a5 · outbound

This paper cites Lawrence Zitnick, and Piotr Dollár.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Lawrence Zitnick, and Piotr Dollár

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.641278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.641278Z digest=sha256:a952f867d62e2a0c5d81e7ae6fb40790648ad6df8eb2acdd0847926849d95225

Observation 734693f4-cdbe-494d-a851-2c3d8c130a24 · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.715662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.715662Z digest=sha256:8699ee0cb0d2d54579477fceb56da6f4d49ee4468948103d6b9a1a8644377bf3

Observation af9be36d-cf66-4a61-ae61-4bb17b033d4f · outbound

This paper cites DeepSeek-V3 Technical Report.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos DeepSeek-V3 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.759586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.759586Z digest=sha256:b72d44b0e52495169c557943e47360647fa1c8c0c09fbf259c0be3597e66c23f

Observation 9b80ee9d-3b88-48a8-aa56-1bae4df04a53 · outbound

This paper cites Gres: Generalized referring expression segmentation.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Gres: Generalized referring expression segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.842353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.842353Z digest=sha256:0f0d38f8ebfbb9c4a2a4c96015f4a9793e7088a951568b8a521155169e697868

Observation 9626591a-cf6b-4f0e-85b6-962987fec048 · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.917393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.917393Z digest=sha256:43f126992cfdc3ac1d86d890f701ec0460ff00c4377c29d8ef8dfe3713f25b55

Observation d60893a0-c1ee-47ba-8b40-24254752769b · outbound

This paper cites Improved baselines with visual instruction tuning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Improved baselines with visual instruction tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.028467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.028467Z digest=sha256:d2ff7e1aa6272acf2b3cce889bfc6b5110f1093feba067cbfeca49cc2b4bbf3e

Observation 2f5a9006-e068-4a88-9059-5b2f4069a2bb · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.090612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.090612Z digest=sha256:ba9af2504ff63471d55efd88a0142620097e085b168eae1b3799c304fb2de056

Observation 40857708-b9b1-43ac-8f63-68cc2bbdbcff · outbound

This paper cites The 2017 davis challenge on video object segmentation, 2018.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos The 2017 davis challenge on video object segmentation, 2018

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.218820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:12.161483Z digest=sha256:0d563b3fb42e24ee97e952b648b70904b8de66ef55d6393aefd54c9d2a4b5f7d

Observation dca6c900-6108-4828-b8be-4fda60a8bde5 · outbound

This paper cites Artemis: Towards referential understanding in complex videos.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Artemis: Towards referential understanding in complex videos

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.045839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:12.252891Z digest=sha256:47bddd5d1c24398cec4d9be04f0fa2690f64aa8432ad8129c758a4c26cce45b3

Observation 0411d034-bbf6-4397-8a7b-2527829345fb · outbound

This paper cites Artemis: Towards referential understanding in complex videos, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Artemis: Towards referential understanding in complex videos, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:18.841719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:12.318885Z digest=sha256:2e361c9719f0669d43e969a5fb96533d47784d5da4ea0d492bd11e104bc1851a

Observation 615de475-4912-4b52-843a-076e6b94c72f · outbound

This paper cites Paco: Parts and attributes of common objects, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Paco: Parts and attributes of common objects, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:18.669927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:12.398651Z digest=sha256:c4cc388db0d813ae86d08518c0070777f0a4b01267735820f8e1d23f94e421b3

Observation 3b52654c-8e87-4765-9961-b934d46bc038 · outbound

This paper cites Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.438893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.438893Z digest=sha256:34968b30944830f40950f744b90a0f356860da222a32cef9bffaeae2aa21af8b

Observation 80503a7d-ce6a-4a73-abf5-43bb8717d469 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos SAM 2: Segment Anything in Images and Videos

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.549714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.549714Z digest=sha256:ef71cbe09e5faea3779a263bb71842b8369fe534fea5337cbdec4fcbaa58d626

Observation 77d3f82b-7679-4289-a451-10f39f150192 · outbound

This paper cites Sam 2: Segment anything in images and videos, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Sam 2: Segment anything in images and videos, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.604590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.604590Z digest=sha256:72aeea4aa14eedf61a7b630a8f2d45cffacf078117e553c7207353d1a295854f

Observation dc7a87ef-9de9-4a0f-a155-21f4c3590af6 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.695638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.695638Z digest=sha256:5c990704a2ef4950ec875ce8f5bf525c3d45c1f6e89e9fb02147a734067ecb4b

Observation 53fdc3f9-85b0-41cd-9010-687b9512139e · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Objects365: A large-scale, high-quality dataset for object detection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.752865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.752865Z digest=sha256:3dc9d9c1ee2e7f05824f0c554a353e6917d2b16b6bc8f6a83b8a6386f185bcd6

Observation bd84cbc5-8476-4f71-9d32-a5ca9a7245e6 · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text, 2021.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text, 2021

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:18.501128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:12.830981Z digest=sha256:9b6694f6924208d25d87e2b43050a8a6954e501fe861e90dc692744891810ccc

Observation 003298e2-5e3e-4286-a1da-025dd32bb251 · outbound

This paper cites Icdar 2019 competition on large-scale street view text with partial labeling – rrc-lsvt, 2019.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Icdar 2019 competition on large-scale street view text with partial labeling – rrc-lsvt, 2019

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:18.357067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:12.899875Z digest=sha256:66c8bd86d4422bf642c8f3105e743c40a3eaeaf2e523e473cf50857b9d5c960f

Observation 97ae4dcd-4b1e-4f1b-90fd-60373548ef85 · outbound

This paper cites Human-centric spatio-temporal video grounding with visual transformers, 2021.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Human-centric spatio-temporal video grounding with visual transformers, 2021

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:18.211229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:12.974590Z digest=sha256:4fad73acc511c52a3ba0cd7e2b8ab3729d215770529ee13bdfa3264208b177af

Observation f81eebbf-4a2b-4617-958f-4e2a0ea98be1 · outbound

This paper cites Human- centric spatio-temporal video grounding with visual transformers.IEEE Transactions on Circuits and Systems for Video Technology, 32(12):8238–8249, 2021.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Human- centric spatio-temporal video grounding with visual transformers.IEEE Transactions on Circuits and Systems for Video Technology, 32(12):8238–8249, 2021

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:18.042690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:13.075724Z digest=sha256:a521fc076a45d3dfb7c4f704093d9622459941930344b4e328ceb1483ced0d7a

Observation a6c7bb8a-736d-4403-b5a3-847e4dbdc50a · outbound

This paper cites Cider: Consensus-based image description evaluation.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Cider: Consensus-based image description evaluation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.106006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.106006Z digest=sha256:16ce6e5024c427c1f21274d06fb31c3b44d1f5d75a946d85d90e86e7a26c1f8d

Observation 6538d8ce-016e-4755-9409-05c67de2a623 · outbound

This paper cites Coco-text: Dataset and benchmark for text detection and recognition in natural images, 2016.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Coco-text: Dataset and benchmark for text detection and recognition in natural images, 2016

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:17.871637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:13.197387Z digest=sha256:ad7a074cb7adb350e4645aaaa79db43f485a5d71033c2c920aa6b5b2bef9b63a

Observation 2be0f5f2-0406-4073-b936-c8820e72116c · outbound

This paper cites Elysium: Exploring object-level perception in videos via mllm, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Elysium: Exploring object-level perception in videos via mllm, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:17.712210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:13.272446Z digest=sha256:16970aa24a16936450e6048a8f1b4ff84313078aaeac91b09a39a0a9cc073e77

Observation 86242f7d-353d-4fce-9401-a217cb7e31d1 · outbound

This paper cites Elysium: Exploring object-level perception in videos via mllm.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Elysium: Exploring object-level perception in videos via mllm

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:17.544629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:13.376288Z digest=sha256:11b3776439a72167161235a986a52470140c6809f5f1106e7d10d919a92782d7

Observation d27a146f-4d23-4b7a-8f8a-50063baddd19 · outbound

This paper cites Towards open-vocabulary video instance segmentation, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Towards open-vocabulary video instance segmentation, 2023

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:17.380342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:13.408456Z digest=sha256:912ea97febe3e9f4a94ffd0b7a2fee3b2750751a86851df089cf9e9912263b71

Observation 24ea961e-be69-4b03-8336-d84a91f841d6 · outbound

This paper cites Git: A generative image-to-text transformer for vision and language, 2022.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Git: A generative image-to-text transformer for vision and language, 2022

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.430957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.430957Z digest=sha256:79778f296f49d7a15df2ec15beafdf7fbc4fcedd7ef9beff855492c3d67ceddd

Observation b4f51c45-574e-48be-87e2-b80093e76b7c · outbound

This paper cites V3det: Vast vocabulary visual detection dataset, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos V3det: Vast vocabulary visual detection dataset, 2023

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:17.243065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:13.449740Z digest=sha256:8dfc08b277a97d2e3db4303df51a9652407535c690b444a29379584498826f1b

Observation 64243a39-bee3-4aca-b170-7f18f7ea9eea · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.474745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.474745Z digest=sha256:3a93c9be1baf1cde859e5dd936c7c9d9c8204003624fd1ea8db71e2d3461cae0

Observation 7b4824b5-1247-46f0-9acb-6b6f537991e6 · outbound

This paper cites The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.496607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.496607Z digest=sha256:821aba0a1f97e379cbfb8c2ec37137a298f33799e8cb31e3ac13ce8fb2004901

Observation ce98c32a-3bee-4a17-b336-b03bff860f17 · outbound

This paper cites Grit: A generative region-to-text transformer for object understanding.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Grit: A generative region-to-text transformer for object understanding

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:17.065248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:13.564740Z digest=sha256:714f0acd9547d1adb16c7b4f4f6920cfa350df8519353bcad913371e246f9d67

Observation feb50aa7-36ce-4526-9e51-e569d7c557cb · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.Advances in Neural Information Processing Systems, 37:69925–69975, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.Advances in Neural Information Processing Systems, 37:69925–69975, 2024

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.586867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.586867Z digest=sha256:68651879ed1c222ac51991053252fd59cffd853be8c19743a57aac8a4d461b73

Observation 4e4340d2-d0db-4de9-ba17-cf974d152d9f · outbound

This paper cites Youtube-vos: A large-scale video object segmentation benchmark, 2018.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Youtube-vos: A large-scale video object segmentation benchmark, 2018

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:16.863847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:13.605063Z digest=sha256:919a12d78ef0225a32f45718fc32b7108b95ba08040c3fc751de0e4dc60790e7

Observation 6306acd4-37f1-40c5-8523-43cfcc258b31 · outbound

This paper cites Qwen2.5 Technical Report.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Qwen2.5 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.643014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.643014Z digest=sha256:fe093bb3222bef8b679eae49232d701be5239cbe224afcc1d8cf250dd3a12a95

Observation bba35fb9-1d5f-43a0-9e47-1b2befb8714e · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Vid2seq: Large-scale pretraining of a visual language model for dense video captioning, 2023

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:16.586598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:13.676752Z digest=sha256:ce2c53cafcdf980d11d5682c1388832f13c557ee793c0c8a1a0328fa0b477ef1

Observation 038aa238-4d5e-452e-897b-8f9fd415f96a · outbound

This paper cites Samurai: Adapting segment anything model for zero-shot visual tracking with motion-aware memory, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Samurai: Adapting segment anything model for zero-shot visual tracking with motion-aware memory, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:16.410525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:13.714771Z digest=sha256:03c9af5bcb547b6dd671f1c65bf50375c9bdb823b5951ca05f656ddbe991d64c

Observation 6f24cf1c-e96e-498b-94ec-f38cf6b5d0a6 · outbound

This paper cites Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.738804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.738804Z digest=sha256:4e0af1a3305d45e4dfc21a3db008dc75428e726240bc0ad74e48834597f857e1

Observation 7b581093-f66b-4a18-8762-233b527d2abf · outbound

This paper cites Detecting texts of arbitrary orientations in natural images.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Detecting texts of arbitrary orientations in natural images

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:16.269910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:13.788564Z digest=sha256:0abf7c7c9a66b160e26b3bf28e496eebae6c03f29b00a2817890f2f04a721aa2

Observation 15915057-4c0d-42e2-9ead-e8bb545702d5 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.842739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.842739Z digest=sha256:0e7245317b468bc051e9ed55d020021fcb2bf0f00118125d03541cf0137b7ff6

Observation 54f1de86-fd8f-4e15-ac4e-e612520875a4 · outbound

This paper cites Merlin: Empowering multimodal llms with foresight minds.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Merlin: Empowering multimodal llms with foresight minds

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.904324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.904324Z digest=sha256:a4ed9fdffef4398e83537982548f6e4e98268163663c84143a9b6a4bd885bc12

Observation 04667787-33fe-4d7a-b681-83be6aea276a · outbound

This paper cites Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos, 2025.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos, 2025

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:16.098427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:13.964588Z digest=sha256:ce22c98bb0766127fbdb022a2c441ad0a724ff705ae661090eede63b0e7225dc

Observation 3532d1b8-250b-4955-bf83-c62da59b9e9e · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Osprey: Pixel understanding with visual instruction tuning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.049510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.049510Z digest=sha256:4c3acb7062bba5ea040fde5ef362d635db78aeef87496faf808c274dd4633279

Observation 60a624db-a6fa-43c6-aedc-ac269ea708c2 · outbound

This paper cites VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.117492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.117492Z digest=sha256:09b625ea2b7edaa5b29cecce095192d1bcd7f610bd39fa1733fd105c78e53813

Observation 86547f57-af2f-49a4-aa87-4cd8e57a6fa5 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.194914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.194914Z digest=sha256:c83f26b72a7338d9edacbd40f8959a5a013016f4b17752e7232a008dfc5a4918

Observation ae642888-c3de-4cbe-aa24-3161505b0301 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.279027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.279027Z digest=sha256:b1483f25f8a9f2e4b4f08dcf47b9b2b55ea587c25f3f5a5ee21a4ef13de05b74

Observation e377b8a5-dc62-43a2-b95e-5499da6488b8 · outbound

This paper cites Gpt4roi: Instruction tuning large language model on region-of-interest, 2025.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Gpt4roi: Instruction tuning large language model on region-of-interest, 2025

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:15.913568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:14.380692Z digest=sha256:60622be2cd527d6e7d03902000810b734505a87c8a196b35a933a765270c9cd0

Observation 614da7e3-d07b-487c-875c-d939bf00e4fa · outbound

This paper cites Where does it exist: Spatio-temporal video grounding for multi-form sentences.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Where does it exist: Spatio-temporal video grounding for multi-form sentences

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:15.744187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:14.500909Z digest=sha256:202c07772fd24b79fc90dd1da38532e3d7440e0e1298dde173265b99fb1f627b

Observation ccdc906f-94a4-4d76-bf34-91859924d5ba · outbound

This paper cites ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.565949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.565949Z digest=sha256:61f4f2cf79e07ff1a071e901232b47a4d45b0e32138fa35320b52d85b3e2bc30

Observation 037adc13-83fb-4e76-a27e-08af8718a73b · outbound

This paper cites Fast Segment Anything.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Fast Segment Anything

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.704495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.704495Z digest=sha256:4dbe8552dd49e1ea23a4eac909bc2c8b3ae03c4b06766d0fad6c5013bb066413

Observation 1b5a5567-5bb8-4335-a8ee-4977daff2dd2 · outbound

This paper cites Controlcap: Controllable region-level captioning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Controlcap: Controllable region-level captioning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:15.590718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:14.776221Z digest=sha256:7d8a07c306a539401aa4ac5b8a8f4dd4a8ca01d613f0ba38200434b668ccb4c9

Observation 103c9006-1579-4b31-b1f1-df970f5937e1 · outbound

This paper cites Streaming dense video captioning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Streaming dense video captioning

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:15.393823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:28:14.907661Z digest=sha256:4f6a7cd627587b8ed563b4db5ec7856335cf8e8c0d4dca5fbfd400a690c2aadc

Pith citing papers

Observation 1309521c-a60a-444f-ba0b-13861ed211f6 · inbound

Describe Anything Model for Visual Question Answering on Text-rich Images cites this paper.

Describe Anything Model for Visual Question Answering on Text-rich Images Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:24.717294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:24.717294Z digest=sha256:043cd48806bd532bd9bfa5a6e3ad89f3275adc4926c0c43f1f574d45cd1ab170

Observation 54b3d678-10f0-4408-b6bb-06bdb6520ea5 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 293

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:11.675850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:11.675850Z digest=sha256:2dbb9d96179e6f92b6743109c5ee87c2c5960d5a321fd493544ccf0c805e9b62

Observation 2b0f3896-8ba2-43c2-ac5a-aa020adf157d · inbound

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data cites this paper.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.722899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.722899Z digest=sha256:6dfa806711003fb33298ffc6ab8207bf0af0b4b5e19db30adbed1aa01bba6360

Observation 5bdff3cf-8f46-4932-b4a7-efc5fd55bbdc · inbound

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension cites this paper.

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T18:15:05.829843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:15:05.829843Z digest=sha256:a3e67ffb5290b8fe5ef36155e9a43c3d85edebee3d63c63b1b0a344ea182a6c6

Observation 64af0091-4638-4aaa-8598-57b8cbcfa627 · inbound

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition cites this paper.

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:21:23.379690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T00:20:58.483350Z digest=sha256:d2a93be0f323a5478f05eed664e55926bfb4e6967a35f4d476a7d2571f236329

Observation f78f9f04-29ac-4021-89e9-4f639e9d59ea · inbound

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling cites this paper.

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:41:19.224758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:39:32.955779Z digest=sha256:e2355077e0b1bf8f8f370ffeee64106f368341a9fddcccfe9ce8b278053623f5

Observation 59ec8268-c8fa-4276-9aee-c8c8bb744edd · inbound

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling cites this paper.

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:37:05.085568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:37:05.085568Z digest=sha256:db264e4adb435338a7494c71a825ef8c654b7fd528a6d3b4b15e0769fc35b247

Observation 4d313a4e-0c65-4f54-b147-e665b3c3d17c · inbound

Enhancing Foundation VLM Robustness to Missing Modality: Scalable Diffusion for Bi-directional Feature Restoration cites this paper.

Enhancing Foundation VLM Robustness to Missing Modality: Scalable Diffusion for Bi-directional Feature Restoration Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:37:37.229588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:33:49.841678Z digest=sha256:2c06843c31ceb0af3ad7dc60b798611b9db59ba2b04ae2b74b2c0ffdc4ea13a8

Observation 728775e5-abdc-40d7-84b4-a2b3a515999e · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:45:49.048227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:93698570d377e7e70d34b24d4b91859e0482538ebbebc5c4cf508ee9e5d228eb

Observation 87f48f3a-8a30-4c2e-baec-3acd0e31e9bc · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:d905d375129a96a1c050991fc57f8c238e3f617419f87de73551406377e8ab7b

Observation c25c039a-fdd0-4fdc-a0c6-8609bd93d72d · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:06.525843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:c0a028a36e1d6c8c0d75bd2965f069806b67252cede469f88107dbe4150a4286

Observation fecadecc-bf3c-4c06-850e-e596473f3e66 · inbound

WOW-Seg: A Word-free Open World Segmentation Model cites this paper.

WOW-Seg: A Word-free Open World Segmentation Model Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:27:48.052668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:23:14.311122Z digest=sha256:940edd03f3fe96f6968367f60b2dc5956391920e7bac45f89bc6c3c7f5dfc8f0

Observation d1ce6cc2-ff2d-43fa-bcdf-2036aa9755d6 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.316928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:77d6f768bc26e0ba8b0b08e929965d9859538aefc91867e5462661dd3bfbf02e

Observation 694d2d02-f883-4042-9944-73eace34e136 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.783743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:c278e7fdebeff7e7a08211df010ccf79f5614dd3d8f3d17311bbb811d1583d39

Observation f067d0b1-ed72-4548-b86f-4a796ce217ed · inbound

FRFDet: Efficient UAV Small Object Detection with Symmetric Sampling and Scalable Fusion cites this paper.

FRFDet: Efficient UAV Small Object Detection with Symmetric Sampling and Scalable Fusion Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T21:31:04.773453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:31:04.773453Z digest=sha256:c164f4ce2b5a56842319431d82a7bf0b2fc06f3f8f8e770ec2b978e18ace9a9b