Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:04:29.255261Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2501.11653.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:04:29.255261Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation aac75950-44cf-4b22-a26d-c667546e7da8 · outbound
Dynamic Scene Understanding from Vision-Language Representations Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 95eef835-4dac-40c8-ad97-aca0a8909734 · outbound
Dynamic Scene Understanding from Vision-Language Representations Learning human- human interactions in images from weak textual supervision
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2e3a36a5-775a-497f-a6de-cf45e09c9500 · outbound
Dynamic Scene Understanding from Vision-Language Representations Zero-shot learning via visual abstraction
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 631fb98c-5f95-4ebc-a278-141f848b5558 · outbound
Dynamic Scene Understanding from Vision-Language Representations Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01422b37-2dda-430c-9b9b-e0508b190b53 · outbound
Dynamic Scene Understanding from Vision-Language Representations On the Opportunities and Risks of Foundation Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f63c90eb-eab9-483f-bdf6-d8980f63670e · outbound
Dynamic Scene Understanding from Vision-Language Representations End-to- end object detection with transformers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bfd6cce4-ae8b-4d02-93b2-e246e6da1e8b · outbound
Dynamic Scene Understanding from Vision-Language Representations Quo vadis, action recognition? a new model and the kinetics dataset
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 57c6dedb-17bf-41ee-b038-dbacff56251d · outbound
Dynamic Scene Understanding from Vision-Language Representations A comprehensive survey of scene graphs: Generation and application
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8b25b575-0c75-45d3-af43-bd3c774dfaf2 · outbound
Dynamic Scene Understanding from Vision-Language Representations Hico: A benchmark for recognizing human-object interactions in images
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb4750d0-d381-41ab-a303-8b1fe6e0c5a8 · outbound
Dynamic Scene Understanding from Vision-Language Representations Learning to detect human-object interactions
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8ecd477b-6191-482f-b123-7c1ebd7e0fde · outbound
Dynamic Scene Understanding from Vision-Language Representations ViStruct: Visual structural knowledge extrac- tion via curriculum guided code-vision representation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 703cb924-8b57-4947-b998-0bb4fd614917 · outbound
Dynamic Scene Understanding from Vision-Language Representations Collab- orative transformers for grounded situation recognition
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c3e88180-8913-47f4-b6cc-6dce707d9aef · outbound
Dynamic Scene Understanding from Vision-Language Representations Imagenet: A large-scale hierarchical image database
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc8823df-6244-410f-8731-49f1861b40b0 · outbound
Dynamic Scene Understanding from Vision-Language Representations An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2a25918-65f4-498b-aa88-02903cfe5efb · outbound
Dynamic Scene Understanding from Vision-Language Representations Background to framenet
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c50a6fcd-c9b3-4f4a-910f-d266a34cc0dd · outbound
Dynamic Scene Understanding from Vision-Language Representations Deep residual learning for image recognition
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d36cdb7d-015d-4418-ac90-b97eb910394b · outbound
Dynamic Scene Understanding from Vision-Language Representations LoRA: Low-Rank Adaptation of Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b4871c5-eb6a-406d-bbe7-8768b5cae342 · outbound
Dynamic Scene Understanding from Vision-Language Representations Language is not all you need: Aligning perception with language mod- els
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aec5140f-b0fe-4f2f-a6f8-e9feafad167f · outbound
Dynamic Scene Understanding from Vision-Language Representations What to look at and where: Se- mantic and spatial refined transformer for detecting human- object interactions
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 98cf085f-0afa-468b-8912-494579e34914 · outbound
Dynamic Scene Understanding from Vision-Language Representations Scaling up visual and vision-language representa- tion learning with noisy text supervision
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a19a440-f42a-4ffd-a385-85b6b616842c · outbound
Dynamic Scene Understanding from Vision-Language Representations Detrs with hybrid matching
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a9f55ce5-6d81-41ca-8df5-214fd0312358 · outbound
Dynamic Scene Understanding from Vision-Language Representations Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f0b8a0b-db11-4d83-b250-48c870406467 · outbound
Dynamic Scene Understanding from Vision-Language Representations Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 69f40340-b426-4e15-8022-a297db112620 · outbound
Dynamic Scene Understanding from Vision-Language Representations Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9b6046be-2011-4ef4-a4bd-bc198685ad87 · outbound
Dynamic Scene Understanding from Vision-Language Representations Gen-vlkt: Simplify association and enhance 9 interaction understanding for hoi detection
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c75e00cd-606a-454b-89dc-adbcbd5198a3 · outbound
Dynamic Scene Understanding from Vision-Language Representations Microsoft coco: Common objects in context
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e587661f-7457-4d75-b58f-b5f7adca23f5 · outbound
Dynamic Scene Understanding from Vision-Language Representations Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 203f3056-d11b-457e-9bc2-9ece97e445d0 · outbound
Dynamic Scene Understanding from Vision-Language Representations Visual instruction tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 173a6cd4-a4b5-47de-8e6f-d99d2fd8b58e · outbound
Dynamic Scene Understanding from Vision-Language Representations Visual instruction tuning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28eadd23-0bcd-4653-ad1c-edd404ef1899 · outbound
Dynamic Scene Understanding from Vision-Language Representations Open scene understanding: Grounded situation recognition meets segment anything for helping people with visual im- pairments
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41189d75-9cf3-48b8-8ba5-e29d2e20b99e · outbound
Dynamic Scene Understanding from Vision-Language Representations Image segmenta- tion using text and image prompts
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1b30bbb4-4a3a-462d-90ea-c3b706b46657 · outbound
Dynamic Scene Understanding from Vision-Language Representations Clip4hoi: Towards adapting clip for practi- cal zero-shot hoi detection
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ff2e719c-5f1c-46f5-b7f8-5eebdb0b2c27 · outbound
Dynamic Scene Understanding from Vision-Language Representations ClipCap: CLIP Prefix for Image Captioning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2de3706-8bd1-40e1-953d-94b84f2f6800 · outbound
Dynamic Scene Understanding from Vision-Language Representations Hoiclip: Efficient knowledge transfer for hoi detection with vision-language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7827cb2b-6ab9-49b2-a6e6-498547c4e7a9 · outbound
Dynamic Scene Understanding from Vision-Language Representations DINOv2: Learning Robust Visual Features without Supervision
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01ecefc0-edab-4d95-bab2-38ae20cf08c2 · outbound
Dynamic Scene Understanding from Vision-Language Representations Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0b6bd39-9a9c-4e0f-b8e5-92328280020d · outbound
Dynamic Scene Understanding from Vision-Language Representations Grounded situation recognition
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 36fd30a8-89be-4a73-ac93-645cafdb9a72 · outbound
Dynamic Scene Understanding from Vision-Language Representations Learn- ing transferable visual models from natural language super- vision
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 64b6a403-d32d-4f5e-8218-9aa542878d80 · outbound
Dynamic Scene Understanding from Vision-Language Representations Robotic vision for human-robot interaction and collaboration: A survey and systematic re- view
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d350ddd5-f586-429d-8707-baa5abe9cd7d · outbound
Dynamic Scene Understanding from Vision-Language Representations High-resolution image synthesis with latent diffusion models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f45f3ff3-3435-493c-87a2-7df579bb0a82 · outbound
Dynamic Scene Understanding from Vision-Language Representations Clip- situ: Effectively leveraging clip for conditional predictions in situation recognition
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 127bb760-d0e2-4bf0-b721-374d14b86ab8 · outbound
Dynamic Scene Understanding from Vision-Language Representations Ut-interaction dataset, icpr contest on semantic description of human activities (sdha)
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 97dbb885-b995-46bc-990c-daac2e0537cd · outbound
Dynamic Scene Understanding from Vision-Language Representations BLEURT: Learning Robust Metrics for Text Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3816ec7c-6c74-4cb7-b9cb-efa396178278 · outbound
Dynamic Scene Understanding from Vision-Language Representations Analyzing human– human interactions: A survey
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 10e15eda-365b-4947-b854-62f591b869b4 · outbound
Dynamic Scene Understanding from Vision-Language Representations Generative multimodal mod- els are in-context learners
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f40e930-6501-479f-aff6-e125cda380c4 · outbound
Dynamic Scene Understanding from Vision-Language Representations Qpic: Query-based pairwise human-object interaction detec- tion with image-wide contextual information
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e9cac3-da5f-49a0-9d3a-05c94c91d1fc · outbound
Dynamic Scene Understanding from Vision-Language Representations Learning latent temporal structure for complex event detection
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 86b04a56-3d0e-4dee-bdb7-443a672c3542 · outbound
Dynamic Scene Understanding from Vision-Language Representations Learning spatiotemporal features with 3d convolutional networks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1d9f3e91-2a1a-4368-9bce-da217fdc3560 · outbound
Dynamic Scene Understanding from Vision-Language Representations ActionCLIP: A New Paradigm for Video Action Recognition
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d29cfbf7-2dcc-4779-aaab-4305283204c6 · outbound
Dynamic Scene Understanding from Vision-Language Representations Cris: Clip- driven referring image segmentation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8357c391-7ff1-48fc-8cb8-d09e3603ec84 · outbound
Dynamic Scene Understanding from Vision-Language Representations Rethinking the two-stage framework for grounded sit- uation recognition
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4d4757c7-835d-4760-bafa-ec89c30a47a7 · outbound
Dynamic Scene Understanding from Vision-Language Representations Rec- ognize complex events from static images by fusing deep channels
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 40e3df64-9735-4a28-aec7-cafe846af708 · outbound
Dynamic Scene Understanding from Vision-Language Representations Recognizing proxemics in personal photos
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d9a2302c-2e38-4d47-8166-9ab607be7c5d · outbound
Dynamic Scene Understanding from Vision-Language Representations Situa- tion recognition: Visual semantic role labeling for image understanding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1fc5b6a0-4560-4c77-94cf-9e8393e3030d · outbound
Dynamic Scene Understanding from Vision-Language Representations Rlip: Rela- tional language-image pre-training for human-object inter- action detection
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a632b5bf-5002-4f11-9363-c366041d74ac · outbound
Dynamic Scene Understanding from Vision-Language Representations Rlipv2: Fast scaling of re- lational language-image pre-training
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a3ee7e89-a783-4f6c-9663-796d404fb7b0 · outbound
Dynamic Scene Understanding from Vision-Language Representations Lit: Zero-shot transfer with locked-image text tuning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a96dc666-14e7-48e0-b7b5-af39e5851057 · outbound
Dynamic Scene Understanding from Vision-Language Representations Zhang, Dylan Campbell, and Stephen Gould
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4fd9d762-a0b8-4a76-a702-5de0f1674ac7 · outbound
Dynamic Scene Understanding from Vision-Language Representations Zhang, Yuhui Yuan, Dylan Campbell, Zhuoyao Zhong, and Stephen Gould
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 903ec1a6-83d9-435b-ae74-c067191ce6b8 · outbound
Dynamic Scene Understanding from Vision-Language Representations OPT: Open Pre-trained Transformer Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e5280f-c78a-4577-9639-a71346507a45 · outbound
Dynamic Scene Understanding from Vision-Language Representations Zegclip: Towards adapting clip for zero-shot se- mantic segmentation
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 93ec1946-844a-4ce5-b39c-6318c576862c · outbound
Dynamic Scene Understanding from Vision-Language Representations Scene Graph Generation: A Comprehensive Survey
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dc5a034-c241-4ee0-b834-9ed607815125 · outbound
Dynamic Scene Understanding from Vision-Language Representations A Comprehensive Study of Deep Video Action Recognition
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6596321f-a215-413d-98bc-57a47de3d03b · outbound
Dynamic Scene Understanding from Vision-Language Representations Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1356fec0-a3f6-4e6c-96b2-c9188483b343 · outbound
Dynamic Scene Understanding from Vision-Language Representations Unresolved cited work
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
No inbound Pith citation observations are available.