Pith. sign in

Paper Citation Record · LEDGER

Dynamic Scene Understanding from Vision-Language Representations

As of 14 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2501.11653.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.11653 v3

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:04:29.255261Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy39
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aac75950-44cf-4b22-a26d-c667546e7da8 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Dynamic Scene Understanding from Vision-Language Representations Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.859186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.058653Z digest=sha256:21ea3a02fe604f629c1af1daf9c9a2a98e00d47f2e73af04053ca5e5df7df34c

Observation 95eef835-4dac-40c8-ad97-aca0a8909734 · outbound

This paper cites Learning human- human interactions in images from weak textual supervision.

Dynamic Scene Understanding from Vision-Language Representations Learning human- human interactions in images from weak textual supervision

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.847754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.062198Z digest=sha256:7b489113ec587b00578b75aa0b18179adcce7809f4767a93796dcc86f1b6e925

Observation 2e3a36a5-775a-497f-a6de-cf45e09c9500 · outbound

This paper cites Zero-shot learning via visual abstraction.

Dynamic Scene Understanding from Vision-Language Representations Zero-shot learning via visual abstraction

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.835676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.065276Z digest=sha256:089c3bf8f03cb01666e2eb395613f379e9943c364f4e07a82de0545aba9947e4

Observation 631fb98c-5f95-4ebc-a278-141f848b5558 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Dynamic Scene Understanding from Vision-Language Representations Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.068523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.068523Z digest=sha256:8ab609d3267888aa122685fd4a8cad41b49878d57be909dc022fccbbffa27221

Observation 01422b37-2dda-430c-9b9b-e0508b190b53 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Dynamic Scene Understanding from Vision-Language Representations On the Opportunities and Risks of Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.072486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.072486Z digest=sha256:4db688517f6834849cef0b94c65b9ee719c6a4d6e14108fe01a582e5acd582e7

Observation f63c90eb-eab9-483f-bdf6-d8980f63670e · outbound

This paper cites End-to- end object detection with transformers.

Dynamic Scene Understanding from Vision-Language Representations End-to- end object detection with transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.824996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.075724Z digest=sha256:cd6733e376d97169e09bd3ac265d522d8ab0e416bab753fb78d35504742474e0

Observation bfd6cce4-ae8b-4d02-93b2-e246e6da1e8b · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Dynamic Scene Understanding from Vision-Language Representations Quo vadis, action recognition? a new model and the kinetics dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.814499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.078754Z digest=sha256:b49f218b696fc4f9ca18b9848db17915c7b3c50d492927e6e4f0971b301353ae

Observation 57c6dedb-17bf-41ee-b038-dbacff56251d · outbound

This paper cites A comprehensive survey of scene graphs: Generation and application.

Dynamic Scene Understanding from Vision-Language Representations A comprehensive survey of scene graphs: Generation and application

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.804672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.081798Z digest=sha256:611a91e24e9cace62bb87d7ac32ab91a5dc4fae03146e5a1089f87c04cae6a63

Observation 8b25b575-0c75-45d3-af43-bd3c774dfaf2 · outbound

This paper cites Hico: A benchmark for recognizing human-object interactions in images.

Dynamic Scene Understanding from Vision-Language Representations Hico: A benchmark for recognizing human-object interactions in images

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.084523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.084523Z digest=sha256:eb631dd023c68b17ed607d022824888759036c771bb6a0f8f3ae8fa3827b4e3e

Observation eb4750d0-d381-41ab-a303-8b1fe6e0c5a8 · outbound

This paper cites Learning to detect human-object interactions.

Dynamic Scene Understanding from Vision-Language Representations Learning to detect human-object interactions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.789523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.087364Z digest=sha256:cc7e2ccbc82f063e41e76c9ac3bf8d6922989cc06a2849d1e888848afc105c15

Observation 8ecd477b-6191-482f-b123-7c1ebd7e0fde · outbound

This paper cites ViStruct: Visual structural knowledge extrac- tion via curriculum guided code-vision representation.

Dynamic Scene Understanding from Vision-Language Representations ViStruct: Visual structural knowledge extrac- tion via curriculum guided code-vision representation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.779720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.090004Z digest=sha256:63e23330a47d9d685c0e977cf9e3e5627b256b2b69bff74618c6917d91f7d862

Observation 703cb924-8b57-4947-b998-0bb4fd614917 · outbound

This paper cites Collab- orative transformers for grounded situation recognition.

Dynamic Scene Understanding from Vision-Language Representations Collab- orative transformers for grounded situation recognition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.770916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.093120Z digest=sha256:7f4db9eead3825a0c2249aa438991b3a63e019dcde66d2974532e54657dc18c9

Observation c3e88180-8913-47f4-b6cc-6dce707d9aef · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Dynamic Scene Understanding from Vision-Language Representations Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.096460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.096460Z digest=sha256:18242b13403d16054b6d1395533381dca3ccf5e2ac5675970ffe06c0e418e664

Observation fc8823df-6244-410f-8731-49f1861b40b0 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Dynamic Scene Understanding from Vision-Language Representations An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.099971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.099971Z digest=sha256:215741d256ad3a67d06843333ca820778c564de6e16702a16711ccf9456bb184

Observation a2a25918-65f4-498b-aa88-02903cfe5efb · outbound

This paper cites Background to framenet.

Dynamic Scene Understanding from Vision-Language Representations Background to framenet

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.755417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.103524Z digest=sha256:8037d145628f30b0c68611a25617d572fa5d951d0dbeb6b4fdc831278d4d02c8

Observation c50a6fcd-c9b3-4f4a-910f-d266a34cc0dd · outbound

This paper cites Deep residual learning for image recognition.

Dynamic Scene Understanding from Vision-Language Representations Deep residual learning for image recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.107173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.107173Z digest=sha256:4bf281980061f517a2d3505c2b994a1f22d14e776c14f65379d90e905f3c794c

Observation d36cdb7d-015d-4418-ac90-b97eb910394b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Dynamic Scene Understanding from Vision-Language Representations LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.111156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.111156Z digest=sha256:3403d91e8c8d9bcbb8a2de9fedf8ab8391912d020f8c80b69eef71c8fe74947d

Observation 2b4871c5-eb6a-406d-bbe7-8768b5cae342 · outbound

This paper cites Language is not all you need: Aligning perception with language mod- els.

Dynamic Scene Understanding from Vision-Language Representations Language is not all you need: Aligning perception with language mod- els

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.740016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.114565Z digest=sha256:870d836f6e78e05cb0a3eb1cefa602b368cf50c8d4da6a280fe5b7b2a9924bfc

Observation aec5140f-b0fe-4f2f-a6f8-e9feafad167f · outbound

This paper cites What to look at and where: Se- mantic and spatial refined transformer for detecting human- object interactions.

Dynamic Scene Understanding from Vision-Language Representations What to look at and where: Se- mantic and spatial refined transformer for detecting human- object interactions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.729962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.117668Z digest=sha256:b313cdd07ef9c33d1c5e92e2d78d5cbbd5fb385da225c0680e7354eedd9e643a

Observation 98cf085f-0afa-468b-8912-494579e34914 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Dynamic Scene Understanding from Vision-Language Representations Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.121047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.121047Z digest=sha256:4a09cf3bf0c78cb26e76da30262ddc43d6b24e143480a0560bc1c31e162c9ee3

Observation 9a19a440-f42a-4ffd-a385-85b6b616842c · outbound

This paper cites Detrs with hybrid matching.

Dynamic Scene Understanding from Vision-Language Representations Detrs with hybrid matching

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.713768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.124643Z digest=sha256:fcf1151d4f83de5b7f2a07b95907af0d1377f65aefdbe3ca6187b7b67ddde421

Observation a9f55ce5-6d81-41ca-8df5-214fd0312358 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

Dynamic Scene Understanding from Vision-Language Representations Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.127943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.127943Z digest=sha256:56d60fbc5808c7188dd21da49f60e675137460f9ff3b6436d40da08ec9b185ec

Observation 0f0b8a0b-db11-4d83-b250-48c870406467 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Dynamic Scene Understanding from Vision-Language Representations Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.704084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.130917Z digest=sha256:1eebdbea440ee6e808a12a4ee1fa1d31d8389a423d490c5b529a6b481ab202b4

Observation 69f40340-b426-4e15-8022-a297db112620 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Dynamic Scene Understanding from Vision-Language Representations Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.693920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.133681Z digest=sha256:14de671813c616d5dc1820f9672ec92e6f626f64d35084504f06471e0e1e7bd3

Observation 9b6046be-2011-4ef4-a4bd-bc198685ad87 · outbound

This paper cites Gen-vlkt: Simplify association and enhance 9 interaction understanding for hoi detection.

Dynamic Scene Understanding from Vision-Language Representations Gen-vlkt: Simplify association and enhance 9 interaction understanding for hoi detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.683874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.136248Z digest=sha256:1caf43f1ac9e085f3707efc87fc296e9081115015984e07b398ae1161b1e7222

Observation c75e00cd-606a-454b-89dc-adbcbd5198a3 · outbound

This paper cites Microsoft coco: Common objects in context.

Dynamic Scene Understanding from Vision-Language Representations Microsoft coco: Common objects in context

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.674358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.139011Z digest=sha256:a336905fcf8b3d15a5e4814f2909e1813733e8ec24963156354eafdac90bfad4

Observation e587661f-7457-4d75-b58f-b5f7adca23f5 · outbound

This paper cites Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation.

Dynamic Scene Understanding from Vision-Language Representations Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.141758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.141758Z digest=sha256:6096d51f0f990401191ed564d9551fdebd4d0b07723b75267f216ebaf6894d56

Observation 203f3056-d11b-457e-9bc2-9ece97e445d0 · outbound

This paper cites Visual instruction tuning.

Dynamic Scene Understanding from Vision-Language Representations Visual instruction tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.144571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.144571Z digest=sha256:6d48168e8f31d75ed0b2869c93b8bf1934ec38a4432601a99d7b39f130386074

Observation 173a6cd4-a4b5-47de-8e6f-d99d2fd8b58e · outbound

This paper cites Visual instruction tuning.

Dynamic Scene Understanding from Vision-Language Representations Visual instruction tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.147466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.147466Z digest=sha256:a3c1205bc41a19d38eaa2c04c4e32896ae9537002edd930dd9d09d2c4c5b35b6

Observation 28eadd23-0bcd-4653-ad1c-edd404ef1899 · outbound

This paper cites Open scene understanding: Grounded situation recognition meets segment anything for helping people with visual im- pairments.

Dynamic Scene Understanding from Vision-Language Representations Open scene understanding: Grounded situation recognition meets segment anything for helping people with visual im- pairments

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.149991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.149991Z digest=sha256:160d198e178d1b9a660f465bf9f7cc9044e79efad34444f39404556d16c70ece

Observation 41189d75-9cf3-48b8-8ba5-e29d2e20b99e · outbound

This paper cites Image segmenta- tion using text and image prompts.

Dynamic Scene Understanding from Vision-Language Representations Image segmenta- tion using text and image prompts

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.641316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.152678Z digest=sha256:86571a9b60dad45680cf44832ac2ec2df5880ddbd802feda1973f4fb48f5a5ec

Observation 1b30bbb4-4a3a-462d-90ea-c3b706b46657 · outbound

This paper cites Clip4hoi: Towards adapting clip for practi- cal zero-shot hoi detection.

Dynamic Scene Understanding from Vision-Language Representations Clip4hoi: Towards adapting clip for practi- cal zero-shot hoi detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.631662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.155309Z digest=sha256:b84aabf142e57d692af9405bb91d3af1eafb1f5c4b2392af396efc89ae4a2110

Observation ff2e719c-5f1c-46f5-b7f8-5eebdb0b2c27 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

Dynamic Scene Understanding from Vision-Language Representations ClipCap: CLIP Prefix for Image Captioning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.158092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.158092Z digest=sha256:dc907e29778abeca34ea83f079ac970a5f431cbeb94b6af327da1cae2121f7ec

Observation c2de3706-8bd1-40e1-953d-94b84f2f6800 · outbound

This paper cites Hoiclip: Efficient knowledge transfer for hoi detection with vision-language models.

Dynamic Scene Understanding from Vision-Language Representations Hoiclip: Efficient knowledge transfer for hoi detection with vision-language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.622766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.161682Z digest=sha256:d3f0202264b3459ccc52cf5e7bbc7de8c86be4e3ea477af1f489f0edb359e6a4

Observation 7827cb2b-6ab9-49b2-a6e6-498547c4e7a9 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Dynamic Scene Understanding from Vision-Language Representations DINOv2: Learning Robust Visual Features without Supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.164848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.164848Z digest=sha256:33f1d1df6cde84263aad94023b2ec8c9cbdd56407cc3e4b015a9214b93e5c2ac

Observation 01ecefc0-edab-4d95-bab2-38ae20cf08c2 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Dynamic Scene Understanding from Vision-Language Representations Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.168429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.168429Z digest=sha256:089c0d40e4299b013665f76bbe8371c8795e59d546aff45973cf3ad7963059c5

Observation c0b6bd39-9a9c-4e0f-b8e5-92328280020d · outbound

This paper cites Grounded situation recognition.

Dynamic Scene Understanding from Vision-Language Representations Grounded situation recognition

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.613873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.172067Z digest=sha256:032b22839df8345a67130844f2f2620741bfbff16898b6d12a8c80e1d490a49e

Observation 36fd30a8-89be-4a73-ac93-645cafdb9a72 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Dynamic Scene Understanding from Vision-Language Representations Learn- ing transferable visual models from natural language super- vision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.604554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.175274Z digest=sha256:3bf2ddd8fbeb66ff6895e6af54a12dd20e8f2f4a42cc7297975d84d1f56bec00

Observation 64b6a403-d32d-4f5e-8218-9aa542878d80 · outbound

This paper cites Robotic vision for human-robot interaction and collaboration: A survey and systematic re- view.

Dynamic Scene Understanding from Vision-Language Representations Robotic vision for human-robot interaction and collaboration: A survey and systematic re- view

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.595466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.178468Z digest=sha256:dc3978ee7a34a2248fcdcdab9316ca5c7a530f7a9d13cf463eb033d7d33df557

Observation d350ddd5-f586-429d-8707-baa5abe9cd7d · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Dynamic Scene Understanding from Vision-Language Representations High-resolution image synthesis with latent diffusion models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.181632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.181632Z digest=sha256:99f3b0c245988fa01663c88e296df20be5f224363afe4f478adf22373392a828

Observation f45f3ff3-3435-493c-87a2-7df579bb0a82 · outbound

This paper cites Clip- situ: Effectively leveraging clip for conditional predictions in situation recognition.

Dynamic Scene Understanding from Vision-Language Representations Clip- situ: Effectively leveraging clip for conditional predictions in situation recognition

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.580244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.184781Z digest=sha256:ef3c69b6d493ea4cadeef3e79fc38574080155744b08dd813e0a7be7e3f7bd6e

Observation 127bb760-d0e2-4bf0-b721-374d14b86ab8 · outbound

This paper cites Ut-interaction dataset, icpr contest on semantic description of human activities (sdha).

Dynamic Scene Understanding from Vision-Language Representations Ut-interaction dataset, icpr contest on semantic description of human activities (sdha)

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.571391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.188041Z digest=sha256:b34fab69315a39bc2c987b5273409654a840d497c1876f3cdb134eaa016ef179

Observation 97dbb885-b995-46bc-990c-daac2e0537cd · outbound

This paper cites BLEURT: Learning Robust Metrics for Text Generation.

Dynamic Scene Understanding from Vision-Language Representations BLEURT: Learning Robust Metrics for Text Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.191112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.191112Z digest=sha256:be0612c716a07330889b7ffd0fa92ab956a6490e6f39a416d9b92696ade915bf

Observation 3816ec7c-6c74-4cb7-b9cb-efa396178278 · outbound

This paper cites Analyzing human– human interactions: A survey.

Dynamic Scene Understanding from Vision-Language Representations Analyzing human– human interactions: A survey

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.561295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.194562Z digest=sha256:98a97736a37d0e18d7e9bdeab8de7de0ba31a95784b5195d7cf693ce80032a1a

Observation 10e15eda-365b-4947-b854-62f591b869b4 · outbound

This paper cites Generative multimodal mod- els are in-context learners.

Dynamic Scene Understanding from Vision-Language Representations Generative multimodal mod- els are in-context learners

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.197813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.197813Z digest=sha256:fcf83f9627ab6416388ccda534fa9f8ee55e598163dc9e4ec3b7ba8909fe9861

Observation 9f40e930-6501-479f-aff6-e125cda380c4 · outbound

This paper cites Qpic: Query-based pairwise human-object interaction detec- tion with image-wide contextual information.

Dynamic Scene Understanding from Vision-Language Representations Qpic: Query-based pairwise human-object interaction detec- tion with image-wide contextual information

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.200832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.200832Z digest=sha256:075542fcc84276234aaa640827d05babb4bdf3a09b4291a75d81748ba5fe5d7e

Observation 87e9cac3-da5f-49a0-9d3a-05c94c91d1fc · outbound

This paper cites Learning latent temporal structure for complex event detection.

Dynamic Scene Understanding from Vision-Language Representations Learning latent temporal structure for complex event detection

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.538545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.203325Z digest=sha256:1b7a778603d70745d43d05cae0e84ba734ce44085c41f25d14b6398b3e91ced4

Observation 86b04a56-3d0e-4dee-bdb7-443a672c3542 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Dynamic Scene Understanding from Vision-Language Representations Learning spatiotemporal features with 3d convolutional networks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.528475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.205878Z digest=sha256:94e5ac655ff832b328382c66ffba2a07f781f60856236ff0adf716a3ddf6115b

Observation 1d9f3e91-2a1a-4368-9bce-da217fdc3560 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Dynamic Scene Understanding from Vision-Language Representations ActionCLIP: A New Paradigm for Video Action Recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.208421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.208421Z digest=sha256:890070edf83a2054e6b5dd3f83b12291420f7efa6f306116a047271b3df799c6

Observation d29cfbf7-2dcc-4779-aaab-4305283204c6 · outbound

This paper cites Cris: Clip- driven referring image segmentation.

Dynamic Scene Understanding from Vision-Language Representations Cris: Clip- driven referring image segmentation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.518750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.211084Z digest=sha256:f4ad20a9fc09018935708d59696ad7d0bb791af749ba91d367fabdd3ff3be517

Observation 8357c391-7ff1-48fc-8cb8-d09e3603ec84 · outbound

This paper cites Rethinking the two-stage framework for grounded sit- uation recognition.

Dynamic Scene Understanding from Vision-Language Representations Rethinking the two-stage framework for grounded sit- uation recognition

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.509425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.213705Z digest=sha256:f73b479fdd888c376d0e387ed68f037660b5f9d48eb45fb6ff4d2ab14f94328f

Observation 4d4757c7-835d-4760-bafa-ec89c30a47a7 · outbound

This paper cites Rec- ognize complex events from static images by fusing deep channels.

Dynamic Scene Understanding from Vision-Language Representations Rec- ognize complex events from static images by fusing deep channels

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.499989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.216364Z digest=sha256:a740e0c5f61138e8d728370602c2a7f872eba0c5aae7b8c3a70eba557cc25968

Observation 40e3df64-9735-4a28-aec7-cafe846af708 · outbound

This paper cites Recognizing proxemics in personal photos.

Dynamic Scene Understanding from Vision-Language Representations Recognizing proxemics in personal photos

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.490090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.219547Z digest=sha256:3d79f2222222b8dbc705c7616a67b0772068a469c8af65662d2c7421b483e563

Observation d9a2302c-2e38-4d47-8166-9ab607be7c5d · outbound

This paper cites Situa- tion recognition: Visual semantic role labeling for image understanding.

Dynamic Scene Understanding from Vision-Language Representations Situa- tion recognition: Visual semantic role labeling for image understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.481079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.222066Z digest=sha256:5eda86ef5d55f118170b0b497fa0a4050c23efe589038fe8fd869924ff9db348

Observation 1fc5b6a0-4560-4c77-94cf-9e8393e3030d · outbound

This paper cites Rlip: Rela- tional language-image pre-training for human-object inter- action detection.

Dynamic Scene Understanding from Vision-Language Representations Rlip: Rela- tional language-image pre-training for human-object inter- action detection

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.463025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.228117Z digest=sha256:6062fc45466a39c3f32473186c3f9ec4c509adb61108f33acceada1c017804ae

Observation a632b5bf-5002-4f11-9363-c366041d74ac · outbound

This paper cites Rlipv2: Fast scaling of re- lational language-image pre-training.

Dynamic Scene Understanding from Vision-Language Representations Rlipv2: Fast scaling of re- lational language-image pre-training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.453244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.230741Z digest=sha256:99a7030d00d1006300d3ae9ce5d39d7b8653a1a981eda590192aad84b7505f5e

Observation a3ee7e89-a783-4f6c-9663-796d404fb7b0 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

Dynamic Scene Understanding from Vision-Language Representations Lit: Zero-shot transfer with locked-image text tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.443341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.233316Z digest=sha256:6f3e55da87711f9bd33a729247c1d37dc1d650219608b9a11f2d6bf3dc0108e5

Observation a96dc666-14e7-48e0-b7b5-af39e5851057 · outbound

This paper cites Zhang, Dylan Campbell, and Stephen Gould.

Dynamic Scene Understanding from Vision-Language Representations Zhang, Dylan Campbell, and Stephen Gould

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.432872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.236106Z digest=sha256:12d8c27810bc55250ca6078e71eefac029a34b18e1e3047ca3c6b2b1472a6d4f

Observation 4fd9d762-a0b8-4a76-a702-5de0f1674ac7 · outbound

This paper cites Zhang, Yuhui Yuan, Dylan Campbell, Zhuoyao Zhong, and Stephen Gould.

Dynamic Scene Understanding from Vision-Language Representations Zhang, Yuhui Yuan, Dylan Campbell, Zhuoyao Zhong, and Stephen Gould

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.423024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.238616Z digest=sha256:0d6b27ed1799e64ded08d0ae10273f8d40e8e7de036e6a6d19b74347eae5eff5

Observation 903ec1a6-83d9-435b-ae74-c067191ce6b8 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Dynamic Scene Understanding from Vision-Language Representations OPT: Open Pre-trained Transformer Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.241423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.241423Z digest=sha256:7bcadd099becf48910829db398c109011e805ce927b959858fc0483ada70bad2

Observation 92e5280f-c78a-4577-9639-a71346507a45 · outbound

This paper cites Zegclip: Towards adapting clip for zero-shot se- mantic segmentation.

Dynamic Scene Understanding from Vision-Language Representations Zegclip: Towards adapting clip for zero-shot se- mantic segmentation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:04:29.412879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.244936Z digest=sha256:cb36793082a3340d1543f35a828aae8845f5941ee904b3132eaed8f23730c9c6

Observation 93ec1946-844a-4ce5-b39c-6318c576862c · outbound

This paper cites Scene Graph Generation: A Comprehensive Survey.

Dynamic Scene Understanding from Vision-Language Representations Scene Graph Generation: A Comprehensive Survey

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.248093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.248093Z digest=sha256:241a91445ea25390de31dc3dbe1f006f6ff2ceadf3660f30897bbf19e3c34c2b

Observation 1dc5a034-c241-4ee0-b834-9ed607815125 · outbound

This paper cites A Comprehensive Study of Deep Video Action Recognition.

Dynamic Scene Understanding from Vision-Language Representations A Comprehensive Study of Deep Video Action Recognition

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:29.251417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:29.251417Z digest=sha256:c4fd68eb267bd080d0c42acb322839ede0aeca32d02eb8b8e246084ea0b40a20

Observation 6596321f-a215-413d-98bc-57a47de3d03b · outbound

This paper cites an unresolved cited work.

Dynamic Scene Understanding from Vision-Language Representations Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:04:29.401574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.255261Z digest=sha256:4846e5a8434f0a73939e751632a79fb20cab42784123044796c5bc1bcd20d0be

Observation 1356fec0-a3f6-4e6c-96b2-c9188483b343 · outbound

This paper cites an unresolved cited work.

Dynamic Scene Understanding from Vision-Language Representations Unresolved cited work

Reference 2016

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:04:29.472198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T18:04:29.224974Z digest=sha256:78fa19f6bafe0a108daca6aa0d1103ab7a66c5d5c5175d35fbe3ad1ccc1ba9ff

Pith citing papers

No inbound Pith citation observations are available.