Pith. sign in

Paper Citation Record · LEDGER

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis

As of 13 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2411.18038.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18038 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:38:34.786612Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2a5237ee-346a-4d9e-8d89-29e96b39bfe1 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.146453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.146453Z digest=sha256:62aa89e0c23ee14c2ef1dbb39f9c39a81b19da5a9ec15379a0d5bfc03003de47

Observation d0e264e1-1968-45a2-a10c-1191df854adc · outbound

This paper cites In: Proceedings of the IEEE international conference on computer vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE international conference on computer vision

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.844543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.150887Z digest=sha256:da673aded6a19e596c84501517a86f0a7579a15d5503799c67bc6bcf29a83d69

Observation 49eafe1f-9012-439b-bebf-cce652cfce60 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence 41(2), 423–443 (2019).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis IEEE Transactions on Pattern Analysis and Machine Intelligence 41(2), 423–443 (2019)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.837010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.153944Z digest=sha256:659e72dbeb0b8719dd3b900f3c5bdb5dedb2e1237f0a6b1f5cc03b1a87f9958a

Observation 88b30466-0e00-46ad-ba7f-467c18b95f20 · outbound

This paper cites Advances in neural information processing systems33, 1877–1901 (2020).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Advances in neural information processing systems33, 1877–1901 (2020)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.182657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.182657Z digest=sha256:d01fcb0c63640e303f894c97524c60b21d396247f747deacc5af3828895208d5

Observation 85040c99-cefb-4f47-835b-4fd33acf2181 · outbound

This paper cites In: European conference on computer vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: European conference on computer vision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.229865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.229865Z digest=sha256:b0a6424508b101b14e1d1a6e09d681acdd1aae8e827e53c5b435e1e5d78cacd7

Observation c5d792de-6252-4937-b5dc-71d1c9c48ea9 · outbound

This paper cites an unresolved cited work.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:38:35.821530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.257749Z digest=sha256:2da20e64dcd72d9f767f96a44c3863946ed5767ece6ab68ca4c193a960327fa7

Observation 67d57506-ec78-4fc6-af1e-002d5bec64c5 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.307670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.307670Z digest=sha256:e598cf99626cea514de1f6a537cb2a94faa5fcdfc277faab53deee939190852d

Observation 5c0fa5d8-16f1-45ed-b17e-e9edc0050153 · outbound

This paper cites In: European conference on computer vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: European conference on computer vision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.310959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.310959Z digest=sha256:12eb3d10e82eaedbed21b60dcbd9ec22db5622037ec1aca4cfffdb7075c1d559

Observation 38c3e4b8-3af2-4bdd-83da-fa6738dc59cb · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis PaLM: Scaling Language Modeling with Pathways

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.313458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.313458Z digest=sha256:2d260e94edf6d9123e4ef22f9e4ab590c2396e335d55ae5ece2338f426a359b2

Observation 20608490-ee45-4f30-acd1-bfd3edbdf1d1 · outbound

This paper cites an unresolved cited work.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.316285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.316285Z digest=sha256:6fbca6dd1e3097df3ad1aa1d276906243aae529efa1a61d9a10e7852b2df538b

Observation baaf330d-1904-45b8-9a6b-72f93c1e3007 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.319094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.319094Z digest=sha256:3fb7a25f70e87fd7f847a92dd454b492c93c6627c15b0af9e53113c161edd207

Observation ce13505b-3646-405b-87d1-3f2fbce601ad · outbound

This paper cites In: British Machine Vision Conference (2018) 16 Kang et al.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: British Machine Vision Conference (2018) 16 Kang et al

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.702688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.322040Z digest=sha256:3072ea20490cb57435c92e16de7b721511b1a6ee4fe840395d339b1604545ea6

Observation 16a34395-73be-4311-b0b8-76a56171dbc7 · outbound

This paper cites Visual Semantic Role Labeling.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Visual Semantic Role Labeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.324732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.324732Z digest=sha256:77d901b82bc5884821e25ad095473d139fec5a4e668d71f3b854d62b4d6f3a91

Observation cfcfafc0-6767-4973-b30c-0997d789c924 · outbound

This paper cites In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XV 16.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XV 16

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.670782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.327667Z digest=sha256:fcfc4810229fec5c17b5aff0839d3c41e75b4a01e4044ebaaa80f149e7c467a4

Observation 3afe03d3-9352-462f-b2e0-99bf7a364685 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.662741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.330845Z digest=sha256:c9d2c4c34a53d62cf19ef21f80853648c985f02a5467456189b998d8a1dfc45a

Observation 49b950cf-6425-4459-8b2e-8763d375c1f4 · outbound

This paper cites In: Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.653126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.333322Z digest=sha256:e7c71425c3531c6c8ef14882ae7db7ccbf2abb6db92b28dcd1a57fd25672ff73

Observation 6cb0126b-990e-4865-a4d0-51c92ddd8914 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.643674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.336324Z digest=sha256:cf360bd6c45f7f57c855eadd18970ffa5e209dd8b70c155b68f2261553a54613

Observation 468cc232-0e42-412d-a22a-152e18e28359 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.634586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.338694Z digest=sha256:86c4a771f6a43ee1399f1d7f76394f1a3e7aa1ac5f2173e4349e9aef542bf8a0

Observation 8c6ff187-c0d1-4a4d-b599-4fd492dc76d0 · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.341241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.341241Z digest=sha256:7a6de3cab18b30f5b154093962b53080be6158d4dff63aaff66b891c6ea30f8c

Observation 9b13d84c-4e55-4138-89f8-c3b2aec89ec1 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.344173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.344173Z digest=sha256:3cf25020a731a84d783ed5663547bb581c83d1b002f81de3607dfa8dadb0ab80

Observation dab8e512-0800-4166-b281-c06f3201a948 · outbound

This paper cites In: International Con- ference on Machine Learning.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: International Con- ference on Machine Learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.588912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.385069Z digest=sha256:dcae6615584751187a2db3407d02018e9f2549fe23a91fec6a8e1654acadd44e

Observation b6a4ca7c-a1a0-47d8-bf5b-48f1acae6504 · outbound

This paper cites Advances in neural information processing systems34, 9694–9705 (2021).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Advances in neural information processing systems34, 9694–9705 (2021)

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.440620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.440620Z digest=sha256:9ad90c10110c712179c5b82e917dc8750a97e874233cc5a4442f2b5bb3dbe840

Observation 7907736c-465b-4605-9b58-b68df20fc304 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.443119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.443119Z digest=sha256:4478bf6fa26af4c3f67185eb92d087782a4ba50a71d9e521b6f0d7cd36df7fef

Observation 68157de3-a025-4aea-b66f-c00cd57e16b2 · outbound

This paper cites Advances in Neural Information Processing Systems33, 5011–5022 (2020).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Advances in Neural Information Processing Systems33, 5011–5022 (2020)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.502440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.446028Z digest=sha256:cfc845ea0cc265aa4a540765c1fd775ec93a3af88904721a68206c8e61e4cd5a

Observation 1d8b1ad2-6a37-4d19-bcee-c45414a361ea · outbound

This paper cites In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.365178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.449299Z digest=sha256:d6d842a181b18f42f5ede5dd06a68638cb13c2ec6b9c4ec735bb1a3606766ec2

Observation 4072d5d1-da5e-4bbb-9427-779fcd30977e · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.356245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.451737Z digest=sha256:46f409c6480345201b6afa9f44ae970b1634580e02936c15123d84808c15d46a

Observation b966a311-15c0-479a-8d53-fcdc31acef21 · outbound

This paper cites In: Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.454100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.454100Z digest=sha256:5832a2107f6312185fe119ec083b2c39ec8039cdcb6344fdddae07150e4b68fe

Observation ad934a85-b42e-47ef-9f5f-3b27c417013b · outbound

This paper cites Visual Instruction Tuning.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Visual Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.457601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.457601Z digest=sha256:61bac0e098fbf4989c4ff0c21232ef3b3c089d3220354da7243d3972f95f56ef

Observation 39f77caa-80f2-428e-94a3-32c7a8f5a8a4 · outbound

This paper cites Decoupled Weight Decay Regularization.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.460943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.460943Z digest=sha256:8a74630424878237c1bdc7f74586a36b5d92aabae7f642fbfe49dc3e46dddf09

Observation b4590073-7e33-49eb-91f9-c9faf21b51ec · outbound

This paper cites In: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.266247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.463850Z digest=sha256:9cf3af4fd5eb5734ccefa790393f37f1472f375e93466bedb3b4215a43dfe341

Observation b3e67d8e-369d-4462-978f-93c180f89bb8 · outbound

This paper cites FGAHOI: Fine-Grained Anchors for Human-Object Interaction Detection.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis FGAHOI: Fine-Grained Anchors for Human-Object Interaction Detection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.466506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.466506Z digest=sha256:bb29b8b193930cdae83ef12104f7ab8deec30366d62602272969b1e164e1a5e0

Observation e0fbbeee-727a-4b37-8f9e-9fde518ce149 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.188249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.469599Z digest=sha256:e8b9b4bab39398a3d3bd2a4e7b1cbdb46f89355e607810b2636f9fa277b1fb16

Observation 2e4a9699-9d60-4685-9290-cc2f3d8290e9 · outbound

This paper cites an unresolved cited work.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.481853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.481853Z digest=sha256:198692d97b290f6202d1640e3fd1bdd05e3291c3c3fda48ab69d08cb1a329326

Observation fb43d5f5-9c30-4e0d-a2dd-3b225be02050 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.175904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.545966Z digest=sha256:7d75ff62b963c0601bca15f09e50a3b1cde90a14ee8dae7184cdaa84d2cab26c

Observation 0f8ebe15-af04-4b86-8431-b3b2de892750 · outbound

This paper cites In: International conference on machine learning.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: International conference on machine learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.572935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.572935Z digest=sha256:5be2737e041db8ba70285ddaa82d601324cc9dd028b6a47566026b3776718221

Observation b4e12a1d-933c-4d96-959e-d9ade6969be9 · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.575819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.575819Z digest=sha256:ef6a34b5a3ad6b3ad765810f522660d81bfc4e952d540a7a80eba83349375176

Observation fe88a09e-a0fb-4683-b71e-33abee063a87 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.578827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.578827Z digest=sha256:8013a6a35efd1bdfa6790b9dddac56637860a1b70c0a920cb9d38f18579c0ceb

Observation b107420d-4051-46e9-80ec-a6028a59b865 · outbound

This paper cites LXMERT: Learning Cross-Modality Encoder Representations from Transformers.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.581837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.581837Z digest=sha256:9db86dc93ba04c1e7e05082e8620cdbb5f172062a9359a3e256f8c36d7acf285

Observation d35a03d6-eefc-4750-8d05-0c575b4f22c2 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.585350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.585350Z digest=sha256:c0ae2bf3d1b67d4e4fab6d636ae2412cb48480bbe1197180d8e586aaacd3db9e

Observation 92a2bf5e-367c-4af0-a9cf-aef629b6e4e3 · outbound

This paper cites In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.155168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.587955Z digest=sha256:1e2a26cc03e21d344ae2d62dc6f7443e26db2a1e3dba6277bcc949b7ddecaef6

Observation 0c545409-b1fa-429c-aec1-bd44cbf335ce · outbound

This paper cites In: Proceedings of the European conference on computer vision (ECCV).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the European conference on computer vision (ECCV)

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.590551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.590551Z digest=sha256:e688220be6a9557b0151cea58a70edcba46e508ad431f53510d0ba6ac3929a3c

Observation 9f4167a1-bedd-4c95-ae2b-021c5018a1d8 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.142769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.592927Z digest=sha256:db757f47227f1d847c5d51ab41c71f0716c20ef9d9258edb67113e74296755c2

Observation 0aa0769e-cb68-4b33-86d1-a54c3a8f39e8 · outbound

This paper cites In: International conference on machine learning.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: International conference on machine learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.135268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.595219Z digest=sha256:c5be93a5a66556df58f762097ac5267ed82cfb5fef9415b10339ec5d8ada8d8c

Observation 4aa3d3be-222d-4f32-ad08-a1709fca3df6 · outbound

This paper cites In: Proceedings of the AAAI conference on ar- tificial intelligence.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the AAAI conference on ar- tificial intelligence

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.597799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.597799Z digest=sha256:3d2fba991a1423d06dfcf8c10cecc477bfc736782e0e023a101370b277d73ed2

Observation e68e1ace-973d-47f8-98b4-5382e19484b6 · outbound

This paper cites In: Proceedings of the European conference on computer vision (ECCV).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the European conference on computer vision (ECCV)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.063951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.600717Z digest=sha256:1225c83ccb63827739afaacb8b063b30a4c8fd3e1324fb285064969cfd79db75

Observation 2259e9ad-ec0a-4263-a383-d3de78fb876f · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.603288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.603288Z digest=sha256:d950aee7c4398630054de3c9987b4ef853f80712e3068ba911f0abcffcf0118f

Observation 51144232-2b26-4a49-814f-47954a3cf100 · outbound

This paper cites Ad- vances in Neural Information Processing Systems35, 37416–37431 (2022).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Ad- vances in Neural Information Processing Systems35, 37416–37431 (2022)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.056817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.605972Z digest=sha256:5f4d40e33cb10c14f366cdb7ad65bcaf962963d172c865591e3ae04c46654f41

Observation e59024cd-59db-416d-9f8d-5aaee0c493b5 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.049340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.609372Z digest=sha256:a52cb2ec6f5f29a8e6838503e4b1fc0943f4fc1b3bdba757963ae79a57f3cdff

Observation 7daa7899-863e-45fb-9a7c-dfced26d0ea8 · outbound

This paper cites Advances in Neural Information Processing Systems 34, 17209–17220 (2021).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Advances in Neural Information Processing Systems 34, 17209–17220 (2021)

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.041544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.611855Z digest=sha256:4dd54fa96eb0ea1d8e79880e0e70a801679bf15dec568e68d22ccf22e740c3aa

Observation e2cc35e1-7672-4cb1-96a8-009b93664ed7 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Con- ference on Computer Vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF International Con- ference on Computer Vision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.033708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.614276Z digest=sha256:205539184f2906f9b89aa2ca374b4380e279ea4d88c63fd5d7f645c9708114eb

Observation e3cb7b2c-7c06-4f1f-b2c5-99a5d09d6a06 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.024223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.617496Z digest=sha256:0ed4f256497966dc28098a5c7326d4d7103b590ef8b1cc34826caf8b15fd09c0

Observation 36ab7426-2f8c-4757-82b5-f92ab22e4fe5 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.014544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.619939Z digest=sha256:8483773814cb9bb34b616e07c8b6e8bcc9efa3f2a970c25fbc07c56c703b4434

Observation 1e058d5a-d39d-46f1-af68-ebd671d072ec · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:34.967395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.647682Z digest=sha256:05b1ba59082d4bb147f5fd83227be3897173473c23443583529499615b2b3215

Observation 7295c9bf-9a7e-4229-9ae1-8e8eb082b4e2 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:34.917075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.669349Z digest=sha256:29560ed890fc2fc461634fb8eaed8fc832afb0c5479f462f0a3636fdc7306a4b

Observation 70d86c36-21bb-43a2-828c-5ea64e462e26 · outbound

This paper cites an unresolved cited work.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:38:34.902737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.723346Z digest=sha256:f47e0a5af00ab23c509e873c4a0f51f23f6a8bd93db34b1df9c57613c923b149

Observation 919cee17-65ca-4fa3-b275-6fa30cb2a7c2 · outbound

This paper cites In: Proceedings of the AAAI conference on artificial intelligence.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the AAAI conference on artificial intelligence

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:34.895432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.746698Z digest=sha256:8552fa3e0ac68c07df99826fe0af84b2719fb8570168f6398b57e3077b5adced

Observation 652a7513-062d-4bf7-8d56-9714bd8b9d39 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:34.887328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.786612Z digest=sha256:31838eda6dd9b9bc2c0ed30ce3d52d056ae06d4182ed9d8733207c6c7ebcc359

Pith citing papers

No inbound Pith citation observations are available.