Pith. sign in

Paper Citation Record · LEDGER

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis

As of 13 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2411.18038.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18038 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:38:34.786612Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2a5237ee-346a-4d9e-8d89-29e96b39bfe1 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.146453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.146453Z digest=sha256:f55e8b79d7ac36e98dd0252b8e7129120f3a512903642f9728cb9850b5b7b4b7

Observation d0e264e1-1968-45a2-a10c-1191df854adc · outbound

This paper cites In: Proceedings of the IEEE international conference on computer vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE international conference on computer vision

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.844543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.150887Z digest=sha256:c547d625f66f2aaea99aaea485fb65e333039748255d5c64ea963d7091b6951a

Observation 49eafe1f-9012-439b-bebf-cce652cfce60 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence 41(2), 423–443 (2019).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis IEEE Transactions on Pattern Analysis and Machine Intelligence 41(2), 423–443 (2019)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.837010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.153944Z digest=sha256:38f1a88beed9328f25f25de3ea6466d4a5d929d1b2cec6a2b01c2d64e2069b84

Observation 88b30466-0e00-46ad-ba7f-467c18b95f20 · outbound

This paper cites Advances in neural information processing systems33, 1877–1901 (2020).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Advances in neural information processing systems33, 1877–1901 (2020)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.182657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.182657Z digest=sha256:ade35ceb08444cce95d238df6fcbb2d480ed08ae85403d4b020fdfbe84198b19

Observation 85040c99-cefb-4f47-835b-4fd33acf2181 · outbound

This paper cites In: European conference on computer vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: European conference on computer vision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.229865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.229865Z digest=sha256:2c73265660dd33b02f203fe07e946cd79241ee9c3f839dafbd8deb2ec86205ea

Observation c5d792de-6252-4937-b5dc-71d1c9c48ea9 · outbound

This paper cites an unresolved cited work.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:38:35.821530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.257749Z digest=sha256:2f98fa3ced46589e254ac463288c9b590d94390fb6a757896ff4cf8fcb649b17

Observation 67d57506-ec78-4fc6-af1e-002d5bec64c5 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.307670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.307670Z digest=sha256:96eb0e591e03cf42a4b24d9ef3041d383586d927d93c45d4e3b1fb5c3ac551a5

Observation 5c0fa5d8-16f1-45ed-b17e-e9edc0050153 · outbound

This paper cites In: European conference on computer vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: European conference on computer vision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.310959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.310959Z digest=sha256:4592a756e77cfd139ecf2f788ac7249560c33326d510a446acea380912bc7b38

Observation 38c3e4b8-3af2-4bdd-83da-fa6738dc59cb · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis PaLM: Scaling Language Modeling with Pathways

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.313458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.313458Z digest=sha256:efdb3048eb37321555183b6492bc1720677b0d4b6f641a3dd7561b0493fea1d3

Observation 20608490-ee45-4f30-acd1-bfd3edbdf1d1 · outbound

This paper cites an unresolved cited work.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.316285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.316285Z digest=sha256:ade7373e4ca142b51b34631eb7bcac50491c9030a09c10936246d4f2eed05ead

Observation baaf330d-1904-45b8-9a6b-72f93c1e3007 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.319094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.319094Z digest=sha256:da3e708e95f0912f49a0bf58d19ddd9449699092c8fdac234aefa025f4dce4c3

Observation ce13505b-3646-405b-87d1-3f2fbce601ad · outbound

This paper cites In: British Machine Vision Conference (2018) 16 Kang et al.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: British Machine Vision Conference (2018) 16 Kang et al

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.702688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.322040Z digest=sha256:b5c11753a4925831683a56c2828db1bfd1789d68774a4217c78b369a05be86b8

Observation 16a34395-73be-4311-b0b8-76a56171dbc7 · outbound

This paper cites Visual Semantic Role Labeling.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Visual Semantic Role Labeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.324732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.324732Z digest=sha256:36fba98f597c1b6b7d6ee3455b7dca4af79ca762ffb9fd58c7e02581df88f0ee

Observation cfcfafc0-6767-4973-b30c-0997d789c924 · outbound

This paper cites In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XV 16.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XV 16

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.670782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.327667Z digest=sha256:376049811aa80985b0e2e734914eaa3f09c0a5c2a521d7134ebdbe937d2412c1

Observation 3afe03d3-9352-462f-b2e0-99bf7a364685 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.662741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.330845Z digest=sha256:209958d23649c13006cd8ccb85cfb1db6952e583d76c4fe82b504e33f9b4414b

Observation 49b950cf-6425-4459-8b2e-8763d375c1f4 · outbound

This paper cites In: Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.653126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.333322Z digest=sha256:be09b0d4274e9365b6f694f009e65dbfe526c5d818c31b6955464ac80c434a01

Observation 6cb0126b-990e-4865-a4d0-51c92ddd8914 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.643674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.336324Z digest=sha256:49e7319b9e5c6b8a7574d651ccec11ec0d4eb71f27a4338011b59137606ded74

Observation 468cc232-0e42-412d-a22a-152e18e28359 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.634586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.338694Z digest=sha256:1ab59d3c2b38be8d1f892897131942221a552a05e335679dc54ee37c076dab53

Observation 8c6ff187-c0d1-4a4d-b599-4fd492dc76d0 · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.341241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.341241Z digest=sha256:5df2f149ec07083336bd2cbf03a8fd10defe0214fcf0f3641264f8a2c95ec879

Observation 9b13d84c-4e55-4138-89f8-c3b2aec89ec1 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.344173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.344173Z digest=sha256:2008ee6f39f46da6b767b0dd33ea4943cc248049166ab45693323d89f95151f8

Observation dab8e512-0800-4166-b281-c06f3201a948 · outbound

This paper cites In: International Con- ference on Machine Learning.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: International Con- ference on Machine Learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.588912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.385069Z digest=sha256:013c4d6a089318b67d520c08b7fb575f547e278fad2b8e529a28af9f083a2b26

Observation b6a4ca7c-a1a0-47d8-bf5b-48f1acae6504 · outbound

This paper cites Advances in neural information processing systems34, 9694–9705 (2021).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Advances in neural information processing systems34, 9694–9705 (2021)

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.440620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.440620Z digest=sha256:777628648a71dcb09874d3cf71ef1162f91eced782df254ead3f8bc6e8253a5e

Observation 7907736c-465b-4605-9b58-b68df20fc304 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.443119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.443119Z digest=sha256:44b0655516240e6446efd993e7d9233df410f17c39216f50f309b5b4b0c23d25

Observation 68157de3-a025-4aea-b66f-c00cd57e16b2 · outbound

This paper cites Advances in Neural Information Processing Systems33, 5011–5022 (2020).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Advances in Neural Information Processing Systems33, 5011–5022 (2020)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.502440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.446028Z digest=sha256:44957713f34732387f23fa2adea98132e24a82a02fdc42109e980f99378792cf

Observation 1d8b1ad2-6a37-4d19-bcee-c45414a361ea · outbound

This paper cites In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.365178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.449299Z digest=sha256:f0aeed77c79e45e1285cd3bd48ff3c9baceb5c7d4d1592fbc01a2f19345f8ca6

Observation 4072d5d1-da5e-4bbb-9427-779fcd30977e · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.356245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.451737Z digest=sha256:eb62eeaa84bf5b29123bc97969bf83306c4907fc4f886bcfa3a555e3b8439df4

Observation b966a311-15c0-479a-8d53-fcdc31acef21 · outbound

This paper cites In: Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.454100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.454100Z digest=sha256:9d842d7db111afccdbbe7b2af65d670195468944540d159378713d452ca181ad

Observation ad934a85-b42e-47ef-9f5f-3b27c417013b · outbound

This paper cites Visual Instruction Tuning.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Visual Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.457601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.457601Z digest=sha256:d3ae3c13a51d270c2b7567809657c1f851bc570408c3c568650a55661ca06944

Observation 39f77caa-80f2-428e-94a3-32c7a8f5a8a4 · outbound

This paper cites Decoupled Weight Decay Regularization.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.460943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.460943Z digest=sha256:5b4b4da0ab018049bd55a16de10070eb91556a3c0fa4dcb651783c7251c84e87

Observation b4590073-7e33-49eb-91f9-c9faf21b51ec · outbound

This paper cites In: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.266247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.463850Z digest=sha256:649cdae2105ea3062bb3363d5cd1c39949ff85f864dd12c85521fc715a9a68bb

Observation b3e67d8e-369d-4462-978f-93c180f89bb8 · outbound

This paper cites FGAHOI: Fine-Grained Anchors for Human-Object Interaction Detection.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis FGAHOI: Fine-Grained Anchors for Human-Object Interaction Detection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.466506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.466506Z digest=sha256:077d4ec7e26a35705ae3a3592614b7f164244f61f34b58e050cce7da2badb75e

Observation e0fbbeee-727a-4b37-8f9e-9fde518ce149 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.188249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.469599Z digest=sha256:769cec8a826148760e242cf7aa3ea96bef3077ef8d466e15f5720b214c0897a4

Observation 2e4a9699-9d60-4685-9290-cc2f3d8290e9 · outbound

This paper cites an unresolved cited work.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.481853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.481853Z digest=sha256:fecad463584ef4379ad5cb0b5be581626b13f1f0f078ec9c31d2b7230a4b8e19

Observation fb43d5f5-9c30-4e0d-a2dd-3b225be02050 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.175904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.545966Z digest=sha256:8a1492eb02fc43174da44c17f801cb17469af95f3b1d8cb44e9d8c9e2de2614e

Observation 0f8ebe15-af04-4b86-8431-b3b2de892750 · outbound

This paper cites In: International conference on machine learning.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: International conference on machine learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.572935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.572935Z digest=sha256:46b3851c50e7451529fe8a86352aa23e13c0251d45d44f7dd13749d90af0a4e3

Observation b4e12a1d-933c-4d96-959e-d9ade6969be9 · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.575819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.575819Z digest=sha256:949926ac3aa5edd37acb14eb79ba23dec222029faec93c9180b36d234f692f71

Observation fe88a09e-a0fb-4683-b71e-33abee063a87 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.578827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.578827Z digest=sha256:3e360405e72d2594e2598e9c53885a3cf7ce1b89d6773009d46bf5d861f8825f

Observation b107420d-4051-46e9-80ec-a6028a59b865 · outbound

This paper cites LXMERT: Learning Cross-Modality Encoder Representations from Transformers.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.581837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.581837Z digest=sha256:b33b8d59297c82501702ec49768bbc8532b5e3a069be21c3f16d10901f0485df

Observation d35a03d6-eefc-4750-8d05-0c575b4f22c2 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.585350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.585350Z digest=sha256:291574b9a595e684f7a0ae70bc7d0c32f312d9a0f0e69d557fc7e1ed9bb015ab

Observation 92a2bf5e-367c-4af0-a9cf-aef629b6e4e3 · outbound

This paper cites In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.155168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.587955Z digest=sha256:dafa102ff7579345decb0f984f267c02c84aead0caf4035a1c06c7aa109de88f

Observation 0c545409-b1fa-429c-aec1-bd44cbf335ce · outbound

This paper cites In: Proceedings of the European conference on computer vision (ECCV).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the European conference on computer vision (ECCV)

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.590551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.590551Z digest=sha256:d78f2c1d73fa120e16c32cca675aff811e2aa600229d1398f9f53df683ad8d9f

Observation 9f4167a1-bedd-4c95-ae2b-021c5018a1d8 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.142769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.592927Z digest=sha256:d06b8a633308093142766864ec363bd29e57091afe1a2eab4a723d9fe678d7ab

Observation 0aa0769e-cb68-4b33-86d1-a54c3a8f39e8 · outbound

This paper cites In: International conference on machine learning.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: International conference on machine learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.135268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.595219Z digest=sha256:cf79ec1af2c868d4c1cc3f180abaa534ef4b4e6d965b90d6f4d18d75f6244d3c

Observation 4aa3d3be-222d-4f32-ad08-a1709fca3df6 · outbound

This paper cites In: Proceedings of the AAAI conference on ar- tificial intelligence.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the AAAI conference on ar- tificial intelligence

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.597799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.597799Z digest=sha256:ce0da07bb31c9777ea2897c11f513181a91a4c746d3c824eb21c988276fb361a

Observation e68e1ace-973d-47f8-98b4-5382e19484b6 · outbound

This paper cites In: Proceedings of the European conference on computer vision (ECCV).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the European conference on computer vision (ECCV)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.063951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.600717Z digest=sha256:ebf5c337a92e3172c010910348fef5c9177ccc70f3c1167ff2acc7efcf9c7d9e

Observation 2259e9ad-ec0a-4263-a383-d3de78fb876f · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.603288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.603288Z digest=sha256:1813e43bf1c07276cf6058f777f8227bdd47edd3502e96ac5f03e44b26ae55f2

Observation 51144232-2b26-4a49-814f-47954a3cf100 · outbound

This paper cites Ad- vances in Neural Information Processing Systems35, 37416–37431 (2022).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Ad- vances in Neural Information Processing Systems35, 37416–37431 (2022)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.056817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.605972Z digest=sha256:2a33b7c8796c26303cb2248dc98167b3e5835b6770fa54becd3aa02eab9141b9

Observation e59024cd-59db-416d-9f8d-5aaee0c493b5 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.049340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.609372Z digest=sha256:d11f5d7c7a02f4cf99a39d8a72a560136ffb94fbb8937ecdaaadc9375ee9484e

Observation 7daa7899-863e-45fb-9a7c-dfced26d0ea8 · outbound

This paper cites Advances in Neural Information Processing Systems 34, 17209–17220 (2021).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Advances in Neural Information Processing Systems 34, 17209–17220 (2021)

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.041544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.611855Z digest=sha256:28ba2296d949590c933daf2a35eaf55b12d04663b55131c9209deacf529e75c2

Observation e2cc35e1-7672-4cb1-96a8-009b93664ed7 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Con- ference on Computer Vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF International Con- ference on Computer Vision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.033708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.614276Z digest=sha256:870ddfa3b356bec9771fd759fa19cbf85301a956911f65102132232f3b989d0a

Observation e3cb7b2c-7c06-4f1f-b2c5-99a5d09d6a06 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.024223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.617496Z digest=sha256:93fc7faaf6850a01a9f022abf2208c65f373bfb44dfa1b9d1ba0cf06b0ecce11

Observation 36ab7426-2f8c-4757-82b5-f92ab22e4fe5 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.014544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.619939Z digest=sha256:db2a2264f758daa593672a1f85f432d0a1cb7a6bfc54a1cd19e75d1cec1a6b16

Observation 1e058d5a-d39d-46f1-af68-ebd671d072ec · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:34.967395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.647682Z digest=sha256:6afcf47b5927ebf365ec8a1e3294408320c7b3976fb1be4f8a53d3fedacd1bc3

Observation 7295c9bf-9a7e-4229-9ae1-8e8eb082b4e2 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:34.917075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.669349Z digest=sha256:da00204ff465981eba508fcb7f7d1e50703711031836b1b66cb10a882050cd72

Observation 70d86c36-21bb-43a2-828c-5ea64e462e26 · outbound

This paper cites an unresolved cited work.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:38:34.902737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.723346Z digest=sha256:923abb2c060b37ef83ef5ceececb920d68cbd297232a2b6368a83e425ea6cee2

Observation 919cee17-65ca-4fa3-b275-6fa30cb2a7c2 · outbound

This paper cites In: Proceedings of the AAAI conference on artificial intelligence.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the AAAI conference on artificial intelligence

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:34.895432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.746698Z digest=sha256:ccb2f5205460c997d11e86e344e3407ebcfb0a26755478ffc2a946fa5b459d50

Observation 652a7513-062d-4bf7-8d56-9714bd8b9d39 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:34.887328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:38:34.786612Z digest=sha256:d30ed06c42306d1b27fa68afd96c696427972e3473e70b705629258c801fe4a6

Pith citing papers

No inbound Pith citation observations are available.