Pith. sign in

Paper Citation Record · LEDGER

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis

As of 21 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2411.18038.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18038 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:38:34.786612Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2a5237ee-346a-4d9e-8d89-29e96b39bfe1 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.146453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.146453Z digest=sha256:7a7e78c172ff4060759b53e38fcf7f43d701d5c60525d68fe14a2cdb5c8a9279

Observation d0e264e1-1968-45a2-a10c-1191df854adc · outbound

This paper cites In: Proceedings of the IEEE international conference on computer vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE international conference on computer vision

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.844543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.150887Z digest=sha256:fca5307a01712f764b8fb3c7178acf12d7adfa6e320017a622cd0725c0f8b8c9

Observation 49eafe1f-9012-439b-bebf-cce652cfce60 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence 41(2), 423–443 (2019).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis IEEE Transactions on Pattern Analysis and Machine Intelligence 41(2), 423–443 (2019)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.837010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.153944Z digest=sha256:764a78bc1af47223084a7de9f205aa4b59020caae83f9a93a682485468915a94

Observation 88b30466-0e00-46ad-ba7f-467c18b95f20 · outbound

This paper cites Advances in neural information processing systems33, 1877–1901 (2020).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Advances in neural information processing systems33, 1877–1901 (2020)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.182657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.182657Z digest=sha256:d48f9497d8caf226be218af34087df5c2b5bdcd1c9139f6dbe2d91307cc63546

Observation 85040c99-cefb-4f47-835b-4fd33acf2181 · outbound

This paper cites In: European conference on computer vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: European conference on computer vision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.229865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.229865Z digest=sha256:b304546e6fe2735fae06fb86d65d47697e6b84c51a5f8b89830f46b311d50961

Observation c5d792de-6252-4937-b5dc-71d1c9c48ea9 · outbound

This paper cites an unresolved cited work.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:38:35.821530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.257749Z digest=sha256:ab4d594c361ba4aa98d9ef7995b32116b257999a098a2dc35ecabcc557260279

Observation 67d57506-ec78-4fc6-af1e-002d5bec64c5 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.307670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.307670Z digest=sha256:7941ee28959b479163b7bb77a26daf99227699e9844afd2f187189ee5e8d6f32

Observation 5c0fa5d8-16f1-45ed-b17e-e9edc0050153 · outbound

This paper cites In: European conference on computer vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: European conference on computer vision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.310959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.310959Z digest=sha256:a0961ab3b4c8b8569612ffcde7fda5424a7a24bb753eda42e30da0c06b354a8e

Observation 38c3e4b8-3af2-4bdd-83da-fa6738dc59cb · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis PaLM: Scaling Language Modeling with Pathways

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.313458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.313458Z digest=sha256:dbb2ec0b5039db858de1e3b356f4a2c53a5ca8d80f9abfa51717ef24eb11cf84

Observation 20608490-ee45-4f30-acd1-bfd3edbdf1d1 · outbound

This paper cites an unresolved cited work.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.316285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.316285Z digest=sha256:bddea9ac689169fe8a4e04069da2cf69b1c6cf1d78241dc84a2228d821ebc22f

Observation baaf330d-1904-45b8-9a6b-72f93c1e3007 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.319094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.319094Z digest=sha256:32d9f453cd1f358bd39d795e90ebaa56f212440a99dffddd86652a98f88c2ec3

Observation ce13505b-3646-405b-87d1-3f2fbce601ad · outbound

This paper cites In: British Machine Vision Conference (2018) 16 Kang et al.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: British Machine Vision Conference (2018) 16 Kang et al

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.702688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.322040Z digest=sha256:0c99dbe7f0ef133f1bb76ec8d5f21dd1620b5d17036da0ec9a0a1b837bb29a11

Observation 16a34395-73be-4311-b0b8-76a56171dbc7 · outbound

This paper cites Visual Semantic Role Labeling.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Visual Semantic Role Labeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.324732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.324732Z digest=sha256:ae28f45965fb05dba98b60b0c95a9cd7ae045e6ce51139c237fa7312033bea9f

Observation cfcfafc0-6767-4973-b30c-0997d789c924 · outbound

This paper cites In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XV 16.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XV 16

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.670782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.327667Z digest=sha256:3e41dba0e827ffe55f6c9bcb361e0d0ae2b5ec086064d07e12c31db861c415af

Observation 3afe03d3-9352-462f-b2e0-99bf7a364685 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.662741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.330845Z digest=sha256:4cc7609e2181819b4043f1f0887636c8cbfc0ff3abb5d94c1f78c0652e5cc09c

Observation 49b950cf-6425-4459-8b2e-8763d375c1f4 · outbound

This paper cites In: Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.653126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.333322Z digest=sha256:03c1af00186f6c58fd9e5f6d728b7d574d1f1a11254c19806d5ce48bd0950e14

Observation 6cb0126b-990e-4865-a4d0-51c92ddd8914 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.643674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.336324Z digest=sha256:3093fcba88d20a6661055c0fbe5e347a955a8d1231a52577e9c5b3726cd27d2a

Observation 468cc232-0e42-412d-a22a-152e18e28359 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.634586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.338694Z digest=sha256:805c0424b6826298233297f9bb45033bd5b6ebc45a99054948a28e68ab2b168a

Observation 8c6ff187-c0d1-4a4d-b599-4fd492dc76d0 · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.341241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.341241Z digest=sha256:20adecd05643aec965f74c720ccc78f50499e97149e2be22125464c94cc72cef

Observation 9b13d84c-4e55-4138-89f8-c3b2aec89ec1 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.344173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.344173Z digest=sha256:16eabbeeb6ce65b6bc3ff366c9a8bb6627104197f2eba8b3a8e3a3f1f439c6ef

Observation dab8e512-0800-4166-b281-c06f3201a948 · outbound

This paper cites In: International Con- ference on Machine Learning.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: International Con- ference on Machine Learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.588912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.385069Z digest=sha256:f921fd169e50690a66fe30c13558f7fe628cb41d2fd8ba49acec0b0359f196ad

Observation b6a4ca7c-a1a0-47d8-bf5b-48f1acae6504 · outbound

This paper cites Advances in neural information processing systems34, 9694–9705 (2021).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Advances in neural information processing systems34, 9694–9705 (2021)

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.440620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.440620Z digest=sha256:166a0f5166c10b6145add549d16026f34e61aa4e5ae7f9106a810b6d44e8fdf9

Observation 7907736c-465b-4605-9b58-b68df20fc304 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.443119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.443119Z digest=sha256:259f32e0faa28877ef9c1563fade912cbb001c19633415210a5ba1e7a1faf703

Observation 68157de3-a025-4aea-b66f-c00cd57e16b2 · outbound

This paper cites Advances in Neural Information Processing Systems33, 5011–5022 (2020).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Advances in Neural Information Processing Systems33, 5011–5022 (2020)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.502440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.446028Z digest=sha256:eff395f11bad1d6ca5ff44a2945900e68fca4b57a29b11e82b594f301eb45548

Observation 1d8b1ad2-6a37-4d19-bcee-c45414a361ea · outbound

This paper cites In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.365178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.449299Z digest=sha256:28682869dbf4e7807053375d4f13cf4f46faab6355ec31323911fc931fe0f975

Observation 4072d5d1-da5e-4bbb-9427-779fcd30977e · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.356245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.451737Z digest=sha256:57454b1425d9a630e048d443cf490e6f7c9c125418a4e9dd45905cb9a8dc8edf

Observation b966a311-15c0-479a-8d53-fcdc31acef21 · outbound

This paper cites In: Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.454100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.454100Z digest=sha256:2920ae5806cb2efc3f27e75e70787577f57881ea97d84380bd9440730b01890b

Observation ad934a85-b42e-47ef-9f5f-3b27c417013b · outbound

This paper cites Visual Instruction Tuning.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Visual Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.457601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.457601Z digest=sha256:3deb793c4715ec66c25a42df4f77d6c770cf4202600486d76d41666778e3dfee

Observation 39f77caa-80f2-428e-94a3-32c7a8f5a8a4 · outbound

This paper cites Decoupled Weight Decay Regularization.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.460943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.460943Z digest=sha256:5c4cd78608db88c28441eaeb1e308f283a11051c94ea20c066f9d05f3c57b6ce

Observation b4590073-7e33-49eb-91f9-c9faf21b51ec · outbound

This paper cites In: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.266247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.463850Z digest=sha256:de434e29ff5f7dbeba392c22bb885fd94296a6192a62975c19978889745267e5

Observation b3e67d8e-369d-4462-978f-93c180f89bb8 · outbound

This paper cites FGAHOI: Fine-Grained Anchors for Human-Object Interaction Detection.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis FGAHOI: Fine-Grained Anchors for Human-Object Interaction Detection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.466506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.466506Z digest=sha256:7a20374f5beb2b8ca2675aadb8f1905aba55bcd7f4ccad71d598e4e69ab1fd6f

Observation e0fbbeee-727a-4b37-8f9e-9fde518ce149 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.188249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.469599Z digest=sha256:eda03002f99cab582aa626a014c53197b85d6a44509b48462d879f5edb848682

Observation 2e4a9699-9d60-4685-9290-cc2f3d8290e9 · outbound

This paper cites an unresolved cited work.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.481853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.481853Z digest=sha256:e49ba4b5b6ad0b729437c1ba40044ff1db32d49e94dff9403b090f2a339eaf38

Observation fb43d5f5-9c30-4e0d-a2dd-3b225be02050 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.175904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.545966Z digest=sha256:0534abfde09bce44863ba07b92224d7d4994e8836bae3c0d85b64ddb8294b328

Observation 0f8ebe15-af04-4b86-8431-b3b2de892750 · outbound

This paper cites In: International conference on machine learning.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: International conference on machine learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.572935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.572935Z digest=sha256:9b5434727f2aa542d663de1ff835a91a607c0b5a12e8c4540a7ba97fcfb696fd

Observation b4e12a1d-933c-4d96-959e-d9ade6969be9 · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.575819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.575819Z digest=sha256:d6ac9ab0bbe097efa5bacf5732d63fc0f29554902e19e4fcba2dc22d43f6825a

Observation fe88a09e-a0fb-4683-b71e-33abee063a87 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.578827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.578827Z digest=sha256:c853d8b3d5091ac34d66c79d3614d2a653f0ed1807f0d67397c74d10b9c93dd6

Observation b107420d-4051-46e9-80ec-a6028a59b865 · outbound

This paper cites LXMERT: Learning Cross-Modality Encoder Representations from Transformers.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.581837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.581837Z digest=sha256:b1f2b5c607d73fd3c5e229e8d9c9226e9c2de1281be49ae65546dfdfb389ef42

Observation d35a03d6-eefc-4750-8d05-0c575b4f22c2 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.585350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.585350Z digest=sha256:cadd515dfb782690eade2e198af90f49e6df3b3bfcbf63433e15250ba5743c34

Observation 92a2bf5e-367c-4af0-a9cf-aef629b6e4e3 · outbound

This paper cites In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.155168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.587955Z digest=sha256:f65efd689d314d32c53da550b6ea3e0682825cade21d6390310d946c7322df11

Observation 0c545409-b1fa-429c-aec1-bd44cbf335ce · outbound

This paper cites In: Proceedings of the European conference on computer vision (ECCV).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the European conference on computer vision (ECCV)

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.590551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.590551Z digest=sha256:792915b89c606029e5d8b2fe03a499e2084e51bb7c688696efad3434a62f2751

Observation 9f4167a1-bedd-4c95-ae2b-021c5018a1d8 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.142769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.592927Z digest=sha256:21b40c83aedfb5f656b1806e851fe2001c681b9615e20d7888b6343203f00720

Observation 0aa0769e-cb68-4b33-86d1-a54c3a8f39e8 · outbound

This paper cites In: International conference on machine learning.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: International conference on machine learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.135268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.595219Z digest=sha256:083d5317b333588078db6cc41be9f41f5f0cf6c5adec4a6eb83c10bc8a154422

Observation 4aa3d3be-222d-4f32-ad08-a1709fca3df6 · outbound

This paper cites In: Proceedings of the AAAI conference on ar- tificial intelligence.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the AAAI conference on ar- tificial intelligence

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.597799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.597799Z digest=sha256:a2b88d290dc6bfca82201d3b1580521ca1cbba7242858090520a8f000e4bef3d

Observation e68e1ace-973d-47f8-98b4-5382e19484b6 · outbound

This paper cites In: Proceedings of the European conference on computer vision (ECCV).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the European conference on computer vision (ECCV)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.063951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.600717Z digest=sha256:9d227414d58bf47e20fdbfd2b53fcece642f1f42cc95f275e51c13b9f2214102

Observation 2259e9ad-ec0a-4263-a383-d3de78fb876f · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T11:38:34.603288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:38:34.603288Z digest=sha256:920d1c544d61918e868b1d8b197721ba2e5f7c94571d9fdc3201cb4251517659

Observation 51144232-2b26-4a49-814f-47954a3cf100 · outbound

This paper cites Ad- vances in Neural Information Processing Systems35, 37416–37431 (2022).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Ad- vances in Neural Information Processing Systems35, 37416–37431 (2022)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.056817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.605972Z digest=sha256:cc846ae6339d78a4eb4eb93944f88dfee594ee85172f807a4b96732bf509e1dd

Observation e59024cd-59db-416d-9f8d-5aaee0c493b5 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.049340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.609372Z digest=sha256:39682c845f666b602819504f822d6e2f3330d1da97bbb1bcd277d2c573b98e6b

Observation 7daa7899-863e-45fb-9a7c-dfced26d0ea8 · outbound

This paper cites Advances in Neural Information Processing Systems 34, 17209–17220 (2021).

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Advances in Neural Information Processing Systems 34, 17209–17220 (2021)

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.041544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.611855Z digest=sha256:8396f1d004c28d286232f5ba43e82f6941c5b8dd4072effc9b1a5ece70da03d4

Observation e2cc35e1-7672-4cb1-96a8-009b93664ed7 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Con- ference on Computer Vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF International Con- ference on Computer Vision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.033708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.614276Z digest=sha256:f531875669bf68608f7bd87b3b8ceb3e34971f8d170ad161f8034aa8e7022822

Observation e3cb7b2c-7c06-4f1f-b2c5-99a5d09d6a06 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.024223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.617496Z digest=sha256:c121f5134d73494758cdb6686bcc61cb70c4ae93d33231f5bee5ae4aca8f7489

Observation 36ab7426-2f8c-4757-82b5-f92ab22e4fe5 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:35.014544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.619939Z digest=sha256:04717baecc56f9d862fc7d53b76d2f5bf048a4c45f7e6434be2f48b9ba830a1a

Observation 1e058d5a-d39d-46f1-af68-ebd671d072ec · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:34.967395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.647682Z digest=sha256:11699bb0c41c9fccf41541a0fe0d06e9874f8efe9db2b3f35591e0d6bf4a1a44

Observation 7295c9bf-9a7e-4229-9ae1-8e8eb082b4e2 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:34.917075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.669349Z digest=sha256:dc94a7235de4c4512a4eefb55cbd9efe090936bd6f549435d16fe6d6499843fd

Observation 70d86c36-21bb-43a2-828c-5ea64e462e26 · outbound

This paper cites an unresolved cited work.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:38:34.902737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.723346Z digest=sha256:6fe25c66d2baa059dd9d4df31d08697387675a9f11c27cb35bfb3cd8b7220ef4

Observation 919cee17-65ca-4fa3-b275-6fa30cb2a7c2 · outbound

This paper cites In: Proceedings of the AAAI conference on artificial intelligence.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the AAAI conference on artificial intelligence

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:34.895432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.746698Z digest=sha256:0be530b240bdd03c2ccd980a423722612b3af238d945ed688ac1246c2c198474

Observation 652a7513-062d-4bf7-8d56-9714bd8b9d39 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition.

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis In: Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:38:34.887328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T11:38:34.786612Z digest=sha256:1314f575e543f1b42c15401d577a61d852b3b82ad39af3c4d7e230c5cd261cf6

Pith citing papers

No inbound Pith citation observations are available.