Pith. sign in

Paper Citation Record · LEDGER

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models

As of 20 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2501.00432.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00432 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:56:23.789593Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87e86ae7-c9a5-44ca-a0f7-d84752eff623 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.603998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.603998Z digest=sha256:c862b9a32853f480c900c699c48d34d3f43b0c22e1179826d23cffc3ab135ebf

Observation 38df0fac-9e36-4d14-92a5-48fe7147a714 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.610325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.610325Z digest=sha256:ac87712e36a7b5385d117290fc71c15226d4787bd626aa4805a9d36b67136cec

Observation b2a1420b-504d-4ce1-8243-97e034835968 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.472616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.615502Z digest=sha256:7fcede79f0aabba701267eb43027dabfccae64211b5764730eff660b8cbd505c

Observation bb9d0b30-a686-48cc-9731-7bcfe6af0515 · outbound

This paper cites Language-grounded dynamic scene graphs for interactive object search with mobile manipulation,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Language-grounded dynamic scene graphs for interactive object search with mobile manipulation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.455214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.620658Z digest=sha256:daba3d26733b3a547bdb078ca853c5ed19ae3a79cfc289a4ea83318381b16266

Observation 4c4088ad-c387-4a9e-86b6-dd42d59b8338 · outbound

This paper cites Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.625900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.625900Z digest=sha256:989eb760bf028e9a333f9659e65fb80e374bd3aab5c2b7ade33fa82da3d3d0cb

Observation 065d66d5-1452-4bf7-a252-c822fe8f44e6 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Code Llama: Open Foundation Models for Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.631491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.631491Z digest=sha256:86d46e031254d7212fb35972e15d6b2b25264c9be6e7e2dadf2cb22040d5a5df

Observation bc779ece-c635-45cb-bb9a-a23e874f2ed6 · outbound

This paper cites CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.637769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.637769Z digest=sha256:86539edf62bfff248b6baa376cf02bb8a604b0f3722d52f182b779bfef5b70b5

Observation 03ec1088-e661-4fce-8483-36849b6b1f6b · outbound

This paper cites Text me the data: Generating ground pressure sequence from textual descriptions for har,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Text me the data: Generating ground pressure sequence from textual descriptions for har,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.437533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.642943Z digest=sha256:040e0953a8635d41e42894f081a7f2e527bbfa3636f53aedc1ca25bd07ec9015

Observation d096f72a-b7db-47f1-8beb-f690a8bd9a96 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.647629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.647629Z digest=sha256:0060668f095bdcbb00b4ce9fbd9ed244d6b32ba0ee13d6f2e6b7259ebe6c3a8d

Observation 72ec3e0b-5e59-4793-b5ea-8f0473dbf639 · outbound

This paper cites Bliva: A simple multimodal llm for better handling of text-rich visual questions,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Bliva: A simple multimodal llm for better handling of text-rich visual questions,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.421483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.652741Z digest=sha256:57136ef82240ddbae87ea671299e9cb42e9ef9062d36c769062d5bbdc2acba35

Observation 6cdbf847-6e40-47f2-885b-ce49ad5fcced · outbound

This paper cites Infogcn: Representation learning for human skeleton-based action recognition,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Infogcn: Representation learning for human skeleton-based action recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.405639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.657569Z digest=sha256:ea670d8c4470e2127bf128f35ce93d36dc8fb63f4ed6920da13dd8040a612674

Observation e2742733-127e-4afb-b79f-0fe78d0b7222 · outbound

This paper cites MuJo: Multimodal Joint Feature Space Learning for Human Activity Recognition.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models MuJo: Multimodal Joint Feature Space Learning for Human Activity Recognition

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:56:24.011915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.662217Z digest=sha256:78d5e7b5a7c8401e9fd915ad23bb984c91010af25fd0805b7ab33145a577e0c9

Observation c0d010bc-eef7-4ae9-850f-5435a80c912e · outbound

This paper cites Decoupled spatial- temporal attention network for skeleton-based action-gesture recogni- tion,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Decoupled spatial- temporal attention network for skeleton-based action-gesture recogni- tion,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.387833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.667838Z digest=sha256:3a4d9034a86b7e62fa1ab3bfc8ee57ac4a3e4ec064d43480bbdae245b173a72e

Observation bde806a5-987e-46eb-a94e-645309497de0 · outbound

This paper cites ALS-HAR: Harnessing Wearable Ambient Light Sensors to Enhance IMU-based Human Activity Recogntion.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models ALS-HAR: Harnessing Wearable Ambient Light Sensors to Enhance IMU-based Human Activity Recogntion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.672238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.672238Z digest=sha256:a844a2a27ed944eb9fe63ca19fc84b96e5258c5e36886992679e7cdbe69b169c

Observation 4617cad2-2272-4829-9dea-f1deff563e5c · outbound

This paper cites Human- to-human interaction detection,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Human- to-human interaction detection,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.371202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.677125Z digest=sha256:59840818b8fe99d71eb6c5b921959d90240138ee862994a1a5da4ec16b75b8ca

Observation 64f0c954-6907-4a33-89d6-f0f86dd997b9 · outbound

This paper cites A Two-stream Hybrid CNN-Transformer Network for Skeleton-based Human Interaction Recognition.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models A Two-stream Hybrid CNN-Transformer Network for Skeleton-based Human Interaction Recognition

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:56:23.971278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.681815Z digest=sha256:4c9508c84c55d43f48b4aa39c291ac441bbc4cc08984b38428f7c769783d1998

Observation 395dfb42-dcd6-46c8-b341-a0f9637baa1d · outbound

This paper cites Hargpt: Are llms zero-shot human activity recognizers?,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Hargpt: Are llms zero-shot human activity recognizers?,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.355129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.686911Z digest=sha256:48d0c7e100c6567bb268be5f44ca61eeee5f7dd86b90ad11775b01b7629805f3

Observation 657d9849-0e6d-4ebf-8902-ccfec0004185 · outbound

This paper cites Unsupervised Human Activity Recognition through Two-stage Prompting with ChatGPT.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Unsupervised Human Activity Recognition through Two-stage Prompting with ChatGPT

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.691724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.691724Z digest=sha256:4627abd05d8301e5326f05828c84fbfd2b1c749294b0a45034aef2f3167852b7

Observation 23e7e97a-c047-4ddd-9efe-647631874ae8 · outbound

This paper cites Segment anything,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Segment anything,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.338026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.696802Z digest=sha256:1eeb052af46c610e8b4f5d73d6e5ac2b8b5ab1e0732f446d089dfcda39191831

Observation 33c1a0c1-e98c-4e43-85b1-977c4e80d32a · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.320721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.701590Z digest=sha256:bca1199dec4f8dfd2691f0a7bf10f52e4cbb909b94a66bb99a632f114740e85b

Observation 5133d8c2-772f-4b07-92b2-f6779863812e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.706672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.706672Z digest=sha256:4af2cd40ec8e6a1a60bfc6c15eade096f7eafd2b047fb207fc81383b0d9286b3

Observation bd55a662-375e-404f-a130-6b7fa18f9898 · outbound

This paper cites GPT-4 Technical Report.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models GPT-4 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.711560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.711560Z digest=sha256:0a5310a554fda01dc4f791a466c10930f3a8941bc171983e533e738b0dba656b

Observation c3bafba6-9543-4549-bc43-8a2c5ec89cca · outbound

This paper cites Track Anything: Segment Anything Meets Videos.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Track Anything: Segment Anything Meets Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.716449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.716449Z digest=sha256:e981b6fa44d98f243244158457a4904ef8292b45d1d4d5d151f95592f00f4ffd

Observation b899e3e7-ec58-4afc-83ad-3c78f8d0bb6b · outbound

This paper cites Actions in context,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Actions in context,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.304524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.722466Z digest=sha256:3cc61a571e986488587301dfe0879ef30d3d1ff80596c8e4b5289e7f023f36fc

Observation f2ed5518-f78d-4e2f-aa19-02a0c6d631ca · outbound

This paper cites Sportshhi: A dataset for human-human interaction detection in sports videos,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Sportshhi: A dataset for human-human interaction detection in sports videos,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.288579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.727378Z digest=sha256:a6a3f113894a6a34fb789b0fc3405a5f74da37e07ba029c8e19f8298673f87d0

Observation 81615868-aa76-4666-8700-2b1e3b646c9e · outbound

This paper cites High five: Recognising human interactions in tv shows.,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models High five: Recognising human interactions in tv shows.,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.272260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.732128Z digest=sha256:f17ba0d27be5c52670fe717b102beb53c2d2fad62bbd2bf835876fc1c3fab9e0

Observation 775c389b-6dff-4080-9af3-82b417bfe29c · outbound

This paper cites Two-person interaction detection using body- pose features and multiple instance learning,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Two-person interaction detection using body- pose features and multiple instance learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.255791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.736755Z digest=sha256:636dc6bced23cc26b1e8d787c17264fe96fac9a7b393aa6c7683d4d32f03c378

Observation edb78eb0-4d28-49be-9d84-56f615613598 · outbound

This paper cites First-person activity recognition: What are they doing to me?,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models First-person activity recognition: What are they doing to me?,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.239362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.741467Z digest=sha256:b567d03752f757f925347e9373f7ae66e2c0ad634a2d4283aefa5a9bdbba85c4

Observation 11898bc6-75b9-4896-8fde-d47ebddf5a3d · outbound

This paper cites Interaction relational net- work for mutual action recognition,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Interaction relational net- work for mutual action recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.222094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.746258Z digest=sha256:a5be592ddd9612357aab908c89689cf97bef2a795ee33fef874eb8dd1082a877

Observation b0c9739a-7184-421c-ac69-12b35725793c · outbound

This paper cites The Kinetics Human Action Video Dataset.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models The Kinetics Human Action Video Dataset

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.750838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.750838Z digest=sha256:709e6a5ecc8add79107f024d2450a47f09f2e5669081a41ed05de6c0cdd7f3fc

Observation aa397625-c615-4051-ad4f-7954a0ce4db9 · outbound

This paper cites Air- act2act: Human–human interaction dataset for teaching non-verbal social behaviors to robots,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Air- act2act: Human–human interaction dataset for teaching non-verbal social behaviors to robots,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.204484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.756151Z digest=sha256:633cabb9f73803e246fdac166932e22acda767313aaa40e6b067baf1b0aec7db

Observation 579d3cd5-eaba-4fd1-82c2-27bad80a2ba5 · outbound

This paper cites Human behavior under- standing,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Human behavior under- standing,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.187189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.760835Z digest=sha256:985f9d3b3bb3dfe36c3cf38c400cea1bf94e7f70c5b204f727688ba6c1f49587

Observation ab50a7af-e2f7-4ff8-bc9e-dc65f055ff5a · outbound

This paper cites Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.169124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.766462Z digest=sha256:74264e2553d6268b5b85f6dbb84c8fd3ade6ce9bcf9260a6dd55dd743590a2eb

Observation b619f521-e1ed-4712-b6f6-6b22246a16ba · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.771088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.771088Z digest=sha256:f1e5c9f92092511a492beb48ae7e479e0fd1d466589893e9ca39b9315d1ecc62

Observation 12207379-a3be-4412-a41d-7ea6de79d4da · outbound

This paper cites Vitpose: Sim- ple vision transformer baselines for human pose estimation,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Vitpose: Sim- ple vision transformer baselines for human pose estimation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.151800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.775670Z digest=sha256:37f421e5c799c813eea401a1cf9e039989cd7da2b025fa9cd52654e9cb3ba5c3

Observation b58fb453-5b87-4516-9d1e-db0f902034fe · outbound

This paper cites Tokens- to-token vit: Training vision transformers from scratch on imagenet,.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models Tokens- to-token vit: Training vision transformers from scratch on imagenet,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:24.133634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:56:23.780097Z digest=sha256:cae75d60d68bd17f0a7fc7fa0404079fd3a1635d2fc04c4b9fbb44367c1d253a

Observation a500a090-c4bf-43d6-acef-9ca4d3e8d58d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.784686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.784686Z digest=sha256:505d22ba82d46234fffc6bc7ffff35941668053cc7372b8ccd5a8b293317aad1

Observation 91384ef2-fb44-493f-8492-a01fe3816933 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:23.789593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:23.789593Z digest=sha256:01c949ae6cc096dcfa22c997ac066cae2d83a90c7ab96ccb5402adcca525f430

Pith citing papers

No inbound Pith citation observations are available.