REVIEW 3 major objections 4 minor 65 references
Dynamic Scene Understanding from Vision-Language Representations
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A frozen vision-language model's embeddings, used either as structured text decoding or as concatenated visual features, are enough to reach state-of-the-art results across four dynamic scene understanding tasks.
desk verdict A genuinely useful empirical result—frozen BLIP-2 features lift four dynamic-scene benchmarks with simple recipes—but the grounded-prediction mechanism is misdescribed and the SOTA claims are slightly overblown. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the frozen BLIP-2 vision-language representation: 32 unpooled embeddings of dimension 768 extracted from the Q-Former output for alternating image patches. It plays two roles. In structured text prediction it is the visual input to a frozen OPT-2.7B text decoder that is lightly adapted with LoRA, trained with a token-wise language-modeling objective on text that serializes semantic frames or human-human interaction descriptions with unambiguous markers such as 'VERB' and role names in capitals. In grounded prediction it is projected to the dimension of the existing backbone's feature map and concatenated feature-wise, so transformer attention in PViC or CoFormer sees the augmented features without any new attention weights, only a linear projection. The paper's analysis instrument is a linear probe on verb prediction, used as a proxy for how much dynamic knowledge a representation encodes.
What would settle it
Shuffle the spatial order of the 32 BLIP-2 embeddings before concatenation and measure GSR or HOI performance; if the score stays roughly flat, the gains come from global semantics rather than spatial alignment, and the paper's no-positional-encoding assumption would not be the source.
Extended reading notes
Core claim
The central discovery is that dynamic scene understanding does not need a dedicated architecture for each sub-task; the semantics needed for all four tasks are already present in the frozen, unpooled patch embeddings of a large vision-language model. Concretely, framing situation recognition and human-human interaction as generation of a single structured text string, followed by deterministic parsing into frames, lets a BLIP-2 decoder with LoRA outperform prior specialized models on imSitu and the Waldo-Wenda benchmark. For human-object interaction detection and grounded situation recognition, concatenating the same frozen embeddings onto the backbone features of PViC and CoFormer, after a linear projection and with no extra positional encodings, pushes both detectors past their un-augmented versions and past prior state of the art on HICO-DET and SWiG. The paper interprets these gains as evidence that recent vision-language representations encode dynamic knowledge, measurable by linear-probe verb prediction accuracy, and that this encoded dynamic knowledge is what makes the unified framework work.
Load-bearing premise
The method assumes the 32 frozen patch embeddings already carry positional information that aligns them with the backbone feature maps, since no extra positional encodings are added.
Editorial extensions
If this is right
- Situation recognition and human-human interaction can be reduced to image captioning with deterministic parsing, so future systems can skip task-specific structured-output heads.
- Adding frozen vision-language features to an existing grounded detector improves it without changing its attention weights, meaning the augmentation is a drop-in upgrade for future human-object interaction and grounded situation recognition models.
- Representation quality now has a cheap proxy: linear-probe verb accuracy predicts downstream dynamic-scene performance, so model selection for these tasks can be guided by that score.
- Because the framework is backbone-agnostic, any future vision-language model with better dynamic semantics should transfer directly to all four benchmarks without new architectures.
Reading between the lines
- A testable extension the authors leave implicit: shuffle the spatial order of the 32 BLIP-2 embeddings before concatenation; if performance holds, the benefit is global semantic content rather than spatial correspondence.
- The same concatenation trick could plausibly extend to other structured-output vision tasks such as scene graph generation, where global situation semantics and local relations are both needed.
- The linear-probe correlation suggests that pretraining on action- or verb-centric captions may be a direct route to improving dynamic scene understanding, something the paper identifies as open but does not test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a unified framework for dynamic scene understanding tasks—situation recognition (SiR), human-human interaction (HHI), grounded situation recognition (GSR), and human-object interaction (HOI)—by leveraging frozen BLIP-2 vision-language representations. For high-level tasks, it frames predictions as structured text and fine-tunes the BLIP-2 OPT-2.7B decoder with LoRA. For grounded tasks, it concatenates frozen BLIP-2 embeddings to existing backbones (CoFormer for GSR, PViC for HOI) via a projection layer. Experiments on imSitu, Waldo & Wenda, SWiG, and HICO-DET report improvements over several prior methods. The paper also analyzes dynamic knowledge in various V&L representations using linear probing of verb prediction.
Significance. If the central claims hold, the paper offers a simple, partially generic recipe for four distinct dynamic-scene tasks, reducing task-specific engineering and using few trainable parameters. The structured-text formulation for SiR and HHI is clean, and the attention feature augmentation improves two existing grounded models. The paper ships qualitative results and a project page, and the experimental design is mostly standard. However, the headline claim of state-of-the-art results across the board is contradicted by the GSR results on a key metric, the grounding mechanism is not validated for spatial alignment, and the lack of code or multiple-seed statistics limits verification of the often modest gains.
major comments (3)
- [Abstract, §5.2, Table 3] The abstract and introduction claim state-of-the-art results across all four tasks, but Table 3 shows that on GSR with ground-truth verb, the proposed method's value-all is 27.28, well below ClipSitu XTF's 33.20, and its top-5 value-all is also lower (24.72 vs 25.22). The claim should be qualified to specific tasks and metrics, or a justification should be given for why the GT-verb value-all deficit does not affect the SOTA claim. As written, the overclaim is load-bearing because the paper's main contribution is a universal recipe with uniform SOTA behavior.
- [§3.2, §5.1] The attention feature augmentation concatenates 32 BLIP-2 Q-Former output embeddings to CNN backbone feature maps without adding positional encodings, 'assuming that these are already present within the existing and newly added features.' This assumption is not established: BLIP-2's Q-Former produces 32 learned, orderless query embeddings attending globally to the image, not a spatial grid of patch tokens. If the embeddings are not spatially aligned with the K spatial locations of the backbone, the concatenation injects global context rather than 'strictly more grounded knowledge for localized predictions.' The paper should provide an alignment check (e.g., shuffling the 32 embeddings or replacing them with grid-aligned ViT patch tokens) and, depending on the outcome, revise the mechanistic claim. This is important because the generality of the recipe depends on whether the augmentation's effect is truly spatial grounding.
- [§5, Tables 1–4] All reported numbers are single runs without error bars, multiple seeds, or statistical significance tests. Some improvements are small (e.g., SiR verb 58.88 vs 58.19; HOI non-rare 46.21 vs 45.64), and the claimed contributions would be more convincing with variance estimates or released code to verify. Given that the paper positions itself as a generic framework replacing task-specific engineering, reproducibility is essential.
minor comments (4)
- [§5.4, Table 5] The table header 'BackboneAttentionFeatureAugmentation' is missing a separating space between 'Backbone' and 'Attention'.
- [§3.1] The likelihood equation is typeset with garbled subscripts and spacing; it should be presented with clean notation for the token-wise language modeling objective.
- [§5.3, Figure 8] The claimed correlation between linear probing verb accuracy and HOI performance is based on only four embedding types. Adding the actual data points and a correlation coefficient would strengthen the claim, or the wording should be softened to 'visual trend'.
- [Appendix C] Typo: 'priovide' should be 'provide'.
Circularity Check
No significant circularity: the framework's predictions are evaluated on external benchmarks and no derived quantity reduces to a fitted input by construction.
full rationale
The paper's central claims rest on empirical evaluations over four benchmarks, including the external imSitu, SWiG, and HICO-DET datasets and the Waldo/Wenda HHI benchmark whose test labels are manually written ground truth. The structured-text prediction method is a standard supervised image-to-text fine-tuning objective (token-wise log-likelihood), and the parser is an explicit bijection between semantic frames and formatted strings, so parsing correctness is a design definition rather than a derived result. The grounded-prediction method's Eq. (3.1), Fconcat = concat(E_backbone, pi(E_V&L)), is an architectural augmentation whose performance is measured on held-out test sets; it does not algebraically force any downstream metric. The linear-probing analysis in Section 5.3 is a correlational proxy (verb prediction accuracy vs. HOI performance) and is not used as a fitted value in the benchmark evaluations. The only self-citation of note is the HHI task's dependence on the Waldo and Wenda benchmark and pseudo-labels from Alper & Averbuch-Elor [2]; however, the test labels are manually written ground truth, all compared methods are trained on the same pseudo-labels, and the other three tasks are fully external, so this provenance does not make the reported gains circular or load-bearing.
Assumptions & free parameters
free parameters (1)
- LoRA rank, alpha, dropout =
128, 256, 0.05
assumptions (5)
- domain assumption Frozen BLIP-2 embeddings encode transferable dynamic scene semantics.
- domain assumption The structured-text format and rule-based parser can represent and recover SiR and HHI outputs without information loss.
- domain assumption Concatenating BLIP-2 features to backbone features requires no extra positional encoding because spatial alignment is already present.
- domain assumption Linear probing verb accuracy on imSitu is a valid proxy for dynamic scene knowledge.
- domain assumption The HHI benchmark, pseudo-labels and evaluation protocol from [2] are reliable and unbiased.
Cite this review
Pith. "Pith review of Dynamic Scene Understanding from Vision-Language Representations." pith.science (2026). https://pith.science/paper/TEMAG5MX
@misc{pith2026250111653,
author = {Pith},
title = {Pith review of: Dynamic Scene Understanding from Vision-Language Representations},
year = {2026},
howpublished = {\url{https://pith.science/paper/TEMAG5MX}},
note = {Machine review of arXiv:2501.11653}
}
read the original abstract
Images depicting complex, dynamic scenes are challenging to parse automatically, requiring both high-level comprehension of the overall situation and fine-grained identification of participating entities and their interactions. Current approaches use distinct methods tailored to sub-tasks such as Situation Recognition and detection of Human-Human and Human-Object Interactions. However, recent advances in image understanding have often leveraged web-scale vision-language (V&L) representations to obviate task-specific engineering. In this work, we propose a framework for dynamic scene understanding tasks by leveraging knowledge from modern, frozen V&L representations. By framing these tasks in a generic manner - as predicting and parsing structured text, or by directly concatenating representations to the input of existing models - we achieve state-of-the-art results while using a minimal number of trainable parameters relative to existing approaches. Moreover, our analysis of dynamic knowledge of these representations shows that recent, more powerful representations effectively encode dynamic scene semantics, making this approach newly possible.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Men- 8 sch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716–23736,
-
[2]
Learning human- human interactions in images from weak textual supervision
Morris Alper and Hadar Averbuch-Elor. Learning human- human interactions in images from weak textual supervision. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 1, 2, 4, 5, 6, 7, 12, 13, 14
work page 2023
-
[3]
Zero-shot learning via visual abstraction
Stanislaw Antol, C Lawrence Zitnick, and Devi Parikh. Zero-shot learning via visual abstraction. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part IV 13, pages 401–416. Springer, 2014. 2
work page 2014
-
[4]
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023. 2
arXiv 2023
-
[5]
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Alt- man, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021. 2
arXiv 2021
-
[6]
End-to- end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to- end object detection with transformers. In European confer- ence on computer vision, pages 213–229. Springer, 2020. 13
work page 2020
-
[7]
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In Pro- ceedings of the IEEE conference on computer vision and pat- tern recognition, page 6299–6308, 2017. 2
work page 2017
-
[8]
A comprehensive survey of scene graphs: Generation and application
Xiaojun Chang, Pengzhen Ren, Pengfei Xu, Zhihui Li, Xiao- jiang Chen, and Alex Hauptmann. A comprehensive survey of scene graphs: Generation and application. IEEE Transac- tions on Pattern Analysis and Machine Intelligence , 45(1): 1–26, 2021. 2
work page 2021
Show all 65 references
-
[9]
Hico: A benchmark for recognizing human-object interactions in images
Yu-Wei Chao, Zhan Wang, Yugeng He, Jiaxuan Wang, and Jia Deng. Hico: A benchmark for recognizing human-object interactions in images. In Proceedings of the IEEE inter- national conference on computer vision , pages 1017–1025,
-
[10]
Learning to detect human-object interactions
Yu-Wei Chao, Yunfan Liu, Xieyang Liu, Huayi Zeng, and Jia Deng. Learning to detect human-object interactions. In 2018 ieee winter conference on applications of computer vision (wacv), pages 381–389. IEEE, 2018. 4, 5, 13
2018
-
[11]
ViStruct: Visual structural knowledge extrac- tion via curriculum guided code-vision representation
Yangyi Chen, Xingyao Wang, Manling Li, Derek Hoiem, and Heng Ji. ViStruct: Visual structural knowledge extrac- tion via curriculum guided code-vision representation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13342–13357, S...
2023
-
[12]
Collab- orative transformers for grounded situation recognition
Junhyeong Cho, Youngseok Yoon, and Suha Kwak. Collab- orative transformers for grounded situation recognition. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 19659–19668, 2022. 2, 5, 6, 7, 12, 13, 14
2022
-
[13]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 12
2009
-
[14]
An image is worth 16x16 words: Trans- formers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. arXiv preprint a...
2010 arXiv
-
[15]
Background to framenet
Charles J Fillmore, Christopher R Johnson, and Miriam RL Petruck. Background to framenet. International journal of lexicography, 16(3):235–250, 2003. 4
2003
-
[16]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 6, 12
2016
-
[17]
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021. 5
2021 arXiv
-
[18]
Language is not all you need: Aligning perception with language mod- els
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, et al. Language is not all you need: Aligning perception with language mod- els. Advances in Neural Information Processing Systems , 36, 2024. 2
2024
-
[19]
What to look at and where: Se- mantic and spatial refined transformer for detecting human- object interactions
ASM Iftekhar, Hao Chen, Kaustav Kundu, Xinyu Li, Joseph Tighe, and Davide Modolo. What to look at and where: Se- mantic and spatial refined transformer for detecting human- object interactions. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognit...
2022
-
[20]
Scaling up visual and vision-language representa- tion learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representa- tion learning with noisy text supervision. In International conference on machine learning, pages 4904–4916. PMLR,
-
[21]
Detrs with hybrid matching
Ding Jia, Yuhui Yuan, Haodi He, Xiaopei Wu, Haojun Yu, Weihong Lin, Lei Sun, Chao Zhang, and Han Hu. Detrs with hybrid matching. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 19702–19712, 2023. 13
2023
-
[22]
Otter: A multi-modal model with in-context instruction tuning
Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. Otter: A multi-modal model with in-context instruction tuning. arXiv preprint arXiv:2305.03726, 2023. 2
2023 arXiv
-
[23]
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Interna- tional conference on machine learning, pages 12888–12900. PMLR, 2022. 7
2022
-
[24]
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In In- ternational conference on machine learning , pages 19730– 19742. PMLR, 2023. 2, 3, 5, 6, 7
2023
-
[25]
Gen-vlkt: Simplify association and enhance 9 interaction understanding for hoi detection
Yue Liao, Aixi Zhang, Miao Lu, Yongliang Wang, Xiaobo Li, and Si Liu. Gen-vlkt: Simplify association and enhance 9 interaction understanding for hoi detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 20123–20132, 2022. 12, 13
2022
-
[26]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceeding...
2014
-
[27]
Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation
Yuqi Lin, Minghao Chen, Wenxiao Wang, Boxi Wu, Ke Li, Binbin Lin, Haifeng Liu, and Xiaofei He. Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...
2023
-
[28]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023. 2
2023
-
[29]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36, 2024. 2
2024
-
[30]
Open scene understanding: Grounded situation recognition meets segment anything for helping people with visual im- pairments
Ruiping Liu, Jiaming Zhang, Kunyu Peng, Junwei Zheng, Ke Cao, Yufan Chen, Kailun Yang, and Rainer Stiefelhagen. Open scene understanding: Grounded situation recognition meets segment anything for helping people with visual im- pairments. In Proceedings of the IEEE/CVF Internat...
2023
-
[31]
Image segmenta- tion using text and image prompts
Timo L ¨uddecke and Alexander Ecker. Image segmenta- tion using text and image prompts. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7086–7096, 2022. 2
2022
-
[32]
Clip4hoi: Towards adapting clip for practi- cal zero-shot hoi detection
Yunyao Mao, Jiajun Deng, Wengang Zhou, Li Li, Yao Fang, and Houqiang Li. Clip4hoi: Towards adapting clip for practi- cal zero-shot hoi detection. Advances in Neural Information Processing Systems, 36, 2024. 2
2024
-
[33]
Clip- cap: Clip prefix for image captioning
Ron Mokady, Amir Hertz, and Amit H Bermano. Clip- cap: Clip prefix for image captioning. arXiv preprint arXiv:2111.09734, 2021. 6, 14
2021 arXiv
-
[34]
Hoiclip: Efficient knowledge transfer for hoi detection with vision-language models
Shan Ning, Longtian Qiu, Yongfei Liu, and Xuming He. Hoiclip: Efficient knowledge transfer for hoi detection with vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 23507–23517, 2023. 2
2023
-
[35]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 7
2023 arXiv
-
[36]
Kosmos-2: Ground- ing multimodal large language models to the world
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Ground- ing multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023. 2
2023 arXiv
-
[37]
Grounded situation recognition
Sarah Pratt, Mark Yatskar, Luca Weihs, Ali Farhadi, and Aniruddha Kembhavi. Grounded situation recognition. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV 16, pages 314–332. Springer, 2020. 1, 2, 4, 5, 6, 7, 12, 13, 14
2020
-
[38]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Ja ck Clark, et al. Learn- ing transferable visual models from natural language super- vision. In International conference on machine learnin...
2021
-
[39]
Robotic vision for human-robot interaction and collaboration: A survey and systematic re- view
Nicole Robinson, Brendan Tidd, Dylan Campbell, Dana Kuli´c, and Peter Corke. Robotic vision for human-robot interaction and collaboration: A survey and systematic re- view. ACM Transactions on Human-Robot Interaction , 12 (1):1–66, 2023. 1
2023
-
[40]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 1
2022
-
[41]
Clip- situ: Effectively leveraging clip for conditional predictions in situation recognition
Debaditya Roy, Dhruv Verma, and Basura Fernando. Clip- situ: Effectively leveraging clip for conditional predictions in situation recognition. In Proceedings of the IEEE/CVF Win- ter Conference on Applications of Computer Vision , pages 444–453, 2024. 2, 6, 12, 13, 14
2024
-
[42]
Ut-interaction dataset, icpr contest on semantic description of human activities (sdha)
Michael S Ryoo and JK Aggarwal. Ut-interaction dataset, icpr contest on semantic description of human activities (sdha). In IEEE International Conference on Pattern Recog- nition Workshops, page 4, 2010. 4
2010
-
[43]
Bleurt: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P Parikh. Bleurt: Learning robust metrics for text generation. arXiv preprint arXiv:2004.04696, 2020. 4
2004 arXiv
-
[44]
Analyzing human– human interactions: A survey
Alexandros Stergiou and Ronald Poppe. Analyzing human– human interactions: A survey. Computer Vision and Image Understanding, 188:102799, 2019. 1, 4
2019
-
[45]
Generative multimodal mod- els are in-context learners
Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiy- ing Yu, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative multimodal mod- els are in-context learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , ...
2024
-
[46]
Qpic: Query-based pairwise human-object interaction detec- tion with image-wide contextual information
Masato Tamura, Hiroki Ohashi, and Tomoaki Yoshinaga. Qpic: Query-based pairwise human-object interaction detec- tion with image-wide contextual information. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10410–10419, 2021. 2
2021
-
[47]
Learning latent temporal structure for complex event detection
Kevin Tang, Li Fei-Fei, and Daphne Koller. Learning latent temporal structure for complex event detection. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 1250–1257, 2012. 2
2012
-
[48]
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE con- ference on computer vision and pattern recognition , pages 4489–4497, 2015
2015
-
[49]
Actionclip: A new paradigm for video action recognition
Mengmeng Wang, Jiazheng Xing, and Yong Liu. Actionclip: A new paradigm for video action recognition. arXiv preprint arXiv:2109.08472, 2021. 2
2021 arXiv
-
[50]
Cris: Clip- driven referring image segmentation
Zhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao, Yandong Guo, Mingming Gong, and Tongliang Liu. Cris: Clip- driven referring image segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11686–11695, 2022. 2 10
2022
-
[51]
Rethinking the two-stage framework for grounded sit- uation recognition
Meng Wei, Long Chen, Wei Ji, Xiaoyu Yue, and Tat-Seng Chua. Rethinking the two-stage framework for grounded sit- uation recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 2651–2658, 2022. 2, 6, 12, 13
2022
-
[52]
Rec- ognize complex events from static images by fusing deep channels
Yuanjun Xiong, Kai Zhu, Dahua Lin, and Xiaoou Tang. Rec- ognize complex events from static images by fusing deep channels. In Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition , pages 1600–1609,
-
[53]
Recognizing proxemics in personal photos
Yi Yang, Simon Baker, Anitha Kannan, and Deva Ramanan. Recognizing proxemics in personal photos. In 2012 IEEE Conference on Computer Vision and Pattern Recognition , pages 3522–3529. IEEE, 2012. 2
2012
-
[54]
Situa- tion recognition: Visual semantic role labeling for image understanding
Mark Yatskar, Luke Zettlemoyer, and Ali Farhadi. Situa- tion recognition: Visual semantic role labeling for image understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5534–5542,
-
[55]
Rlip: Rela- tional language-image pre-training for human-object inter- action detection
Hangjie Yuan, Jianwen Jiang, Samuel Albanie, Tao Feng, Ziyuan Huang, Dong Ni, and Mingqian Tang. Rlip: Rela- tional language-image pre-training for human-object inter- action detection. In Advances in Neural Information Pro- cessing Systems (NeurIPS), 2022. 12
2022
-
[56]
Rlipv2: Fast scaling of re- lational language-image pre-training
Hangjie Yuan, Shiwei Zhang, Xiang Wang, Samuel Al- banie, Yining Pan, Tao Feng, Jianwen Jiang, Dong Ni, Yingya Zhang, and Deli Zhao. Rlipv2: Fast scaling of re- lational language-image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision , p...
2023
-
[57]
Lit: Zero-shot transfer with locked-image text tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer. Lit: Zero-shot transfer with locked-image text tuning. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 18123–18133, 2022. 2
2022
-
[58]
Zhang, Dylan Campbell, and Stephen Gould
Frederic Z. Zhang, Dylan Campbell, and Stephen Gould. Efficient two-stage detection of human-object interactions with a novel unary-pairwise transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20104–20112, 2022. 2
2022
-
[59]
Zhang, Yuhui Yuan, Dylan Campbell, Zhuoyao Zhong, and Stephen Gould
Frederic Z. Zhang, Yuhui Yuan, Dylan Campbell, Zhuoyao Zhong, and Stephen Gould. Exploring predicate visual con- text in detecting human–object interactions. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), pages 10411–10421, 2023. 1, 2, 5, 6, 8, 12
2023
-
[60]
Opt: Open pre-trained trans- former language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained trans- former language models. arXiv preprint arXiv:2205.01068,
-
[61]
Zegclip: Towards adapting clip for zero-shot se- mantic segmentation
Ziqin Zhou, Yinjie Lei, Bowen Zhang, Lingqiao Liu, and Yifan Liu. Zegclip: Towards adapting clip for zero-shot se- mantic segmentation. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 11175–11185, 2023. 2
2023
-
[62]
Scene graph generation: A comprehensive survey
Guangming Zhu, Liang Zhang, Youliang Jiang, Yixuan Dang, Haoran Hou, Peiyi Shen, Mingtao Feng, Xia Zhao, Qiguang Miao, Syed Afaq Ali Shah, et al. Scene graph generation: A comprehensive survey. arXiv preprint arXiv:2201.00443, 2022. 2
2022 arXiv
-
[63]
A comprehensive study of deep video action recognition
Yi Zhu, Xinyu Li, Chunhui Liu, Mohammadreza Zolfaghari, Yuanjun Xiong, Chongruo Wu, Zhi Zhang, Joseph Tighe, R Manmatha, and Mu L. A comprehensive study of deep video action recognition. arXiv preprint arXiv:2012.06567, 2020. 2 11 A. Additional Evaluations and Comparisons A.1....
2012 arXiv
-
[65]
features specifically designed for the human-object in- teraction (HOI) task. C.2. Situation Recognition (SiR) Experimental details . Experiments over this task were trained and evaluated on the imSitu dataset [54]. The imSitu dataset contains 75K, 25K and 25K images for train...
-
[2016]
1, 2, 4, 5, 6, 7, 12, 13, 14
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.