{"id":"c117c139-9c93-4572-a6a7-55326eb41151","arxiv_id":"2501.00432","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A multi-branch video-language model with person-specific ViTPose encoders and a frozen LLaMA 2 generates open-vocabulary descriptions of human interactions, and a merged 103-class benchmark is introduced.","lead":"This paper proposes a video-language system that splits footage of two people into separate person crops and a background, feeds each through its own encoder, and lets a frozen large language model write an open-ended sentence describing their interaction. The authors also merge ten public interaction datasets into one benchmark, and report that their system beats VideoLLaMA variants on caption similarity and beats per-dataset classifiers on recognition.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table II lacks any stated train/test split and the paper itself calls cosine similarity misleading, so the quantitative SOTA claim is unsupported.","rationale":"The paper's central claim of open-vocabulary superiority over SOTA rests primarily on Table II. The missing train/test split makes those numbers uninterpretable: if the evaluation is in-sample, the reported gains are an artifact of training on the same videos. The authors' own admission in Section III-B that cosine similarity is misleading further weakens the metric's validity. The reader's verdict of REJECT is appropriate, as the quantitative foundation for the main claim is absent. I agree with the reader's identification of the data/protocol assumption as the weakest point, and I have narrowed it to the single most concrete issue: the undefined evaluation split. Without that, the paper's other contributions, including the unified dataset and the open-vocabulary evaluation, cannot rescue the core claim of superior performance.","tokens_in":7981,"tokens_out":5714,"duration_ms":57443,"concrete_test":"Request the exact train/test split used for Table II. Then rerun the comparison on a held-out 20% of HHIRChat that was not used for any model's fine-tuning, training all models (VideoLLaMA LoRA, VideoLLaMA-2 LoRA, OV-HHIR) on the same 80% training split. If OV-HHIR's cosine-similarity advantage over VideoLLaMA-2 LoRA shrinks to within one standard deviation or reverses, Table II's SOTA claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section III-A and Table II present the paper's headline quantitative result: OV-HHIR achieves 0.628 mean cosine similarity versus 0.414 for VideoLLaMA-2 LoRA on the HHIRChat dataset. The paper does not state whether the videos used for this evaluation were held out from training. This matters because OV-HHIR is trained with cross-entropy between LLaMA-generated and GPT-4 ground-truth embeddings (Section III), so an in-sample evaluation would trivially favor it over baselines that were not trained on the same examples. Moreover, Section III-B explicitly concedes that 'purely relying on cosine similarity can lead to incorrect evaluations' and that it is 'inaccurate,' so the metric used for the main comparison is disowned by the authors. Without a defined held-out split and with a metric the authors themselves distrust, the claim of large-margin improvement over SOTA is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes OV-HHIR, an open-vocabulary human-to-human interaction recognition framework. Videos are first decomposed into Person 1, Person 2, and background streams using Track Anything; each stream is fed to a separate vision-language branch using ViTPose or ViT encoders, Q-Formers, and linear projection layers; the fused embeddings are passed to a frozen LLaMA-2 13B chat model that is trained to generate GPT-4-produced textual descriptions. The authors also introduce HHIRChat, a unified dataset of 86,623 video sequences covering 103 interaction classes assembled from ten existing datasets. Quantitative evaluation reports cosine similarity between generated and reference embeddings on HHIRChat, macro-F1 scores on per-dataset classification, and an open-set experiment with six unseen activity classes.","tokens_in":8100,"tokens_out":4816,"duration_ms":48693,"significance":"If the claimed results are correct, the paper makes two useful contributions: a large unified human-interaction dataset with GPT-4-generated soft labels, and a multi-branch architecture that decouples segmentation, feature extraction, and language generation for interaction recognition. The open-set evaluation idea is also valuable, as fixed-vocabulary classifiers are indeed limited for emerging interaction types. However, the quantitative evidence in the current manuscript is not sufficient to support the central claims. The headline cosine-similarity comparison is computed without a stated train/test split, and the paper itself concedes that this metric can be misleading. Because the training objective is to reproduce the GPT-4 captions on the same HHIRChat data, an in-sample evaluation would strongly favor OV-HHIR over baselines that did not train on those captions. The open-set experiment likewise lacks essential details. The dataset and architecture are plausible resources, but the current evaluation does not establish superiority over existing methods.","major_comments":[{"comment":"The paper does not state whether the videos used for the cosine-similarity evaluation in Table II are held out from training. Since OV-HHIR is trained with cross-entropy between LLaMA-generated embeddings and GPT-4 ground-truth embeddings on the same HHIRChat data, an in-sample evaluation would be heavily biased in its favor. The problem is compounded by Section III-B, which explicitly states that 'Purely relying on cosine similarity can lead to incorrect evaluations' and calls the metric 'inaccurate.' Therefore the headline comparison (0.628 vs 0.414) is not established as a fair measure of generalization. The authors must specify the exact split, report which videos or classes are used for evaluation, and either justify cosine similarity as a valid metric despite their own caveat or replace it with a metric they consider reliable.","section":"Section III-A, Table II"},{"comment":"The procedure for merging ten datasets with heterogeneous label vocabularies into 103 unified interaction classes is not described. The paper simply states that GPT-4 was used to convert hard labels into descriptive soft labels, but does not provide the taxonomy mapping, the prompt used, or any validation of the generated descriptions. Since HHIRChat is both a claimed contribution and the training signal for all experiments, inconsistent or erroneous soft labels would undermine every reported result. The authors should provide a detailed label-merging protocol, examples of original and converted labels, and an analysis of label consistency and errors.","section":"Section II-A, Table I"},{"comment":"The open-vocabulary claim is supported only by a macro-F1 score of 0.315 +/- 0.023 on 'a set of 6 uncommon activities' collected from YouTube. The paper does not report the number of videos per activity, the class distribution, or whether these videos were used in any way during development of the classifier prompt. Moreover, the assertion that 'The same approach would always result in 0 for any classification model' is not demonstrated and is not generally true, since a fixed-vocabulary classifier can be designed to reject out-of-vocabulary inputs. To support the central contribution, the authors should provide per-class results, dataset statistics, and a comparison against a reasonable open-set-capable baseline.","section":"Section III-A, open-set evaluation"},{"comment":"The qualitative example in Fig. 3 illustrates a deeper problem with the chosen evaluation metric: in the second example, OV-HHIR achieves a lower cosine similarity (0.42) than VideoLLaMA-2 LoRA (0.58), even though the OV-HHIR caption is semantically correct and the LoRA caption is wrong. This directly contradicts the use of cosine similarity as the primary quantitative metric in Table II. The paper should either adopt an evaluation metric that ranks semantic correctness appropriately or explain why Table II's cosine-similarity ranking is still meaningful despite this counterexample.","section":"Section III-B, Fig. 3"}],"minor_comments":[{"comment":"The JSON example contains a malformed field: `\"bg:\"0_bg.mp4\"` is missing the colon-space and proper quoting; it should be `\"bg\": \"0_bg.mp4\"`.","section":"Figure 1"},{"comment":"The text lists '6 uncommon activities, i.e., bowing down, ear whispering, forehead kissing, lifting, and taunting' but only five activities are listed; the sixth should be added.","section":"Section III-A"},{"comment":"The row for DeepMind Kinetics appears to have only one action count and one sample count despite including multiple kinetics classes, and the NTU RGB+D row is formatted ambiguously; please clarify the table entries.","section":"Table I"},{"comment":"Reference [14] is titled 'Als-har' in the bibliography but the text refers to 'ASL-HAR'; the inconsistency should be fixed.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The central problem is the evaluation protocol: without a documented held-out split and with a metric the authors themselves disown, the reported quantitative gains are not interpretable. I would not reject outright, because the dataset and the multi-branch architecture are potentially useful contributions that could be salvaged by a substantially revised evaluation. However, the revision must include concrete details of the split, the label-merging process, and the open-set protocol before the claims can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague —\n\nThe paper has a real idea buried under an evaluation that can't support its claims. The novelty is modest but genuine: chopping video into person-level streams with ViTPose and a background branch, then feeding the aligned embeddings into a frozen LLaMA 2. That is a sensible decomposition for interaction recognition, and it's not in VideoLLaMA. The HHIRChat unification of ten datasets into 103 classes is also useful in principle, though no code or data is released.\n\nThe soft spot is the numbers. Table II reports cosine similarity on an unspecified split. Since the model is trained to match GPT-4 captions, in-sample similarity would be a tautology, and the paper never says the videos were held out. Worse, Section III-B states that cosine similarity can be misleading and inaccurate, so the headline metric is disowned by the authors. The open-vocabulary test is only six classes, and a macro-F1 of 0.315 is above chance but nothing to shout about. The dataset table lists NTU RGB+D 120 as having 11 actions, which is plainly wrong, and the baselines are trained per-dataset while OV-HHIR trains on everything, making the comparison lopsided. No ablations, no significance tests, no released code.\n\nTo be fair, the limitations section is honest about nonsense generation and memory overhead, and the authors don't hide that the cosine metric is imperfect. That candor suggests they are not trying to deceive. But the abstract and conclusion make claims that the evidence doesn't back. The architecture might be worth a follow-up, and the dataset could be a resource if properly released.\n\nWho is this for? Someone working on video-language models for interaction recognition will find the branch-decomposition idea worth discussing. But they'll need to redo the evaluation themselves. I'd bring it to a reading group as a case study in how a plausible architecture can be undercut by weak protocol.\n\nRecommendation: send to peer review, but with a strong expectation of heavy revision. The editor should ensure the reviewers demand a held-out split, a defended metric, ablations, and dataset release. As submitted, I'd reject, but the core idea deserves a second chance.\n\n— [Your name]","headline":"A genuine multi-branch architecture and a useful dataset idea are undermined by an evaluation that lacks a held-out split and uses a metric the authors themselves call misleading.","tokens_in":8672,"tokens_out":3275,"would_cite":false,"duration_ms":28643,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes OV-HHIR, a three-branch vision-language model that feeds person and background video streams to LLaMA 2, generating open-vocabulary interaction descriptions and recognizing interactions never seen in training.","keywords":["human activity recognition","human-to-human interaction","open-vocabulary recognition","video understanding","large language models","cross-modal learning","video segmentation","LLaMA"],"falsifier":"Run the exact OV-HHIR pipeline on HHIRChat with an explicit, documented train/test split (for instance, 80/20 at the video level) and compute Table II's cosine similarity only on held-out videos; if the score drops to the level of VideoLLaMA-2 LoRA or below, the advertised open-vocabulary advantage is not demonstrated. Independently, have a small panel of human annotators rate whether the GPT-4 soft labels match the video content; a high mismatch rate would indicate that the training signal itself is unreliable.","tokens_in":7754,"feed_emoji":"🤝","tokens_out":10068,"duration_ms":89378,"temperature":0.7,"pith_summary":"This paper is trying to establish that an open-vocabulary, LLM-backed video model can recognize human-to-human interactions, including interactions it has never seen, by converting each video into separate person and background streams and producing free-form textual descriptions. The motivation is that real-world surveillance and monitoring demand systems that adapt to novel, unpredictable interactions rather than being confined to a predetermined label set. To make this trainable, the paper pools ten existing datasets into one 86,623-video benchmark, HHIRChat, with 103 interaction classes and GPT-4-generated descriptive labels. The reported results put the model ahead of VideoLLaMA-2 LoRA (cosine similarity 0.628 vs. 0.414) and ahead of fixed-vocabulary classifiers on several datasets, and it is the only method that can attempt the unseen-class evaluation at all. If these numbers hold up under a clean evaluation protocol, the practical consequence is a single, general-purpose interaction recognizer for open-world video.","feed_headline":"Open-vocabulary video model recognizes human interactions it never saw","feed_subtitle":"It feeds person and background streams to a large language model, naming known and novel two-person actions.","key_machinery":"The machinery is a multi-branch cross-modal alignment built on Video-LLaMA. Track Anything first segments each video into a Person 1 stream, a Person 2 stream, and an optional background stream. Each stream is encoded by a frozen feature extractor—ViTPose for the person branches, a vanilla ViT for the background—then passed through a video Q-Former (borrowed from BLIP-2) and linear projection layers that map the video embeddings into the same space as text embeddings. The three projected embeddings are concatenated and fed into the frozen LLaMA 2 13B chat model, which generates open-vocabulary captions. Training uses cross-entropy between the model's embedding and the GPT-4-generated soft labels, updating only the Q-Formers and the linear layers; all encoders and the LLM stay frozen.","core_discovery":"The paper's central claim is that splitting an interaction video into three branches—Person 1, Person 2, and background—and feeding the separately encoded streams into a frozen LLaMA 2 13B chat model unlocks open-vocabulary interaction recognition. On the HHIRChat benchmark, the model obtains a mean cosine similarity of $0.628 \\pm 0.019$ between its generated captions and ground-truth descriptions, against $0.414 \\pm 0.033$ for VideoLLaMA-2 LoRA, and it outperforms or matches the fixed-vocabulary baselines on most of the ten constituent datasets. When evaluated through a prompt-engineered classifier on six interaction classes that were deliberately excluded from training, it reaches a macro-F1 of $0.315 \\pm 0.023$; the paper notes that any fixed-vocabulary classifier would necessarily score zero on such classes. The qualitative examples further indicate that the model's captions track the actual interaction more closely than vanilla VideoLLaMA-2 or its LoRA adaptation.","pith_inferences":["Inference: A sharper test of the open-vocabulary claim would be a zero-shot evaluation where the classifier is trained only on a randomly chosen subset of the 103 classes and tested on the rest, rather than on six hand-picked YouTube clips; this would give a controlled measure of how recognition degrades as the label gap widens.","Inference: Because the GPT-4 soft labels are not human-validated, a useful extension is a human-annotation study on a random sample of HHIRChat to measure label noise; if label noise is high, some of the reported similarity gain may reflect the labels rather than the architecture.","Inference: The three-branch design suggests a direct route to group interactions (three or more people) by adding one branch per tracked individual; the paper lists this as future work but the mechanism is already modular.","Inference: The model's natural-language output could be consumed by downstream surveillance queries (e.g., 'flag any video where one person strikes another'), which the paper motivates but does not demonstrate."],"forward_implications":["A fixed-vocabulary classifier cannot assign any label to an interaction it has not been trained on, whereas the OV-HHIR classifier reports a macro-F1 of 0.315 on six such unseen classes.","Because the model outputs free-form text, it can be used to retrieve or filter surveillance footage with natural-language queries instead of a fixed label list.","The unified HHIRChat dataset (86,623 sequences, 103 categories) gives the field a single common benchmark, making results from different labs comparable.","The per-person branches reduce interference in crowded scenes, since the model sees each individual as a separate stream rather than one fused frame.","The reported cosine similarity of 0.628 over caption embeddings suggests the generated descriptions are semantically close to ground truth, not just label-matched."],"supporting_citations":[{"why":"The architecture OV-HHIR builds on, providing the video-LLM base that the three-branch design modifies.","marker":"[9]"},{"why":"Supplies the video Q-Former cross-modal alignment used in each branch to project video features toward text embeddings.","marker":"[3]"},{"why":"The frozen LLaMA 2 13B chat model that generates the open-vocabulary interaction descriptions.","marker":"[21]"},{"why":"Produces the soft, descriptive labels for the 103 classes that constitute the training and evaluation signal.","marker":"[22]"},{"why":"Segments each video into Person 1, Person 2, and background streams required by the multi-branch design.","marker":"[23]"},{"why":"Frozen ViTPose encoder used in the two person branches to capture pose-specific features.","marker":"[35]"},{"why":"The main state-of-the-art baseline (VideoLLaMA 2) that OV-HHIR is compared against in Table II.","marker":"[37]"},{"why":"SportsHHI is the largest source dataset in HHIRChat, contributing most of the 86,623 training sequences.","marker":"[25]"},{"why":"NTU RGB+D 120 supplies the mutual-action samples and is one of the ten datasets folded into HHIRChat.","marker":"[33]"},{"why":"LoRA is the adaptation method used for the VideoLLaMA baselines, providing the 0.414 comparison point.","marker":"[38]"}],"fun_headline_variants":["Open-vocabulary model names unseen human interactions from video","LLM-powered system recognizes new human interactions it never saw","Video model uses large language model to describe unseen interactions","Cross-modal LLM spots novel two-person actions in surveillance video","Open-vocabulary interaction recognition beats fixed-label baselines"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 103-class unified label set produced by GPT-4 is accurate and consistent across the ten source datasets, and that the reported scores are computed on a held-out split; the paper does not document the label-merging process, validate the generated descriptions, or state whether the cosine-similarity numbers come from training or test videos.","fun_headline_variants_meta":{"raw":{"variants":["Open-vocabulary model names unseen human interactions from video","LLM-powered system recognizes new human interactions it never saw","Video model uses large language model to describe unseen interactions","Cross-modal LLM spots novel two-person actions in surveillance video","Open-vocabulary interaction recognition beats fixed-label baselines"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00047,"raw_usage":{"total_tokens":2330,"prompt_tokens":928,"completion_tokens":1402,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":1321}},"tokens_in":544,"tokens_out":1402,"duration_ms":10235,"temperature":1.0,"reasoning_tokens":1321,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:51:03.092654+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the exact OV-HHIR pipeline on HHIRChat with an explicit, documented train/test split (for instance, 80/20 at the video level) and compute Table II's cosine similarity only on held-out videos; if the score drops to the level of VideoLLaMA-2 LoRA or below, the advertised open-vocabulary advantage is not demonstrated. Independently, have a small panel of human annotators rate whether the GPT-4 soft labels match the video content; a high mismatch rate would indicate that the training signal itself is unreliable.","supporting_citations":[{"cited_title":"Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,","cited_arxiv_id":null,"evidence_quote":"Supplies the video Q-Former cross-modal alignment used in each branch to project video features toward text embeddings."},{"cited_title":"Vitpose: Sim- ple vision transformer baselines for human pose estimation,","cited_arxiv_id":null,"evidence_quote":"Frozen ViTPose encoder used in the two person branches to capture pose-specific features."},{"cited_title":"Sportshhi: A dataset for human-human interaction detection in sports videos,","cited_arxiv_id":null,"evidence_quote":"SportsHHI is the largest source dataset in HHIRChat, contributing most of the 86,623 training sequences."},{"cited_title":"Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,","cited_arxiv_id":null,"evidence_quote":"NTU RGB+D 120 supplies the mutual-action samples and is one of the ten datasets folded into HHIRChat."}],"review_version":1}