REVIEW 3 major objections 5 minor 63 references
Step-level visual grounding faithfulness predicts out-of-distribution generalization in long-horizon vision-language models.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 02:33 UTC pith:6RUQJ2X3
load-bearing objection A genuinely new evaluation axis with strong internal controls, but the headline r=0.83 rests on pooled non-independent samples; needs cluster-robust stats before the 'behavioral law' claim holds. the 3 major comments →
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The discovery is a behavioral law: models that maintain temporally grounded beliefs generalize better. The authors formalize behavioral faithfulness over long horizons with four diagnostics—SGR, temporal consistency, hallucination rate, and visual reliance score—and show that in-distribution SGR strongly predicts out-of-distribution retention (r=0.83, permutation p=0.003). The relationship replicates on all three benchmarks (STAR, R2R, TEACh), survives controlling for accuracy (partial r=0.68), and holds within a parameter-matched 7B cluster (r=0.78, N=15), establishing grounding as a capability dimension independent of scale and in-distribution accuracy.
What carries the argument
The central object is the Step Grounding Rate (SGR), the share of intermediate reasoning claims verified as supported by the visual evidence at that time, computed by parsing reasoning traces into claims, aligning temporal references to frames, and verifying via object detection, tracking, and action recognition. The framework treats reasoning traces as behavioral artifacts rather than trusted explanations, and accompanies SGR with three other diagnostics: Temporal Consistency Score (TCS) for belief maintenance or justified revision, Hallucination Rate (HR) for steps containing any unsupported claim, and Visual Reliance Score (VRS) for sensitivity to task-relevant versus irrelevant perturbat
Load-bearing premise
The correlation of 0.83 is computed across 24 data points formed by eight models on three benchmarks, and the permutation test treats these as exchangeable; if the true independent sample is the eight model families, the reported significance and confidence intervals are likely overconfident.
What would settle it
Run the same SGR-to-OOD analysis on a larger set of, say, 20 or more distinct model families (not just more benchmarks on the same 8 families), and compute the correlation with clustering or bootstrap at the model-family level. If the cluster-robust correlation drops substantially below 0.83 or its confidence interval includes zero, the claimed behavioral law does not generalize.
If this is right
- Standard accuracy-only benchmarks should be supplemented with step-level grounding metrics, since final accuracy can mask large grounding failures.
- SGR measured in-distribution could serve as a cheap leading indicator of how much accuracy a model will lose under distribution shift.
- The within-capacity spread in SGR suggests that model selection among similarly sized models should consider grounding quality, not just task accuracy.
- If the law holds, training or fine-tuning methods that explicitly improve step-level grounding may also improve out-of-distribution robustness.
- The relationship extends beyond video QA to embodied navigation and instruction following, suggesting a general principle for long-horizon visual tasks.
Where Pith is reading between the lines
- The headline correlation is computed across 24 model-benchmark points, but the 8 model families may be the true independent sample; a re-analysis at the model-family level would likely yield wider confidence intervals and is a sharper test of the claimed law.
- The causal interpretation (grounding causes robustness) is plausible but not directly proven; a training intervention that improves SGR and then measures OOD generalization would test it directly.
- If grounding degrades with task progress, then the gap between short and long tasks should widen for low-SGR models, an observable prediction that could be tested on tasks of varying length.
- SGR could be adopted as a practical diagnostic in model development pipelines, for instance to detect when a model is relying on language priors or dataset statistics rather than visual evidence.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a family of step-level faithfulness metrics—Step Grounding Rate (SGR), Temporal Consistency Score (TCS), Hallucination Rate (HR), and Visual Reliance Score (VRS)—for long-horizon vision-language models. It reports that SGR measured in-distribution predicts out-of-distribution retention across eight models and three benchmarks (r=0.83, N=24, permutation p=0.003), that this relationship holds within a parameter-matched 7B cluster, and that grounding quality is an independent axis of capability not reducible to scale or in-distribution accuracy. The paper includes multiple controls: counterfactual trace perturbations, cross-architecture verifier agreement, a random-reasoning floor, and a non-disclosure SGR-lite variant.
Significance. If the central statistical claim survives appropriate cluster-robust inference, the result is significant: it offers a mechanistic, step-level indicator of OOD robustness that standard accuracy evaluation misses, with practical implications for VLM evaluation and model selection. The paper's strengths are its concrete operationalization (Eqs. 1–4), the breadth of internal controls (counterfactual traces, verifier diversity, SGR-lite, perturbation analyses), and the transparent reporting of model-level tables. The main concern is inferential: the headline p-value treats repeated model-benchmark measurements as independent, and the paper must demonstrate that the correlation is not an artifact of benchmark-level differences before the 'behavioral law' framing is justified.
major comments (3)
- [§5.2, Table 2, Finding 4] The headline correlation pools 8 models × 3 benchmarks (N=24) and the permutation test/95% CI treat these 24 points as exchangeable. However, both SGR and OOD retention vary systematically by benchmark in Tables 1–2 (for every model, TEACh is lowest), so the pooled estimate can be inflated by benchmark-level differences. The paper reports per-benchmark r values (0.81–0.85) but not their p-values or CIs; with N=8 these are the more appropriate inferential units. Please add a model-level analysis (e.g., average SGR and OOD retention per model, N=8), a cluster-bootstrap or permutation-by-model test, or a mixed model with random intercepts for model and benchmark. The model-level averages from the tables appear strongly monotone, so the effect may survive, but as written p=0.003 is not calibrated.
- [§5.2, Finding 5(iii)] The partial correlation r_partial=0.68 (p<0.05) is reported without specifying the covariate set, the sample size, or whether benchmark is included as a fixed effect. If this is a partial correlation on the pooled N=24 controlling only for in-distribution accuracy, it inherits the clustering problem of Finding 4. Please report the partial correlation within the 7B cluster or with benchmark and model as random effects, and state the effective degrees of freedom.
- [§7.4 Limitations] The Limitations section correctly acknowledges that the number of distinct model families is small (8), but the main result is still presented with a permutation p-value computed over 24 non-independent model-benchmark points. This is internally inconsistent. The manuscript should either present cluster-robust inference as the primary result or explicitly frame the N=24 correlation as a descriptive pooled statistic, with the per-benchmark correlations (and their p-values) carrying the inferential weight.
minor comments (5)
- [§5.1, Finding 1] The claim that 'task accuracy consistently exceeds visual grounding (SGR)' is not supported by Table 1 for TEACh: for 7 of 8 models, SGR exceeds the task success rate (e.g., GPT-4o: SR=59.2 vs SGR=63.7). Qualify this finding by benchmark or restrict it to STAR/R2R.
- [§3.4, Eq. (1) example] The worked example for g(r_i) is unclear: a step mentioning a red chair, a blue table, and a green bulb is said to score 0.67 even if 'correctly identified in the visual input'; if all three are verified the score should be 1.0. Please rewrite the example to specify which claims are and are not supported, and change 'Concurrently' to 'In contrast'.
- [Abstract vs §7.3] The counterfactual SGR drop is reported as 26–41 percentage points in the abstract and 26–39 pp in §7.3. Reconcile the numbers.
- [Throughout] Several references to 'supplementary Section??' and 'supplementary Table??' are unfilled placeholders. Fill these before submission.
- [Table 2] The column headers are garbled (e.g., 'Acc ∆ SR ∆ SR ∆'). Since R2R and TEACh use success rate rather than accuracy, label the columns unambiguously for each benchmark.
Circularity Check
No circularity: SGR is independently measured and correlated with OOD retention, with no parameter fitted to the outcome.
full rationale
No load-bearing circular step was found. The central claim is an empirical correlation between Step Grounding Rate (SGR), defined in Eq. 1 as the fraction of verified visual claims in reasoning traces, and OOD retention, defined as accuracy on held-out splits (STAR-OOD, R2R-OOD, TEACh-OOD). SGR is computed from model outputs, an external pipeline (spaCy, Faster R-CNN, DeepSORT, SlowFast), and thresholds calibrated on human annotations (e.g., tau=0.8), not on OOD accuracy. The paper explicitly tests the alternative explanations that SGR could reflect grammatical coherence, verbosity, or prompt compliance (Sec. 7.2, 7.3), and these controls are independent of the OOD outcome. No parameter in the SGR definition is fit to the OOD result; hyperparameters such as epsilon, clip, and the similarity threshold are either fixed by human labels or shown via sensitivity analysis not to affect the correlation. The paper uses no self-citations as load-bearing premises and imports no uniqueness theorem or ansatz from its own prior work. The Limitations (Sec. 7.4) acknowledge the small number of model families, which is a statistical-power and independence concern rather than a circularity concern. Therefore the derivation is self-contained: SGR is measured in-distribution, OOD retention is measured separately, and the reported r=0.83 is a genuine empirical association.
Axiom & Free-Parameter Ledger
free parameters (3)
- TCS belief-transition threshold τ =
0.8
- VRS smoothing constant ε =
0.01
- VRS upper clip =
10
axioms (4)
- domain assumption OOD splits are genuinely out-of-distribution and difficulty-matched
- domain assumption The automatic verifier (spaCy, Faster R-CNN, DeepSORT, SlowFast) gives sufficiently unbiased support labels
- domain assumption The 24 measurements are treated as independent in the main permutation test
- domain assumption CoT traces elicited by the structured prompt are a valid view of the model's reasoning
read the original abstract
We uncover a behavioral law of long-horizon vision-language models: models that maintain temporally grounded beliefs generalize better. Standard benchmarks measure only final-answer accuracy, which obscures how models use visual information; a model can guess correctly while its step-by-step reasoning is entirely unanchored to the visual input. We formalize this as behavioral faithfulness over long horizons, an empirically measurable property that quantifies whether a model's intermediate reasoning remains consistent with the evolving visual state. Across eight models on three long-horizon benchmarks, we demonstrate that temporal grounding quality is a leading indicator of robustness: the Step Grounding Rate (SGR) predicts out-of-distribution retention with $r = 0.83$ (permutation test $p = 0.003$), a relationship that holds within capacity-matched models and cannot be explained by scale or in-distribution accuracy. Critically, grounding quality varies by up to 10.8 percentage points within parameter-matched 7B models despite similar accuracy, revealing it as an independent axis of model capability. Multiple robustness checks confirm the signal reflects genuine visual reliance: counterfactual traces drop SGR by 26--41 percentage points, cross-architecture verifiers agree at $\rho = 0.96$, random reasoning scores near chance ($\sim 18\%$), and the predictor remains strong even without explicit reasoning disclosure ($r = 0.78$).
Figures
Reference graph
Works this paper leans on
-
[1]
In: IEEE Conf
Agrawal, A., Batra, D., Parikh, D., Kembhavi, A.: Don’t just assume; look and answer: Overcoming priors for visual question answering. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 4971–4980 (2018) 2
2018
-
[2]
In: IEEE Conf
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L.: Bottom-up and top-down attention for image captioning and visual question answering. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 6077–6086 (2018) 3
2018
-
[3]
In: IEEE Conf
Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., Van Den Hengel, A.: Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 3674–3683 (2018) 1, 2, 7
2018
-
[4]
arXiv preprint arXiv:2006.13171 (2020) 2
Batra, D., Gokaslan, A., Kembhavi, A., Maksymets, O., Mottaghi, R., Savva, M., Toshev, A., Wijmans, E.: Objectnav revisited: On evaluation of embodied agents navigating to objects. arXiv preprint arXiv:2006.13171 (2020) 2
Pith/arXiv arXiv 2006
-
[5]
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End- to-end object detection with transformers. In: Eur. Conf. Comput. Vis. pp. 213–229 (2020) 4
2020
-
[6]
In: IEEE Winter Conf
Chattopadhay, A., Sarkar, A., Howlader, P., Balasubramanian, V.N.: Grad- cam++: Generalized gradient-based visual explanations for deep convolutional networks. In: IEEE Winter Conf. Appl. Comput. Vis. pp. 839–847 (2018) 3
2018
-
[7]
In: IEEE Conf
Chen, L., Yan, X., Xiao, J., Zhang, H., Pu, S., Zhuang, Y.: Counterfactual samples synthesizing for robust visual question answering. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 10800–10809 (2020) 2
2020
-
[8]
arXiv preprint arXiv:2312.14238 (2024) 7
Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., et al.: InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. arXiv preprint arXiv:2312.14238 (2024) 7
Pith/arXiv arXiv 2024
-
[9]
Chen, Z., Li, S., Li, P., Gao, C., Shan, Y., Zheng, L.: How far are we from intelligent video understanding? arXiv preprint arXiv:2406.07689 (2024) 3
Pith/arXiv arXiv 2024
-
[10]
In: IEEE Conf
Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., Batra, D.: Embodied question answering. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1–10 (2018) 1, 2
2018
-
[11]
arXiv preprint arXiv:2406.14515 (2024) 3
Fang, X., Mao, K., Duan, H., Zhao, X., Zhang, Y., Lin, D., Chen, K.: MMBench- Video: A long-form multi-shot benchmark for holistic video understanding. arXiv preprint arXiv:2406.14515 (2024) 3
Pith/arXiv arXiv 2024
-
[12]
Feichtenhofer, C., Fan, H., Malik, J., He, K.: Slowfast networks for video recogni- tion. In: Int. Conf. Comput. Vis. pp. 6202–6211 (2019) 4, 7
2019
-
[13]
In: IEEE Conf
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D.: Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 6904–6913 (2017) 2
2017
-
[14]
Hendricks, L.A., Akata, Z., Rohrbach, M., Donahue, J., Schiele, B., Darrell, T.: Generating visual explanations. In: Eur. Conf. Comput. Vis. pp. 3–19 (2016) 3
2016
-
[15]
In: IEEE Conf
Hudson, D.A., Manning, C.D.: Gqa: A new dataset for real-world visual reason- ing and compositional question answering. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 6700–6709 (2019) 2
2019
-
[16]
Jacovi, A., Goldberg, Y.: Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness? In: Proc. Annu. Meet. Assoc. Comput. Lin- guist. pp. 4198–4205 (2020) 3
2020
-
[17]
In: IEEE Conf
Jang, Y., Song, Y., Yu, Y., Kim, Y., Kim, G.: Tgif-qa: Toward spatio-temporal rea- soning in visual question answering. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 2758–2766 (2017) 2 16 M. A. Rahman et al
2017
-
[18]
In: IEEE Conf
Johnson, J., Hariharan, B., Van Der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., Girshick, R.: Clevr: A diagnostic dataset for compositional language and elemen- tary visual reasoning. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 2901–2910 (2017) 2
2017
-
[19]
Frontiers in Artificial Intelligence2, 28 (2019) 2
Kafle, K., Kanan, C.: Challenges and prospects in vision and language research. Frontiers in Artificial Intelligence2, 28 (2019) 2
2019
-
[20]
Kaushik, D., Hovy, E., Lipton, Z.: Learning the difference that makes a difference with counterfactually-augmented data. In: Int. Conf. Learn. Represent. (2020) 2
2020
-
[21]
In: IJCAI
Kim, K.M., Heo, M.O., Choi, S.H., Zhang, B.T.: Deepstory: Video story qa by deep embedded memory networks. In: IJCAI. pp. 2016–2022 (2017) 2
2016
-
[22]
Kojima, T., Gu, S.S., Reid, M., Matsuo, Y., Iwasawa, Y.: Large language models are zero-shot reasoners. Adv. Neural Inform. Process. Syst.35, 22199–22213 (2022) 3
2022
-
[23]
Krantz, J., Wijmans, E., Majumdar, A., Batra, D., Lee, S.: Beyond the nav-graph: Vision-and-language navigation in continuous environments. In: Eur. Conf. Com- put. Vis. pp. 104–120 (2020) 2
2020
-
[24]
arXiv preprint arXiv:2307.13702 (2023) 3
Lanham, T., Chen, A., Radhakrishnan, A., Steiner, B., Denison, C., Hernandez, D., Li, D., Durmus, E., Hubinger, E., Kernion, J., et al.: Measuring faithfulness in chain-of-thought reasoning. arXiv preprint arXiv:2307.13702 (2023) 3
Pith/arXiv arXiv 2023
-
[25]
In: Proc
Lei, J., Berg, T.L., Bansal, M.: Revealing single frame bias for video-and-language learning. In: Proc. Annu. Meet. Assoc. Comput. Linguist. pp. 1029–1041 (2022) 2
2022
-
[26]
In: IEEE Conf
Lei, J., Li, L., Zhou, L., Gan, Z., Berg, T.L., Bansal, M., Liu, J.: Less is more: Clip- bert for video-and-language learning via sparse sampling. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 7331–7341 (2021) 1, 2
2021
-
[27]
arXiv preprint arXiv:2305.06355 (2023) 7
Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., Qiao, Y.: Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355 (2023) 7
Pith/arXiv arXiv 2023
-
[28]
arXiv preprint arXiv:2311.17005 (2023) 3
Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Liu, Y., Wang, Z., Xu, J., Chen, G., Luo, P., et al.: MVBench: A comprehensive multi-modal video understanding benchmark. arXiv preprint arXiv:2311.17005 (2023) 3
Pith/arXiv arXiv 2023
-
[29]
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W.X., Wen, J.R.: Evaluating object hal- lucination in large vision-language models. Proc. Conf. Empirical Methods Natural Language Process. (2023) 3
2023
-
[30]
arXiv preprint arXiv:2305.20050 (2023) 3
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., Cobbe, K.: Let’s verify step by step. arXiv preprint arXiv:2305.20050 (2023) 3
Pith/arXiv arXiv 2023
-
[31]
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., Lee, Y.J.: LLaVA-NeXT: Im- proved reasoning, ocr, and world knowledge.https://llava-vl.github.io/blog/ 2024-01-30-llava-next/(2024) 7
2024
-
[32]
arXiv preprint arXiv:2310.14566 (2023) 3
Liu, T., Guan, T., Zhu, Q., Zhang, X., et al.: Hallusionbench: An advanced di- agnostic suite for entangled language hallucination and visual illusion in large vision-language models. arXiv preprint arXiv:2310.14566 (2023) 3
Pith/arXiv arXiv 2023
-
[33]
(2021) 2
Min, S.Y., Chaplot, D.S., Ravikumar, P., Bisk, Y., Salakhutdinov, R.: Film: Follow- inginstructionsinlanguagewithmodularmethods.In:Int.Conf.Learn.Represent. (2021) 2
2021
-
[34]
OpenAI: Gpt-4v(ision) system card.https://openai.com/research/gpt- 4v- system-card(2023) 7
2023
-
[35]
OpenAI: Hello GPT-4o.https://openai.com/index/hello-gpt-4o/(2024) 7
2024
-
[36]
In: AAAI
Padmakumar, A., Thomason, J., Shrivastava, A., Lange, P., Narayan-Chen, A., Piramuthu, S., Yin, D., Ramanathan, D.H.T.: Teach: Task-driven embodied agents that chat. In: AAAI. pp. 2017–2025 (2022) 7 Step-Level Visual Grounding Faithfulness in Long-Horizon VLMs 17
2017
-
[37]
In: IEEE Conf
Park, D.H., Hendricks, L.A., Akata, Z., Rohrbach, A., Schiele, B., Darrell, T., Rohrbach, M.: Multimodal explanations: Justifying decisions and pointing to the evidence. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 8779–8788 (2018) 3
2018
-
[38]
Plummer, B.A., Wang, L., Cervantes, C.M., Caicedo, J.C., Hockenmaier, J., Lazeb- nik, S.: Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In: Int. Conf. Comput. Vis. pp. 2641–2649 (2015) 3
2015
-
[39]
Ramakrishnan, S., Agrawal, A., Lee, S.: Overcoming language priors in visual ques- tion answering with adversarial regularization. In: Adv. Neural Inform. Process. Syst. pp. 1541–1551 (2018) 2
2018
-
[40]
In: Proc
Reimers, N., Gurevych, I.: Sentence-bert: Sentence embeddings using siamese bert- networks. In: Proc. Conf. Empirical Methods Natural Language Process. pp. 3982– 3992 (2019) 5
2019
-
[41]
Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object de- tection with region proposal networks. In: Adv. Neural Inform. Process. Syst. pp. 91–99 (2015) 4, 7
2015
-
[42]
In: Proc
Ribeiro,M.T.,Wu,T.,Guestrin,C.,Singh,S.:Beyondaccuracy:Behavioraltesting of nlp models with checklist. In: Proc. Annu. Meet. Assoc. Comput. Linguist. pp. 4902–4912 (2020) 2
2020
-
[43]
Rohrbach, A., Rohrbach, M., Hu, R., Darrell, T., Schiele, B.: Grounding of textual phrases in images by reconstruction. In: Eur. Conf. Comput. Vis. pp. 817–834 (2016) 3
2016
-
[44]
Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., et al.: Habitat: A platform for embodied ai research. In: Int. Conf. Comput. Vis. pp. 9339–9347 (2019) 2
2019
-
[45]
Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., Batra, D.: Grad- cam: Visual explanations from deep networks via gradient-based localization. In: Int. Conf. Comput. Vis. pp. 618–626 (2017) 2, 3, 8
2017
-
[46]
Shen, S., Li, L., Tan, H., Bansal, M., Rohrbach, A., Chang, K.W., Yao, Z., Keutzer, K.: How much can clip benefit vision-and-language tasks? In: Int. Conf. Learn. Represent. (2022) 7
2022
-
[47]
In: IEEE Conf
Shridhar, M., Thomason, J., Gordon, D., Bisk, Y., Han, W., Mottaghi, R., Zettle- moyer, L., Fox, D.: Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 10740–10749 (2020) 1, 2
2020
-
[48]
In: Proc
Suhr, A., Yan, C., Schluger, J., Yu, S., Khader, H., Mouallem, M., Zhang, I., Artzi, Y.: Situated mapping of sequential instructions to actions with single-step reward observation. In: Proc. Annu. Meet. Assoc. Comput. Linguist. pp. 2072–2082 (2019) 2
2072
-
[49]
In: IEEE Conf
Tapaswi, M., Zhu, Y., Stiefelhagen, R., Torralba, A., Urtasun, R., Fidler, S.: Movieqa: Understanding stories in movies through question-answering. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 4631–4640 (2016) 2
2016
-
[50]
In: Conf
Thomason, J., Murray, M., Cakmak, M., Zettlemoyer, L.: Vision-and-dialog navi- gation. In: Conf. Robot Learn. pp. 394–406 (2020) 2
2020
-
[51]
Turpin, M., Michael, J., Perez, E., Bowman, S.: Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. Adv. Neural Inform. Process. Syst. (2023) 3
2023
-
[52]
arXiv preprint arXiv:2312.17225 (2024) 3 18 M
Wang, K., He, Y., Wang, Y., Li, Y., Li, K., Qian, R., Luo, P., Wang, Y., Qiao, Y.: VideoInstruct: Towards detailed video understanding via instruction tuning. arXiv preprint arXiv:2312.17225 (2024) 3 18 M. A. Rahman et al
Pith/arXiv arXiv 2024
-
[53]
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V., Zhou, D., et al.: Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inform. Process. Syst.35, 24824–24837 (2022) 3, 4
2022
-
[54]
Wiegreffe, S., Marasović, A., Smith, N.A.: Measuring association between labels andfree-textrationales.Proc.Conf.EmpiricalMethodsNaturalLanguageProcess. pp. 10266–10284 (2021) 3
2021
-
[55]
In: IEEE Int
Wojke, N., Bewley, A., Paulus, D.: Simple online and realtime tracking with a deep association metric. In: IEEE Int. Conf. Image Process. pp. 3645–3649 (2017) 4, 7
2017
-
[56]
Wu,B.,Yu,S.,Chen,Z.,Tenenbaum,J.B.,Gan,C.:Star:Abenchmarkforsituated reasoning in real-world videos. In: Adv. Neural Inform. Process. Syst. (2021) 7
2021
-
[57]
arXiv preprint arXiv:2304.08958 (2023) 2
Wu, J., Li, Y., Chen, A., Liu, Z.: REvAL: A benchmark for robust evaluation of vision-language models. arXiv preprint arXiv:2304.08958 (2023) 2
Pith/arXiv arXiv 2023
-
[58]
In: IEEE Conf
Wu, Z., Wang, Y., Lu, J., Zhou, J.: Visual question answering with attention-based explanation. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 4358–4367 (2021) 3
2021
-
[59]
In: ACM Int
Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., Zhuang, Y.: Video question answering via gradually refined attention over appearance and motion. In: ACM Int. Conf. Multimedia. pp. 1645–1653 (2017) 2
2017
-
[60]
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., Bengio, Y.: Show, attend and tell: Neural image caption generation with visual attention. In: Int. Conf. Mach. Learn. pp. 2048–2057 (2015) 2, 3, 8
2048
-
[61]
Yi, K., Gan, C., Li, Y., Kohli, P., Wu, J., Torralba, A., Tenenbaum, J.B.: Clevrer: Collision events for video representation and reasoning. In: Int. Conf. Learn. Rep- resent. (2020) 2
2020
-
[62]
arXiv preprint arXiv:2311.10122 (2023) 7
Zhang, B., Lin, B., Yang, X., Zhang, S., Liu, X., Wang, J.: Video-llava: Learn- ing united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122 (2023) 7
Pith/arXiv arXiv 2023
-
[63]
arXiv preprint arXiv:2306.02858 (2023) 7
Zhang, H., Li, X., Bing, L.: Video-llama: An instruction-tuned audio-visual lan- guage model for video understanding. arXiv preprint arXiv:2306.02858 (2023) 7
Pith/arXiv arXiv 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.