Pith. sign in

REVIEW 2 major objections 3 minor 39 references

Contact-Rich and Deformable Foot Modeling for Locomotion Control of the Human Musculoskeletal System

T0 review · 2 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Replacing the rigid foot with a deformable, contact-rich model inside a full musculoskeletal simulation yields more human-like walking kinematics, kinetics, and stability, and reproduces measured human gait.

desk verdict The abstract and body are two different papers; the foot model is never presented, so the submission cannot be evaluated as is. read the letter →

arxiv 2508.11885 v1 pith:WGLXVW6E submitted 2025-08-16 cs.RO

classification cs.RO
keywords musculoskeletalsimulationdeformablefootmodelfoot-groundcontactlocomotioncontrolpolicytraininggaitanalysishumanoidroboticsbiomechanics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that oversimplified rigid foot-ground contact is a key limit in musculoskeletal walking simulation, and that a deformable, contact-rich foot model fixes it. The authors embed this foot model in a complete musculoskeletal system and use a two-stage policy training strategy to handle the resulting multi-point contacts and tissue deformation. They report improvements over rigid musculoskeletal models in kinematic, kinetic, and gait-stability metrics, and say their simulation closely reproduces real-world biomechanical measurements of walking. If this holds, biomechanical simulation and humanoid robotics gain a more faithful account of how feet actually interact with the ground.

What carries the argument

The load-bearing object is the deformable foot model: a contact-rich representation of the foot that computes multi-point, deformable interaction with the ground, integrated into a complete musculoskeletal body. Around it, a two-stage policy training strategy makes control tractable by first learning a natural walking pattern and then refining it under the full multi-contact dynamics. The foot model carries the argument because it is the main difference between the proposed system and the rigid-baseline comparison, while the two-stage training is the mechanism that lets that difference be exploited.

What would settle it

Run the trained deformable-foot simulation through the same walking trials used for human motion capture and compare the predicted vertical ground-reaction-force profile, center-of-pressure path, and plantar pressure distribution against force-plate data. If the deformable-foot model fails to beat a rigid-foot model on those measurements by a clear margin, or if its values fall outside the spread of the human subjects, the central claim is refuted.

Watch

Extended reading notes

Core claim

The central claim is that foot-ground interaction must be modeled as deformable and contact-rich rather than as a few rigid contact points in order to reproduce human walking dynamics. The paper develops a deformable foot model integrated into a complete musculoskeletal system, and shows that this interface-enhanced model outperforms conventional rigid musculoskeletal models on kinematic, kinetic, and gait-stability measures. The authors further validate against human subject data, reporting that the simulation closely reproduced real biomechanical measurements. On the paper's own terms, the deformable interface is what closes the gap between simulated and measured gait.

Load-bearing premise

The deformable foot's tissue properties and contact equations must faithfully represent real biological foot-ground interaction, because everything about the learned gait and the claimed match to human data depends on that realism.

Editorial extensions

If this is right

  • Simulations using the deformable foot outperform rigid-foot models on kinematic, kinetic, and gait-stability metrics during walking.
  • A two-stage policy training strategy is sufficient to control a full musculoskeletal model with multi-point deformable contacts, removing a control bottleneck.
  • The simulation's walking output closely tracks human biomechanical measurements, supporting its use as a surrogate for gait experiments.
  • The same foot-ground interaction modeling and training framework can be extended to humanoid robots that need precise foot-ground control.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves open whether the gains come from tissue deformation per se or from the richer multi-point contact geometry; a variant with a rigid foot but many contact points could separate the two.
  • If tissue parameters are varied, the same model could predict gait changes in conditions such as flatfoot or aging, an extension the paper does not test.
  • For humanoid robotics, the practical promise is a sim-to-real transfer path: a policy trained on this contact-rich foot may transfer more reliably to hardware because the ground-reaction feedback is more realistic.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The manuscript arXiv:2508.11885 is submitted under the title "Contact-Rich and Deformable Foot Modeling for Locomotion Control of the Human Musculoskeletal System," and its abstract claims a novel deformable foot model integrated into a musculoskeletal system, a two-stage policy training strategy, improvements over rigid musculoskeletal models in kinematic, kinetic, and gait stability metrics, and validation against human walking data. The submitted full text, however, is an entirely different paper: "EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models," with a different author list, its own abstract, method, experiments, references, and appendices. The body contains no foot model, no musculoskeletal simulation, no foot-ground contact mechanics, no deformable tissue parameters, no control policy, no gait stability analysis, and no human-subject validation. The only link between the title and the full text is the arXiv header, which displays identifier 2508.11886v1 rather than 2508.11885. Because every load-bearing component of the claimed contribution is absent from the manuscript, the central claim cannot be evaluated on the submitted content.

Significance. If substantiated, a contact-rich deformable foot model integrated with a musculoskeletal system and validated against human gait data would be a useful contribution to biomechanics simulation and humanoid locomotion control. However, this manuscript provides no evidence toward that contribution: the full text is a visual-token-pruning paper whose methods, equations, experiments, and references are unrelated to locomotion. The EVTP-IVS portion appears to contain an internally coherent empirical study, including coverage-based pruning experiments and speedup measurements, but that work cannot be credited toward the foot-model claims made in the abstract. There are no machine-checked proofs, reproducible code, parameter-free derivations, or falsifiable predictions relevant to the abstract's central claim within the submitted content.

major comments (2)
  1. [Full text, all sections] The central claim of the abstract—that a novel contact-rich and deformable foot model integrated within a musculoskeletal system improves kinematic, kinetic, and gait stability metrics over rigid models and closely reproduces human walking measurements—is entirely unsupported by the submitted manuscript. The full text is EVTP-IVS, a paper on visual token pruning for multimodal large language models. It contains no musculoskeletal model, no deformable foot geometry or tissue properties, no contact model, no two-stage policy training, no gait simulation, and no comparison against human subject data. This is not a debatable modeling assumption but the complete absence of the claimed work, so the manuscript cannot be evaluated scientifically for the stated contribution.
  2. [EVTP-IVS Sections 4–6 and Tables 1–4] The methods and experimental sections of the submitted text concern k-center token selection with spatial augmentation, FLOPs estimation, and instructed visual segmentation benchmarks on RefCOCO, ReasonSeg, ReVOS, and related datasets. Equations (1)–(6) define pruning objectives, not foot-ground contact mechanics, and Tables 1–4 report segmentation metrics, not gait kinematics or kinetics. These contents cannot serve as the derivation, simulation setup, or validation for the abstract's claims about locomotion control, so the abstract's assertions about comparative gait improvements and human-subject validation have no evidentiary basis in this manuscript.
minor comments (3)
  1. [Header/metadata] The arXiv identifier printed in the full-text header is 2508.11886v1 [cs.CV] dated 16 Aug 2025, whereas the submission identifier is 2508.11885 (cs.RO); this mismatch should be reconciled if the correct manuscript is resubmitted.
  2. [References] The reference list contains no entries on biomechanics, musculoskeletal modeling, foot anatomy, contact simulation, or human gait; all cited works are about vision-language models and visual token pruning, which further confirms that the body text does not correspond to the abstract's topic.
  3. [EVTP-IVS Section 8] The concluding section of the submitted text states that the work presents "the first study on visual token pruning for IVS" and discusses inference acceleration; this is irreconcilable with the abstract's claim of a deformable foot model for locomotion control, and the mismatch should be corrected by submitting the intended paper.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity detectable: the submitted full text is an unrelated visual-token-pruning paper, so the claimed foot-model derivation chain has no content to audit.

full rationale

The abstract claims a 'novel contact-rich and deformable model of the human foot integrated within a complete musculoskeletal system,' a two-stage policy training strategy, and validation against human gait data. However, the submitted full text is a different paper: 'EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models,' arXiv:2508.11886v1 [cs.CV], by different authors. The body contains no musculoskeletal model, no foot-ground contact mechanics, no deformable tissue parameters, no control policy, no gait stability analysis, and no human-subject gait validation. There is therefore no derivation chain of the form 'X derives Y' in which one could exhibit a reduction of a prediction to a fitted input, a load-bearing self-citation, a uniqueness theorem imported from the authors' prior work, an ansatz smuggled in via citation, or a renamed empirical pattern. The enumerated circularity patterns cannot be substantiated with quoted equations because the relevant equations do not appear in the manuscript. Under the hard rule that circularity may only be claimed when the paper itself exhibits the specific reduction, no circularity finding is possible. The manuscript-content mismatch is a serious integrity and correctness problem, but it is not a circularity problem: a missing derivation is not the same as a derivation that reduces to its own inputs. If the intended foot-model paper existed, the abstract's promised external validation against human subject data would, if present, be independent non-circular evidence; but that content is absent here. Accordingly, the appropriate circularity score is 0, with the caveat that this score reflects absence of a circular chain rather than scientific validity of the claimed foot-model contribution.

Assumptions & free parameters 0 free parameters · 3 assumptions · 1 invented entities

Only the abstract is available, so the ledger records assumptions stated there. No free parameters can be identified from the abstract alone.

assumptions (3)
  • domain assumption A deformable foot model captures complex biomechanical interactions during walking.
    The abstract asserts this without providing a physical derivation or empirical validation.
  • ad hoc to paper Two-stage policy training can learn natural walking patterns for multi-point contact and deformable models.
    This is the proposed solution to the control challenge; its existence and sufficiency are assumed in the abstract.
  • domain assumption The musculoskeletal model is a complete representation of the human body for gait.
    The abstract states 'complete musculoskeletal system' but no details are given.
invented entities (1)
  • Deformable foot model
    purpose: Modeling foot-ground contact to improve gait simulation
    It is a new modeling construct claimed in the abstract, but the full text provides no falsifiable details or validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Contact-Rich and Deformable Foot Modeling for Locomotion Control of the Human Musculoskeletal System." pith.science (2026). https://pith.science/paper/WGLXVW6E

@misc{pith2026250811885,
  author       = {Pith},
  title        = {Pith review of: Contact-Rich and Deformable Foot Modeling for Locomotion Control of the Human Musculoskeletal System},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WGLXVW6E}},
  note         = {Machine review of arXiv:2508.11885}
}
read the original abstract

The human foot serves as the critical interface between the body and environment during locomotion. Existing musculoskeletal models typically oversimplify foot-ground contact mechanics, limiting their ability to accurately simulate human gait dynamics. We developed a novel contact-rich and deformable model of the human foot integrated within a complete musculoskeletal system that captures the complex biomechanical interactions during walking. To overcome the control challenges inherent in modeling multi-point contacts and deformable material, we developed a two-stage policy training strategy to learn natural walking patterns for this interface-enhanced model. Comparative analysis between our approach and conventional rigid musculoskeletal models demonstrated improvements in kinematic, kinetic, and gait stability metrics. Validation against human subject data confirmed that our simulation closely reproduced real-world biomechanical measurements. This work advances contact-rich interface modeling for human musculoskeletal systems and establishes a robust framework that can be extended to humanoid robotics applications requiring precise foot-ground interaction control.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

39 extracted references · 27 canonical work pages

  1. [1]

    Flamingo: a visual language model for few-shot learning.Advances in Neural Informa- tion Processing Systems, 35:23716–23736, 2022

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning.Advances in Neural Informa- tion Processing Systems, 35:23716–23736, 2022. 2

  2. [2]

    Deep variational information bottle- neck.arXiv preprint arXiv:1612.00410, 2016

    Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy. Deep variational information bottle- neck.arXiv preprint arXiv:1612.00410, 2016. 5

  3. [3]

    Divprune: Diversity-based visual token pruning for large multimodal models

    Saeed Ranjbar Alvar, Gursimran Singh, Mohammad Akbari, and Yong Zhang. Divprune: Diversity-based visual token pruning for large multimodal models. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 9392–9401, 2025. 2, 4, 6, 12

  4. [4]

    A closer look at referring expressions for video object segmentation.Multimedia Tools and Applications, 82 (3):4419–4438, 2023

    Miriam Bellver, Carles Ventura, Carina Silberer, Ioan- nis Kazakos, Jordi Torres, and Xavier Giro-i Nieto. A closer look at referring expressions for video object segmentation.Multimedia Tools and Applications, 82 (3):4419–4438, 2023. 2

  5. [5]

    Token merging: Your vit but faster.arXiv preprint arXiv:2210.09461, 2022

    Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. Token merging: Your vit but faster.arXiv preprint arXiv:2210.09461, 2022. 2, 6, 12

  6. [6]

    Matryoshka multimodal models

    Mu Cai, Jianwei Yang, Jianfeng Gao, and Yong Jae Lee. Matryoshka multimodal models. InWorkshop on Video-Language Models@ NeurIPS 2024, 2024. 2, 6, 12

  7. [7]

    An im- age is worth 1/2 tokens after layer 2: Plug-and-play in- ference acceleration for large vision-language models

    Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. An im- age is worth 1/2 tokens after layer 2: Plug-and-play in- ference acceleration for large vision-language models. InEuropean Conference on Computer Vision, pages 19–35. Springer, 2024. 1, 2, 6, 12

  8. [8]

    Sequence complementor: Complementing transformers for time series forecasting with learnable sequences

    Xiwen Chen, Peijie Qiu, Wenhui Zhu, Huayu Li, Hao Wang, Aristeidis Sotiras, Yalin Wang, and Abol- fazl Razi. Sequence complementor: Complementing transformers for time series forecasting with learnable sequences. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 15913– 15921, 2025. 5

Show all 39 references
  1. [9]

    Masked- attention mask transformer for universal image seg- mentation

    Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked- attention mask transformer for universal image seg- mentation. InProceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, pages 1290–1299, 2022. 2, 3, 11

  2. [10]

    Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model

    Ho Kei Cheng and Alexander G Schwing. Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model. InEuropean Conference on Computer Vision, pages 640–658. Springer, 2022. 2

  3. [11]

    John Wiley & Sons, 2nd edition,

    Thomas M Cover and Joy A Thomas.Elements of information theory. John Wiley & Sons, 2nd edition,

  4. [12]

    Phi-2: The surprising power of small language models.Microsoft Research Blog, 1:3, 2023

    Mojan Javaheripi, S ´ebastien Bubeck, Marah Abdin, Jyoti Aneja, Sebastien Bubeck, Caio C ´esar Teodoro Mendes, Weizhu Chen, Allie Del Giorno, Ronen El- dan, Sivakanth Gopi, et al. Phi-2: The surprising power of small language models.Microsoft Research Blog, 1:3, 2023. 11

  5. [13]

    Segment anything

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 4015–4026, 2023. 2, 3

  6. [14]

    Lisa: Reasoning segmentation via large language model

    Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. Lisa: Reasoning segmentation via large language model. ArXiv, abs/2308.00692, 2023. URL����� � ����������������������������������� ���������. 1, 2, 6

  7. [15]

    Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large lan- guage models.arXiv preprint arXiv:2301.12597,

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large lan- guage models.arXiv preprint arXiv:2301.12597,

  8. [16]

    Referring transformer: A one-step approach to multi-task visual grounding

    Muchen Li and Leonid Sigal. Referring transformer: A one-step approach to multi-task visual grounding. Advances in neural information processing systems, 34:19652–19664, 2021. 2

  9. [17]

    To- kenpacker: Efficient visual projector for multimodal llm.International Journal of Computer Vision, pages 1–19, 2025

    Wentong Li, Yuqian Yuan, Jian Liu, Dongqi Tang, Song Wang, Jie Qin, Jianke Zhu, and Lei Zhang. To- kenpacker: Efficient visual projector for multimodal llm.International Journal of Computer Vision, pages 1–19, 2025. 2, 6, 12

  10. [18]

    Boosting multimodal large language models with visual tokens withdrawal for rapid inference

    Zhihang Lin, Mingbao Lin, Luxi Lin, and Rongrong Ji. Boosting multimodal large language models with visual tokens withdrawal for rapid inference. InPro- ceedings of the AAAI Conference on Artificial Intelli- gence, volume 39, pages 5334–5342, 2025. 1, 2, 6

  11. [19]

    Swin transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. InProceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021. 11

  12. [20]

    Modeling context between objects for referring ex- pression understanding

    Varun K Nagaraja, Vlad I Morariu, and Larry S Davis. Modeling context between objects for referring ex- pression understanding. InComputer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14, pages 792–807. Sprin...

  13. [21]

    Perceptiongpt: Effectively fusing visual perception into llm.arXiv preprint arXiv:2311.06612,

    Renjie Pi, Lewei Yao, Jiahui Gao, Jipeng Zhang, and Tong Zhang. Perceptiongpt: Effectively fusing visual perception into llm.arXiv preprint arXiv:2311.06612,

  14. [22]

    Pixellm: Pixel reasoning with large multimodal model.ArXiv, abs/2312.02228, 2023

    Zhongwei Ren, Zhicheng Huang, Yunchao Wei, Yao Zhao, Dongmei Fu, Jiashi Feng, and Xiaojie Jin. Pixellm: Pixel reasoning with large multimodal model.ArXiv, abs/2312.02228, 2023. URL������ ����������������������������������� ���������. 1, 2

  15. [23]

    Urvos: Unified referring video object segmentation network with a large-scale benchmark

    Seonguk Seo, Joon-Young Lee, and Bohyung Han. Urvos: Unified referring video object segmentation network with a large-scale benchmark. InCom- puter Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XV 16, pages 208–223. Springer, 2020. 6

  16. [24]

    Llava-prumerge: Adaptive token re- duction for efficient large multimodal models.ICCV,

    Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, and Yan Yan. Llava-prumerge: Adaptive token re- duction for efficient large multimodal models.ICCV,

  17. [25]

    Contrastive grouping with transformer for referring image segmentation

    Jiajin Tang, Ge Zheng, Cheng Shi, and Sibei Yang. Contrastive grouping with transformer for referring image segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, pages 23570–23580, 2023. 2

  18. [26]

    Deep learning and the information bottleneck principle

    Naftali Tishby and Noga Zaslavsky. Deep learning and the information bottleneck principle. In2015 ieee information theory workshop (itw), pages 1–5. Ieee,

  19. [27]

    Cris: Clip-driven referring image segmentation

    Zhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao, Yandong Guo, Mingming Gong, and Tongliang Liu. Cris: Clip-driven referring image segmentation. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition, pages 11686– 11695, 2022. 2

  20. [28]

    Lasagna: Language-based segmenta- tion assistant for complex queries.arXiv preprint arXiv:2404.08506, 2024

    Cong Wei, Haoxian Tan, Yujie Zhong, Yujiu Yang, and Lin Ma. Lasagna: Language-based segmenta- tion assistant for complex queries.arXiv preprint arXiv:2404.08506, 2024. 6

  21. [29]

    Instructseg: Unifying instructed visual segmentation with multi- modal large language models.ICCV 2025, 2025

    Cong Wei, Yujie Zhong, Haoxian Tan, Yingsen Zeng, Yong Liu, Zheng Zhao, and Yujiu Yang. Instructseg: Unifying instructed visual segmentation with multi- modal large language models.ICCV 2025, 2025. 1, 2, 4, 6, 11, 12

  22. [30]

    Gsva: Generalized segmentation via multimodal large language models

    Zhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan, Shiji Song, and Gao Huang. Gsva: Generalized segmentation via multimodal large language models. arXiv preprint arXiv:2312.10103, 2023. 2

  23. [31]

    Visa: Reasoning video object seg- mentation via large language models.arXiv preprint arXiv:2407.11325, 2024

    Cilin Yan, Haochen Wang, Shilin Yan, Xiaolong Jiang, Yao Hu, Guoliang Kang, Weidi Xie, and Ef- stratios Gavves. Visa: Reasoning video object seg- mentation via large language models.arXiv preprint arXiv:2407.11325, 2024. 1, 2, 6

  24. [32]

    Visionzip: Longer is better but not necessary in vision language models

    Senqiao Yang, Yukang Chen, Zhuotao Tian, Chengyao Wang, Jingyao Li, Bei Yu, and Jiaya Jia. Visionzip: Longer is better but not necessary in vision language models. InProceedings of the Com- puter Vision and Pattern Recognition Conference, pages 19792–19802, 2025. 2, 6, 12

  25. [33]

    Fit and prune: Fast and training-free visual token pruning for multi-modal large language models

    Weihao Ye, Qiong Wu, Wenhao Lin, and Yiyi Zhou. Fit and prune: Fast and training-free visual token pruning for multi-modal large language models. In Proceedings of the AAAI Conference on Artificial In- telligence, volume 39, pages 22128–22136, 2025. 1, 2

  26. [34]

    Modeling context in re- ferring expressions

    Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling context in re- ferring expressions. InComputer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Nether- lands, October 11-14, 2016, Proceedings, Part II 14, pages 69–85. Springer, 2016. 6

  27. [35]

    Sigmoid loss for language im- age pre-training

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language im- age pre-training. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 11975–11986, 2023. 11

  28. [36]

    Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms.arXiv preprint arXiv:2412.01818,

    Qizhe Zhang, Aosong Cheng, Ming Lu, Renrui Zhang, Zhiyong Zhuo, Jiajun Cao, Shaobo Guo, Qi She, and Shanghang Zhang. Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms.arXiv preprint arXiv:2412.01818,

  29. [37]

    Psalm: Pixelwise segmentation with large multi- modal model.arXiv preprint arXiv:2403.14598, 2024

    Zheng Zhang, Yeyao Ma, Enming Zhang, and Xiang Bai. Psalm: Pixelwise segmentation with large multi- modal model.arXiv preprint arXiv:2403.14598, 2024. 2

  30. [38]

    Minigpt-4: Enhancing vision- language understanding with advanced large language models.arXiv preprint arXiv:2304.10592, 2023

    Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision- language understanding with advanced large language models.arXiv preprint arXiv:2304.10592, 2023. 2

  31. [39]

    dog with its mouth open,

    Xueyan Zou, Zi-Yi Dou, Jianwei Yang, Zhe Gan, Lin- jie Li, Chunyuan Li, Xiyang Dai, Harkirat Behl, Jian- feng Wang, Lu Yuan, et al. Generalized decoding for pixel, image, and language. InProceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.