REVIEW 3 major objections 1 minor 61 references
Multi-layer Abstraction for Nested Generation of Options (MANGO) in Hierarchical Reinforcement Learning
T0 review · 3 major / 1 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The MANGO abstract claims a nested-options HRL framework, but the submitted body is a different paper on CLIP bias transfer.
desk verdict The abstract and body are two different papers; the MANGO claims have no supporting content, so this is a desk reject, not a scientific evaluation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
For the intended MANGO framework, the central object is a multi-layer options hierarchy: each layer defines an abstract state space, options serve as macro-actions, intra-layer policies guide transitions within the abstract space, and task actions carry task-specific reward components. This machinery is supposed to enable nesting and reuse of learned movement primitives across layers. In the submitted body text, the corresponding machinery is instead the CLIP bias-analysis setup: measuring pre-training bias on global and local subsets of data, then computing correlations between that bias and downstream bias in tasks such as VQA and captioning. No equations or training procedure for the inte
What would settle it
A concrete check is to inspect the supplied PDF's title page and footer: if the title and identifier there are not the same as those in the abstract, the MANGO claims have no supporting text. A second check is to search the body for the term MANGO; its absence outside the abstract would confirm that no algorithm or experiment is presented.
Extended reading notes
Core claim
The paper intended under this identifier aims to establish that decomposing a reinforcement learning task into multiple abstraction layers, where each layer defines an abstract state space and generates nested options, allows an agent to reuse learned macro-actions and thereby improve sample efficiency, generalization, and interpretability. The author's position would be that this nested-option hierarchy outperforms standard RL in procedurally generated sparse-reward grid environments. That position, however, is not demonstrated in the supplied body text: the body consists of a different manuscript, so the central claim is currently unsupported by any algorithmic description or experimental
Load-bearing premise
The load-bearing premise is that the body text is the MANGO manuscript; if that is not true, then every claim in the abstract is without evidence in the submission.
Editorial extensions
If this is right
- If the intended MANGO claim held, agents would solve long-horizon sparse-reward tasks by reusing nested macro-actions across abstraction layers rather than relearning from scratch.
- Sample efficiency and generalization on procedurally generated grid worlds would improve relative to standard RL baselines.
- Layered options would make the agent's decision process transparent, which would matter for safety-critical and industrial deployments.
- The abstract's proposed future work—automated abstraction discovery, continuous or fuzzy environments, and robust multi-layer training—would be the natural next tests of the framework.
Reading between the lines
- A reader who wants to evaluate MANGO should seek a corrected manuscript; the submitted text cannot support or refute the framework's claims as it stands.
- The CLIP article that actually occupies the body suggests that aggregate bias metrics may be misleading, since bias scores vary substantially between global and local data views.
- The discrepancy between abstract and body points to a simple, testable safeguard: a consistency check that the supplied full text matches the abstract's title, author list, and identifier would have caught this mismatch before distribution.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript is titled "Multi-layer Abstraction for Nested Generation of Options (MANGO) in Hierarchical Reinforcement Learning" and its abstract claims a new HRL framework with nested options, intra-layer policies, task actions, and experiments in procedurally-generated grid environments that improve sample efficiency and generalization. However, the full text is an entirely different paper: "From Global to Local: Social Bias Transfer in CLIP" by Ramos et al., with its own abstract, introduction, references, and page footer reading arXiv:2508.17750v1 [cs.CV]. None of the promised MANGO formalism, algorithm definitions, environment descriptions, baseline comparisons, or experimental results appears anywhere in the body. As submitted, the document is not self-consistent and cannot support the abstract's claims.
Significance. If MANGO were actually specified and evaluated as described, the claimed improvements in sample efficiency and generalization for hierarchical RL would be a useful contribution to the field, particularly for sparse-reward and safety-critical applications. However, the submission provides no content that can be assessed for correctness, novelty, or reproducibility. The abstract alone is not a scientific paper, and the body text addresses a disjoint topic. The significance of the work therefore cannot be evaluated on the evidence presented.
major comments (3)
- [Full text (entire body)] The body of the submission is not the MANGO paper. The title, author list, abstract, introduction, and references all correspond to "From Global to Local: Social Bias Transfer in CLIP" (Ramos et al.), and the page footer reads arXiv:2508.17750v1 [cs.CV]. None of the MANGO framework—multi-layer abstraction, nested options, intra-layer policies, task actions—is defined in the text, and no grid-environment experiments, baselines, or result tables are reported. The central claim of the abstract is therefore entirely unsupported by the submitted manuscript.
- [Abstract vs. body consistency] Even taken on its own terms, the body text is a paper about bias transfer in CLIP models, with no connection to hierarchical reinforcement learning or option-based methods. There is no equation, algorithm, environment definition, or experimental protocol for MANGO anywhere in the document. The submission is internally inconsistent at the most basic level, making it impossible to review the proposed method for scientific soundness.
- [Reproducibility and empirical support] The abstract claims "substantial improvements in both sample efficiency and generalization capabilities compared to standard RL methods." No experimental setup, hyperparameters, environment generator, or code is provided. Even if the body text were the correct MANGO paper, the absence of any empirical artifact would prevent verification of the claimed results. As it stands, the claim is not merely unverified but unverifiable from the submitted text.
minor comments (1)
- [Page footer / metadata] The page footer lists arXiv:2508.17750v1 [cs.CV], which is inconsistent with the manuscript number 2508.17751 stated in the submission title. This should be corrected if the wrong file was uploaded; otherwise it signals a metadata mismatch that needs clarification.
Circularity Check
No circularity can be identified: the abstract's MANGO claims have no derivation chain in the body, which is an unrelated CLIP paper.
full rationale
Circularity analysis requires exhibiting a specific reduction in which a claimed prediction, fitted parameter, uniqueness theorem, or ansatz reduces by construction to the paper's own inputs. The submitted manuscript provides no such reduction. The abstract promises a hierarchical RL framework (MANGO) with nested options, intra-layer policies, and grid-world experiments, but the full text is 'From Global to Local: Social Bias Transfer in CLIP' by Ramos et al., with footer 'arXiv:2508.17750v1 [cs.CV]'. None of MANGO's formalism, equations, training objectives, or ablations appear anywhere in the body. Consequently there is no fitted-input-called-prediction step to exhibit, no self-citation chain to trace, and no definitional equivalence to expose. The mismatch is a serious document-integrity failure and leaves the abstract's empirical claims entirely unsupported, but unsupported is not the same as circular. Per the hard rules, one may only flag circularity with a quotation and a specific reduction; no such reduction can be exhibited for a manuscript whose substantive content is absent. Score 0 reflects the absence of detected circularity, not an endorsement of the submission's correctness or completeness.
Assumptions & free parameters
free parameters (2)
- Number of abstraction layers (depth of option nesting)
- Option granularity, termination conditions, and intra-layer policy hyperparameters
assumptions (2)
- standard math Options / semi-Markov decision process formalism is the correct model for temporally extended actions
- domain assumption Temporal abstraction improves learning in long-horizon sparse-reward tasks
invented entities (2)
-
Intra-layer policies
-
Task actions
Cite this review
Pith. "Pith review of Multi-layer Abstraction for Nested Generation of Options (MANGO) in Hierarchical Reinforcement Learning." pith.science (2026). https://pith.science/paper/2UHQO5GZ
@misc{pith2026250817751,
author = {Pith},
title = {Pith review of: Multi-layer Abstraction for Nested Generation of Options (MANGO) in Hierarchical Reinforcement Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/2UHQO5GZ}},
note = {Machine review of arXiv:2508.17751}
}
read the original abstract
This paper introduces MANGO (Multilayer Abstraction for Nested Generation of Options), a novel hierarchical reinforcement learning framework designed to address the challenges of long-term sparse reward environments. MANGO decomposes complex tasks into multiple layers of abstraction, where each layer defines an abstract state space and employs options to modularize trajectories into macro-actions. These options are nested across layers, allowing for efficient reuse of learned movements and improved sample efficiency. The framework introduces intra-layer policies that guide the agent's transitions within the abstract state space, and task actions that integrate task-specific components such as reward functions. Experiments conducted in procedurally-generated grid environments demonstrate substantial improvements in both sample efficiency and generalization capabilities compared to standard RL methods. MANGO also enhances interpretability by making the agent's decision-making process transparent across layers, which is particularly valuable in safety-critical and industrial applications. Future work will explore automated discovery of abstractions and abstract actions, adaptation to continuous or fuzzy environments, and more robust multi-layer training strategies.
Reference graph
Works this paper leans on
-
[1]
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadal- lah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone. Technical report, Microsoft, 2024. 7
work page 2024
-
[2]
Evaluating CLIP: Towards characterization of broader capabilities and downstream implications, 2021
Sandhini Agarwal, Gretchen Krueger, Jack Clark, Alec Rad- ford, Jong Wook Kim, and Miles Brundage. Evaluating CLIP: Towards characterization of broader capabilities and downstream implications, 2021. 1, 2
work page 2021
-
[3]
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. In NeurIPS,
-
[4]
VQA: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. VQA: Visual question answering. In ICCV, 2015. 3, 7
work page 2015
-
[5]
Fair- ness and Machine Learning: Limitations and Opportunities
Solon Barocas, Moritz Hardt, and Arvind Narayanan. Fair- ness and Machine Learning: Limitations and Opportunities. MIT Press, 2023. 3
work page 2023
-
[6]
A prompt array keeps the bias away: Debiasing vision-language models with adversar- ial learning
Hugo Berg, Siobhan Hall, Yash Bhalgat, Hannah Kirk, Alek- sandar Shtedritski, and Max Bain. A prompt array keeps the bias away: Debiasing vision-language models with adversar- ial learning. In AACL-IJCNLP, 2022. 4
work page 2022
-
[7]
Multimodal datasets: misogyny, pornography, and ma- lignant stereotypes, 2021
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahem- bwe. Multimodal datasets: misogyny, pornography, and ma- lignant stereotypes, 2021. 2
work page 2021
-
[8]
In- structPix2Pix: Learning to follow image editing instructions
Tim Brooks, Aleksander Holynski, and Alexei A Efros. In- structPix2Pix: Learning to follow image editing instructions. In CVPR, 2023. 1
work page 2023
Show all 61 references
-
[9]
Evaluating bias and fairness in gender- neutral pretrained vision-and-language models
Laura Cabello, Emanuele Bugliarello, Stephanie Brandl, and Desmond Elliott. Evaluating bias and fairness in gender- neutral pretrained vision-and-language models. In EMNLP,
-
[10]
Microsoft COCO captions: Data collection and evaluation server, 2015
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedan- tam, Saurabh Gupta, Piotr Doll ´ar, and C Lawrence Zitnick. Microsoft COCO captions: Data collection and evaluation server, 2015. 4
2015
-
[11]
Reproducible scal- ing laws for contrastive language-image learning
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuh- mann, Ludwig Schmidt, and Jenia Jitsev. Reproducible scal- ing laws for contrastive language-image learning. In CVPR,
-
[12]
Debiasing vision-language models via biased prompts, 2023
Ching-Yao Chuang, Varun Jampani, Yuanzhen Li, Antonio Torralba, and Stefanie Jegelka. Debiasing vision-language models via biased prompts, 2023. 4
2023
-
[13]
Utility-fairness trade-offs and how to find them
Sepehr Dehdashtian, Bashir Sadeghi, and Vishnu Naresh Boddeti. Utility-fairness trade-offs and how to find them. In CVPR, 2024. 2
2024
-
[14]
FairerCLIP: Debiasing clip’s zero-shot predictions us- ing functions in RKHSs
Sepehr Dehdashtian, Lan Wang, and Vishnu Naresh Bod- deti. FairerCLIP: Debiasing clip’s zero-shot predictions us- ing functions in RKHSs. In ICLR, 2024. 2, 4
2024
-
[15]
BERT: Pre-training of deep bidirectional Trans- formers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional Trans- formers for language understanding. In NAACL, 2019. 4
2019
-
[16]
From captions to vi- sual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh K Sri- vastava, Li Deng, Piotr Doll´ar, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C Platt, et al. From captions to vi- sual concepts and back. In CVPR, 2015. 3
2015
-
[17]
Certifying and removing disparate impact
Michael Feldman, Sorelle A Friedler, John Moeller, Carlos Scheidegger, and Suresh Venkatasubramanian. Certifying and removing disparate impact. In KDD, 2015. 2
2015
-
[18]
Interpreting CLIP’s image representation via text-based de- composition
Yossi Gandelsman, Alexei A Efros, and Jacob Steinhardt. Interpreting CLIP’s image representation via text-based de- composition. In ICLR, 2024. 2
2024
-
[19]
Uncurated image-text datasets: Shedding light on demographic bias
Noa Garcia, Yusuke Hirota, Yankun Wu, and Yuta Nakashima. Uncurated image-text datasets: Shedding light on demographic bias. In CVPR, 2023. 2, 4
2023
-
[20]
Fairness-aware ranking in search & recommendation systems with application to linkedin talent search
Sahin Cem Geyik, Stuart Ambler, and Krishnaram Kentha- padi. Fairness-aware ranking in search & recommendation systems with application to linkedin talent search. In KDD,
-
[21]
Biases propagate in encoder-based vision- language models: A systematic analysis from intrinsic mea- sures to zero-shot retrieval outcomes, 2025
Kshitish Ghate, Tessa Charlesworth, Mona Diab, and Aylin Caliskan. Biases propagate in encoder-based vision- language models: A systematic analysis from intrinsic mea- sures to zero-shot retrieval outcomes, 2025. 1, 2, 8
2025
-
[22]
Making the V in VQA matter: El- evating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Ba- tra, and Devi Parikh. Making the V in VQA matter: El- evating the role of image understanding in visual question answering. In CVPR, 2017. 4, 7
2017
-
[23]
Equality of op- portunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro. Equality of op- portunity in supervised learning. In NeurIPS, 2016. 2
2016
-
[24]
The bias of harmful label associations in vision- language models
Caner Hazirbas, Alicia Sun, Yonathan Efroni, and Mark Ibrahim. The bias of harmful label associations in vision- language models. In ICLRW, 2024. 2
2024
-
[25]
Gaussian Error Linear Units (GELUs), 2023
Dan Hendrycks and Kevin Gimpel. Gaussian Error Linear Units (GELUs), 2023. 7
2023
-
[26]
Quantify- ing societal bias amplification in image captioning
Yusuke Hirota, Yuta Nakashima, and Noa Garcia. Quantify- ing societal bias amplification in image captioning. InCVPR,
-
[27]
GQA: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning. GQA: A new dataset for real-world visual reasoning and compositional question answering. In CVPR, 2019. 7
2019
-
[28]
FairFace: Face at- tribute dataset for balanced race, gender, and age for bias measurement and mitigation
Kimmo Karkkainen and Jungseock Joo. FairFace: Face at- tribute dataset for balanced race, gender, and age for bias measurement and mitigation. In WACV, 2021. 4
2021
-
[29]
Deep visual-semantic align- ments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. Deep visual-semantic align- ments for generating image descriptions. In CVPR, 2015. 3
2015
-
[30]
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014. 4, 7
2014
-
[31]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023. 1, 7
2023
-
[32]
Learning adversarially fair and transferable represen- tations
David Madras, Elliot Creager, Toniann Pitassi, and Richard Zemel. Learning adversarially fair and transferable represen- tations. In ICML, 2018. 2 9
2018
-
[33]
Gen- der bias in multimodal models: A transnational feminist ap- proach considering geographical region and culture, 2023
Abhishek Mandal, Suzanne Little, and Susan Leavy. Gen- der bias in multimodal models: A transnational feminist ap- proach considering geographical region and culture, 2023. 1, 2
2023
-
[34]
OK-VQA: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. OK-VQA: A visual question answering benchmark requiring external knowledge. In CVPR, 2019. 7
2019
-
[35]
Provably fair representations, 2017
Daniel McNamara, Cheng Soon Ong, and Robert C Williamson. Provably fair representations, 2017. 2
2017
-
[36]
Gender arti- facts in visual datasets
Nicole Meister, Dora Zhao, Angelina Wang, Vikram V Ra- maswamy, Ruth Fong, and Olga Russakovsky. Gender arti- facts in visual datasets. In ICCV, 2023. 3
2023
-
[37]
Null-text inversion for editing real images using guided diffusion models
Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Null-text inversion for editing real images using guided diffusion models. In CVPR, 2023. 1
2023
-
[38]
Im2Text: Describing images using 1 million captioned pho- tographs
Vicente Ordonez, Girish Kulkarni, and Tamara Berg. Im2Text: Describing images using 1 million captioned pho- tographs. In NeurIPS, 2011. 7
2011
-
[39]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In ICML, 2021. 1
2021
-
[40]
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In ICML, 2021. 1
2021
-
[41]
High-resolution image syn- thesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. In CVPR, 2022. 1
2022
-
[42]
Distributionally robust neural networks for group shifts: On the importance of regularization for worst- case generalization
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang. Distributionally robust neural networks for group shifts: On the importance of regularization for worst- case generalization. In ICLR, 2020. 2
2020
-
[43]
LAION-5B: An open large-scale dataset for train- ing next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, et al. LAION-5B: An open large-scale dataset for train- ing next generation image-text models. In NeurIPS, 2022. 7
2022
-
[44]
A-OKVQA: A benchmark for visual question answering using world knowl- edge
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-OKVQA: A benchmark for visual question answering using world knowl- edge. In ECCV, 2022. 7
2022
-
[45]
Conceptual Captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual Captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning. In ACL,
-
[46]
Fair representation: guaranteeing approximate multiple group fairness for unknown tasks
Xudong Shen, Yongkang Wong, and Mohan Kankanhalli. Fair representation: guaranteeing approximate multiple group fairness for unknown tasks. PAMI, 2022. 2
2022
-
[47]
Towards VQA models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards VQA models that can read. In CVPR,
-
[48]
Upstream mitigation is not all you need: Testing the bias transfer hypothesis in pre-trained language models
Ryan Steed, Swetasudha Panda, Ari Kobren, and Michael Wick. Upstream mitigation is not all you need: Testing the bias transfer hypothesis in pre-trained language models. In ACL, 2022. 1, 2, 4, 7, 8
2022
-
[49]
CIDEr: Consensus-based image description evalu- ation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. CIDEr: Consensus-based image description evalu- ation. In CVPR, 2015. 4
2015
-
[50]
Show and tell: A neural image caption gen- erator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Du- mitru Erhan. Show and tell: A neural image caption gen- erator. In CVPR, 2015. 3
2015
-
[51]
Directional bias am- plification
Angelina Wang and Olga Russakovsky. Directional bias am- plification. In ICML, 2021. 4
2021
-
[52]
Overwriting pre- trained bias with finetuning data
Angelina Wang and Olga Russakovsky. Overwriting pre- trained bias with finetuning data. In ICCV, 2023. 1, 2, 8
2023
-
[53]
Are gender-neutral queries really gender-neutral? Mitigating gender bias in im- age search
Jialu Wang, Yang Liu, and Xin Wang. Are gender-neutral queries really gender-neutral? Mitigating gender bias in im- age search. In EMNLP, 2021. 1, 2
2021
-
[54]
American == white in multimodal language-and-image AI
Robert Wolfe and Aylin Caliskan. American == white in multimodal language-and-image AI. In AIES, 2022. 1, 2
2022
-
[55]
Markedness in visual se- mantic AI
Robert Wolfe and Aylin Caliskan. Markedness in visual se- mantic AI. In FAccT, 2022. 1, 2
2022
-
[56]
Contrastive language-vision AI models pretrained on web- scraped multimodal data exhibit sexual objectification bias
Robert Wolfe, Yiwei Yang, Bill Howe, and Aylin Caliskan. Contrastive language-vision AI models pretrained on web- scraped multimodal data exhibit sexual objectification bias. In FAccT, 2023. 2
2023
-
[57]
mPLUG-Owl: Modularization empowers large language models with multimodality, 2024
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. mPLUG-Owl: Modularization empowers large language models with multimodality, 2024. 1
2024
-
[58]
Yi: Open foundation models by 01.ai,
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al. Yi: Open foundation models by 01.ai,
-
[59]
Under- standing and evaluating racial biases in image captioning
Dora Zhao, Angelina Wang, and Olga Russakovsky. Under- standing and evaluating racial biases in image captioning. In ICCV, 2021. 4
2021
-
[60]
Inherent tradeoffs in learning fair representations
Han Zhao and Geoffrey J Gordon. Inherent tradeoffs in learning fair representations. JMLR, 2022. 2
2022
-
[61]
TinyLLaV A: A framework of small-scale large multimodal models, 2024
Baichuan Zhou, Ying Hu, Xi Weng, Junlong Jia, Jie Luo, Xien Liu, Ji Wu, and Lei Huang. TinyLLaV A: A framework of small-scale large multimodal models, 2024. 6 10
2024
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.