Pith. sign in

REVIEW 3 major objections 1 minor 61 references

Multi-layer Abstraction for Nested Generation of Options (MANGO) in Hierarchical Reinforcement Learning

T0 review · 3 major / 1 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The MANGO abstract claims a nested-options HRL framework, but the submitted body is a different paper on CLIP bias transfer.

desk verdict The abstract and body are two different papers; the MANGO claims have no supporting content, so this is a desk reject, not a scientific evaluation. read the letter →

arxiv 2508.17751 v1 pith:2UHQO5GZ submitted 2025-08-25 cs.LG

classification cs.LG
keywords hierarchicalreinforcementlearningmulti-layerabstractionoptionsframeworknestedmacro-actionssparserewardtaskssampleefficiencygeneralizationmanuscriptmismatch
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The submission's abstract announces MANGO, a hierarchical reinforcement learning framework that stacks abstraction layers, nests options as macro-actions, and promises better sample efficiency and generalization on grid-world tasks. The body of the submitted text, however, is not the MANGO paper: it is a separate article on how social biases transfer from CLIP models to downstream tasks, with a different title, a different author list, and a footer that does not match the abstract's identifier. This means the manuscript, as presented, contains no algorithm, equations, or experiments that would support the MANGO claims. The one coherent claim that can be read in the full text is the CLIP bias-transfer article's empirical finding that bias measurement depends on the data subset considered.

What carries the argument

For the intended MANGO framework, the central object is a multi-layer options hierarchy: each layer defines an abstract state space, options serve as macro-actions, intra-layer policies guide transitions within the abstract space, and task actions carry task-specific reward components. This machinery is supposed to enable nesting and reuse of learned movement primitives across layers. In the submitted body text, the corresponding machinery is instead the CLIP bias-analysis setup: measuring pre-training bias on global and local subsets of data, then computing correlations between that bias and downstream bias in tasks such as VQA and captioning. No equations or training procedure for the inte

What would settle it

A concrete check is to inspect the supplied PDF's title page and footer: if the title and identifier there are not the same as those in the abstract, the MANGO claims have no supporting text. A second check is to search the body for the term MANGO; its absence outside the abstract would confirm that no algorithm or experiment is presented.

Watch

Extended reading notes

Core claim

The paper intended under this identifier aims to establish that decomposing a reinforcement learning task into multiple abstraction layers, where each layer defines an abstract state space and generates nested options, allows an agent to reuse learned macro-actions and thereby improve sample efficiency, generalization, and interpretability. The author's position would be that this nested-option hierarchy outperforms standard RL in procedurally generated sparse-reward grid environments. That position, however, is not demonstrated in the supplied body text: the body consists of a different manuscript, so the central claim is currently unsupported by any algorithmic description or experimental

Load-bearing premise

The load-bearing premise is that the body text is the MANGO manuscript; if that is not true, then every claim in the abstract is without evidence in the submission.

Editorial extensions

If this is right

  • If the intended MANGO claim held, agents would solve long-horizon sparse-reward tasks by reusing nested macro-actions across abstraction layers rather than relearning from scratch.
  • Sample efficiency and generalization on procedurally generated grid worlds would improve relative to standard RL baselines.
  • Layered options would make the agent's decision process transparent, which would matter for safety-critical and industrial deployments.
  • The abstract's proposed future work—automated abstraction discovery, continuous or fuzzy environments, and robust multi-layer training—would be the natural next tests of the framework.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A reader who wants to evaluate MANGO should seek a corrected manuscript; the submitted text cannot support or refute the framework's claims as it stands.
  • The CLIP article that actually occupies the body suggests that aggregate bias metrics may be misleading, since bias scores vary substantially between global and local data views.
  • The discrepancy between abstract and body points to a simple, testable safeguard: a consistency check that the supplied full text matches the abstract's title, author list, and identifier would have caught this mismatch before distribution.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The manuscript is titled "Multi-layer Abstraction for Nested Generation of Options (MANGO) in Hierarchical Reinforcement Learning" and its abstract claims a new HRL framework with nested options, intra-layer policies, task actions, and experiments in procedurally-generated grid environments that improve sample efficiency and generalization. However, the full text is an entirely different paper: "From Global to Local: Social Bias Transfer in CLIP" by Ramos et al., with its own abstract, introduction, references, and page footer reading arXiv:2508.17750v1 [cs.CV]. None of the promised MANGO formalism, algorithm definitions, environment descriptions, baseline comparisons, or experimental results appears anywhere in the body. As submitted, the document is not self-consistent and cannot support the abstract's claims.

Significance. If MANGO were actually specified and evaluated as described, the claimed improvements in sample efficiency and generalization for hierarchical RL would be a useful contribution to the field, particularly for sparse-reward and safety-critical applications. However, the submission provides no content that can be assessed for correctness, novelty, or reproducibility. The abstract alone is not a scientific paper, and the body text addresses a disjoint topic. The significance of the work therefore cannot be evaluated on the evidence presented.

major comments (3)
  1. [Full text (entire body)] The body of the submission is not the MANGO paper. The title, author list, abstract, introduction, and references all correspond to "From Global to Local: Social Bias Transfer in CLIP" (Ramos et al.), and the page footer reads arXiv:2508.17750v1 [cs.CV]. None of the MANGO framework—multi-layer abstraction, nested options, intra-layer policies, task actions—is defined in the text, and no grid-environment experiments, baselines, or result tables are reported. The central claim of the abstract is therefore entirely unsupported by the submitted manuscript.
  2. [Abstract vs. body consistency] Even taken on its own terms, the body text is a paper about bias transfer in CLIP models, with no connection to hierarchical reinforcement learning or option-based methods. There is no equation, algorithm, environment definition, or experimental protocol for MANGO anywhere in the document. The submission is internally inconsistent at the most basic level, making it impossible to review the proposed method for scientific soundness.
  3. [Reproducibility and empirical support] The abstract claims "substantial improvements in both sample efficiency and generalization capabilities compared to standard RL methods." No experimental setup, hyperparameters, environment generator, or code is provided. Even if the body text were the correct MANGO paper, the absence of any empirical artifact would prevent verification of the claimed results. As it stands, the claim is not merely unverified but unverifiable from the submitted text.
minor comments (1)
  1. [Page footer / metadata] The page footer lists arXiv:2508.17750v1 [cs.CV], which is inconsistent with the manuscript number 2508.17751 stated in the submission title. This should be corrected if the wrong file was uploaded; otherwise it signals a metadata mismatch that needs clarification.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity can be identified: the abstract's MANGO claims have no derivation chain in the body, which is an unrelated CLIP paper.

full rationale

Circularity analysis requires exhibiting a specific reduction in which a claimed prediction, fitted parameter, uniqueness theorem, or ansatz reduces by construction to the paper's own inputs. The submitted manuscript provides no such reduction. The abstract promises a hierarchical RL framework (MANGO) with nested options, intra-layer policies, and grid-world experiments, but the full text is 'From Global to Local: Social Bias Transfer in CLIP' by Ramos et al., with footer 'arXiv:2508.17750v1 [cs.CV]'. None of MANGO's formalism, equations, training objectives, or ablations appear anywhere in the body. Consequently there is no fitted-input-called-prediction step to exhibit, no self-citation chain to trace, and no definitional equivalence to expose. The mismatch is a serious document-integrity failure and leaves the abstract's empirical claims entirely unsupported, but unsupported is not the same as circular. Per the hard rules, one may only flag circularity with a quotation and a specific reduction; no such reduction can be exhibited for a manuscript whose substantive content is absent. Score 0 reflects the absence of detected circularity, not an endorsement of the submission's correctness or completeness.

Assumptions & free parameters 2 free parameters · 2 assumptions · 2 invented entities

Reconstructed from the abstract only because the submitted body is an unrelated paper (CLIP bias transfer). MANGO's components and all numerical choices (layer count, option lengths, learning rates, discount factors) are unspecified, so every load-bearing constant is an unknown free parameter awaiting specification in a corrected submission.

free parameters (2)
  • Number of abstraction layers (depth of option nesting)
    The abstract describes multiple layers of abstraction but does not state how many, how they are selected, or whether the count is tuned per environment; in HRL this choice materially affects sample efficiency.
  • Option granularity, termination conditions, and intra-layer policy hyperparameters
    Not specified anywhere in the submission (the body is a different paper); these would be hand-chosen and would influence the claimed gains over standard RL.
assumptions (2)
  • standard math Options / semi-Markov decision process formalism is the correct model for temporally extended actions
    The abstract's use of 'options' and 'macro-actions' presumes the standard HRL formalism without restating or deriving it.
  • domain assumption Temporal abstraction improves learning in long-horizon sparse-reward tasks
    The motivating premise of the abstract; asserted rather than demonstrated in the submitted text.
invented entities (2)
  • Intra-layer policies
    purpose: Guide the agent's transitions within each abstract state space layer
    Introduced as a MANGO component in the abstract; no definition, equations, or ablation appear in the submitted text, and the abstract does not distinguish it from existing intra-option/hierarchical policies.
  • Task actions
    purpose: Integrate task-specific components such as reward functions into the nested option structure
    Named in the abstract as a MANGO component; its mechanism is undefined in the submission, and no evidence is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multi-layer Abstraction for Nested Generation of Options (MANGO) in Hierarchical Reinforcement Learning." pith.science (2026). https://pith.science/paper/2UHQO5GZ

@misc{pith2026250817751,
  author       = {Pith},
  title        = {Pith review of: Multi-layer Abstraction for Nested Generation of Options (MANGO) in Hierarchical Reinforcement Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2UHQO5GZ}},
  note         = {Machine review of arXiv:2508.17751}
}
read the original abstract

This paper introduces MANGO (Multilayer Abstraction for Nested Generation of Options), a novel hierarchical reinforcement learning framework designed to address the challenges of long-term sparse reward environments. MANGO decomposes complex tasks into multiple layers of abstraction, where each layer defines an abstract state space and employs options to modularize trajectories into macro-actions. These options are nested across layers, allowing for efficient reuse of learned movements and improved sample efficiency. The framework introduces intra-layer policies that guide the agent's transitions within the abstract state space, and task actions that integrate task-specific components such as reward functions. Experiments conducted in procedurally-generated grid environments demonstrate substantial improvements in both sample efficiency and generalization capabilities compared to standard RL methods. MANGO also enhances interpretability by making the agent's decision-making process transparent across layers, which is particularly valuable in safety-critical and industrial applications. Future work will explore automated discovery of abstractions and abstract actions, adaptation to continuous or fuzzy environments, and more robust multi-layer training strategies.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

61 extracted references · 59 canonical work pages

  1. [1]

    Phi-3 technical report: A highly capable language model locally on your phone

    Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadal- lah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone. Technical report, Microsoft, 2024. 7

  2. [2]

    Evaluating CLIP: Towards characterization of broader capabilities and downstream implications, 2021

    Sandhini Agarwal, Gretchen Krueger, Jack Clark, Alec Rad- ford, Jong Wook Kim, and Miles Brundage. Evaluating CLIP: Towards characterization of broader capabilities and downstream implications, 2021. 1, 2

  3. [3]

    Flamingo: a visual language model for few-shot learning

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. In NeurIPS,

  4. [4]

    VQA: Visual question answering

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. VQA: Visual question answering. In ICCV, 2015. 3, 7

  5. [5]

    Fair- ness and Machine Learning: Limitations and Opportunities

    Solon Barocas, Moritz Hardt, and Arvind Narayanan. Fair- ness and Machine Learning: Limitations and Opportunities. MIT Press, 2023. 3

  6. [6]

    A prompt array keeps the bias away: Debiasing vision-language models with adversar- ial learning

    Hugo Berg, Siobhan Hall, Yash Bhalgat, Hannah Kirk, Alek- sandar Shtedritski, and Max Bain. A prompt array keeps the bias away: Debiasing vision-language models with adversar- ial learning. In AACL-IJCNLP, 2022. 4

  7. [7]

    Multimodal datasets: misogyny, pornography, and ma- lignant stereotypes, 2021

    Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahem- bwe. Multimodal datasets: misogyny, pornography, and ma- lignant stereotypes, 2021. 2

  8. [8]

    In- structPix2Pix: Learning to follow image editing instructions

    Tim Brooks, Aleksander Holynski, and Alexei A Efros. In- structPix2Pix: Learning to follow image editing instructions. In CVPR, 2023. 1

Show all 61 references
  1. [9]

    Evaluating bias and fairness in gender- neutral pretrained vision-and-language models

    Laura Cabello, Emanuele Bugliarello, Stephanie Brandl, and Desmond Elliott. Evaluating bias and fairness in gender- neutral pretrained vision-and-language models. In EMNLP,

  2. [10]

    Microsoft COCO captions: Data collection and evaluation server, 2015

    Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedan- tam, Saurabh Gupta, Piotr Doll ´ar, and C Lawrence Zitnick. Microsoft COCO captions: Data collection and evaluation server, 2015. 4

  3. [11]

    Reproducible scal- ing laws for contrastive language-image learning

    Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuh- mann, Ludwig Schmidt, and Jenia Jitsev. Reproducible scal- ing laws for contrastive language-image learning. In CVPR,

  4. [12]

    Debiasing vision-language models via biased prompts, 2023

    Ching-Yao Chuang, Varun Jampani, Yuanzhen Li, Antonio Torralba, and Stefanie Jegelka. Debiasing vision-language models via biased prompts, 2023. 4

  5. [13]

    Utility-fairness trade-offs and how to find them

    Sepehr Dehdashtian, Bashir Sadeghi, and Vishnu Naresh Boddeti. Utility-fairness trade-offs and how to find them. In CVPR, 2024. 2

  6. [14]

    FairerCLIP: Debiasing clip’s zero-shot predictions us- ing functions in RKHSs

    Sepehr Dehdashtian, Lan Wang, and Vishnu Naresh Bod- deti. FairerCLIP: Debiasing clip’s zero-shot predictions us- ing functions in RKHSs. In ICLR, 2024. 2, 4

  7. [15]

    BERT: Pre-training of deep bidirectional Trans- formers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional Trans- formers for language understanding. In NAACL, 2019. 4

  8. [16]

    From captions to vi- sual concepts and back

    Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh K Sri- vastava, Li Deng, Piotr Doll´ar, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C Platt, et al. From captions to vi- sual concepts and back. In CVPR, 2015. 3

  9. [17]

    Certifying and removing disparate impact

    Michael Feldman, Sorelle A Friedler, John Moeller, Carlos Scheidegger, and Suresh Venkatasubramanian. Certifying and removing disparate impact. In KDD, 2015. 2

  10. [18]

    Interpreting CLIP’s image representation via text-based de- composition

    Yossi Gandelsman, Alexei A Efros, and Jacob Steinhardt. Interpreting CLIP’s image representation via text-based de- composition. In ICLR, 2024. 2

  11. [19]

    Uncurated image-text datasets: Shedding light on demographic bias

    Noa Garcia, Yusuke Hirota, Yankun Wu, and Yuta Nakashima. Uncurated image-text datasets: Shedding light on demographic bias. In CVPR, 2023. 2, 4

  12. [20]

    Fairness-aware ranking in search & recommendation systems with application to linkedin talent search

    Sahin Cem Geyik, Stuart Ambler, and Krishnaram Kentha- padi. Fairness-aware ranking in search & recommendation systems with application to linkedin talent search. In KDD,

  13. [21]

    Biases propagate in encoder-based vision- language models: A systematic analysis from intrinsic mea- sures to zero-shot retrieval outcomes, 2025

    Kshitish Ghate, Tessa Charlesworth, Mona Diab, and Aylin Caliskan. Biases propagate in encoder-based vision- language models: A systematic analysis from intrinsic mea- sures to zero-shot retrieval outcomes, 2025. 1, 2, 8

  14. [22]

    Making the V in VQA matter: El- evating the role of image understanding in visual question answering

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Ba- tra, and Devi Parikh. Making the V in VQA matter: El- evating the role of image understanding in visual question answering. In CVPR, 2017. 4, 7

  15. [23]

    Equality of op- portunity in supervised learning

    Moritz Hardt, Eric Price, and Nati Srebro. Equality of op- portunity in supervised learning. In NeurIPS, 2016. 2

  16. [24]

    The bias of harmful label associations in vision- language models

    Caner Hazirbas, Alicia Sun, Yonathan Efroni, and Mark Ibrahim. The bias of harmful label associations in vision- language models. In ICLRW, 2024. 2

  17. [25]

    Gaussian Error Linear Units (GELUs), 2023

    Dan Hendrycks and Kevin Gimpel. Gaussian Error Linear Units (GELUs), 2023. 7

  18. [26]

    Quantify- ing societal bias amplification in image captioning

    Yusuke Hirota, Yuta Nakashima, and Noa Garcia. Quantify- ing societal bias amplification in image captioning. InCVPR,

  19. [27]

    GQA: A new dataset for real-world visual reasoning and compositional question answering

    Drew A Hudson and Christopher D Manning. GQA: A new dataset for real-world visual reasoning and compositional question answering. In CVPR, 2019. 7

  20. [28]

    FairFace: Face at- tribute dataset for balanced race, gender, and age for bias measurement and mitigation

    Kimmo Karkkainen and Jungseock Joo. FairFace: Face at- tribute dataset for balanced race, gender, and age for bias measurement and mitigation. In WACV, 2021. 4

  21. [29]

    Deep visual-semantic align- ments for generating image descriptions

    Andrej Karpathy and Li Fei-Fei. Deep visual-semantic align- ments for generating image descriptions. In CVPR, 2015. 3

  22. [30]

    Microsoft COCO: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014. 4, 7

  23. [31]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023. 1, 7

  24. [32]

    Learning adversarially fair and transferable represen- tations

    David Madras, Elliot Creager, Toniann Pitassi, and Richard Zemel. Learning adversarially fair and transferable represen- tations. In ICML, 2018. 2 9

  25. [33]

    Gen- der bias in multimodal models: A transnational feminist ap- proach considering geographical region and culture, 2023

    Abhishek Mandal, Suzanne Little, and Susan Leavy. Gen- der bias in multimodal models: A transnational feminist ap- proach considering geographical region and culture, 2023. 1, 2

  26. [34]

    OK-VQA: A visual question answering benchmark requiring external knowledge

    Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. OK-VQA: A visual question answering benchmark requiring external knowledge. In CVPR, 2019. 7

  27. [35]

    Provably fair representations, 2017

    Daniel McNamara, Cheng Soon Ong, and Robert C Williamson. Provably fair representations, 2017. 2

  28. [36]

    Gender arti- facts in visual datasets

    Nicole Meister, Dora Zhao, Angelina Wang, Vikram V Ra- maswamy, Ruth Fong, and Olga Russakovsky. Gender arti- facts in visual datasets. In ICCV, 2023. 3

  29. [37]

    Null-text inversion for editing real images using guided diffusion models

    Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Null-text inversion for editing real images using guided diffusion models. In CVPR, 2023. 1

  30. [38]

    Im2Text: Describing images using 1 million captioned pho- tographs

    Vicente Ordonez, Girish Kulkarni, and Tamara Berg. Im2Text: Describing images using 1 million captioned pho- tographs. In NeurIPS, 2011. 7

  31. [39]

    Learn- ing transferable visual models from natural language super- vision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In ICML, 2021. 1

  32. [40]

    Zero-shot text-to-image generation

    Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In ICML, 2021. 1

  33. [41]

    High-resolution image syn- thesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. In CVPR, 2022. 1

  34. [42]

    Distributionally robust neural networks for group shifts: On the importance of regularization for worst- case generalization

    Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang. Distributionally robust neural networks for group shifts: On the importance of regularization for worst- case generalization. In ICLR, 2020. 2

  35. [43]

    LAION-5B: An open large-scale dataset for train- ing next generation image-text models

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, et al. LAION-5B: An open large-scale dataset for train- ing next generation image-text models. In NeurIPS, 2022. 7

  36. [44]

    A-OKVQA: A benchmark for visual question answering using world knowl- edge

    Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-OKVQA: A benchmark for visual question answering using world knowl- edge. In ECCV, 2022. 7

  37. [45]

    Conceptual Captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

    Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual Captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning. In ACL,

  38. [46]

    Fair representation: guaranteeing approximate multiple group fairness for unknown tasks

    Xudong Shen, Yongkang Wong, and Mohan Kankanhalli. Fair representation: guaranteeing approximate multiple group fairness for unknown tasks. PAMI, 2022. 2

  39. [47]

    Towards VQA models that can read

    Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards VQA models that can read. In CVPR,

  40. [48]

    Upstream mitigation is not all you need: Testing the bias transfer hypothesis in pre-trained language models

    Ryan Steed, Swetasudha Panda, Ari Kobren, and Michael Wick. Upstream mitigation is not all you need: Testing the bias transfer hypothesis in pre-trained language models. In ACL, 2022. 1, 2, 4, 7, 8

  41. [49]

    CIDEr: Consensus-based image description evalu- ation

    Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. CIDEr: Consensus-based image description evalu- ation. In CVPR, 2015. 4

  42. [50]

    Show and tell: A neural image caption gen- erator

    Oriol Vinyals, Alexander Toshev, Samy Bengio, and Du- mitru Erhan. Show and tell: A neural image caption gen- erator. In CVPR, 2015. 3

  43. [51]

    Directional bias am- plification

    Angelina Wang and Olga Russakovsky. Directional bias am- plification. In ICML, 2021. 4

  44. [52]

    Overwriting pre- trained bias with finetuning data

    Angelina Wang and Olga Russakovsky. Overwriting pre- trained bias with finetuning data. In ICCV, 2023. 1, 2, 8

  45. [53]

    Are gender-neutral queries really gender-neutral? Mitigating gender bias in im- age search

    Jialu Wang, Yang Liu, and Xin Wang. Are gender-neutral queries really gender-neutral? Mitigating gender bias in im- age search. In EMNLP, 2021. 1, 2

  46. [54]

    American == white in multimodal language-and-image AI

    Robert Wolfe and Aylin Caliskan. American == white in multimodal language-and-image AI. In AIES, 2022. 1, 2

  47. [55]

    Markedness in visual se- mantic AI

    Robert Wolfe and Aylin Caliskan. Markedness in visual se- mantic AI. In FAccT, 2022. 1, 2

  48. [56]

    Contrastive language-vision AI models pretrained on web- scraped multimodal data exhibit sexual objectification bias

    Robert Wolfe, Yiwei Yang, Bill Howe, and Aylin Caliskan. Contrastive language-vision AI models pretrained on web- scraped multimodal data exhibit sexual objectification bias. In FAccT, 2023. 2

  49. [57]

    mPLUG-Owl: Modularization empowers large language models with multimodality, 2024

    Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. mPLUG-Owl: Modularization empowers large language models with multimodality, 2024. 1

  50. [58]

    Yi: Open foundation models by 01.ai,

    Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al. Yi: Open foundation models by 01.ai,

  51. [59]

    Under- standing and evaluating racial biases in image captioning

    Dora Zhao, Angelina Wang, and Olga Russakovsky. Under- standing and evaluating racial biases in image captioning. In ICCV, 2021. 4

  52. [60]

    Inherent tradeoffs in learning fair representations

    Han Zhao and Geoffrey J Gordon. Inherent tradeoffs in learning fair representations. JMLR, 2022. 2

  53. [61]

    TinyLLaV A: A framework of small-scale large multimodal models, 2024

    Baichuan Zhou, Ying Hu, Xi Weng, Junlong Jia, Jie Luo, Xien Liu, Ji Wu, and Lei Huang. TinyLLaV A: A framework of small-scale large multimodal models, 2024. 6 10

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.