Pith. sign in

REVIEW 5 major objections 6 minor 58 references

SEMANTIC SEE-THROUGH GOGGLES: Wearing Linguistic Virtual Reality in (Artificial) Intelligence

T0 review · 5 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Goggles that caption a live camera view and redraw it as an image let wearers feel, in the first person, the biases and information losses of AI's linguistic mediation of the world.

desk verdict A genuinely novel experiential prototype wrapped in a rough, unfinished manuscript; the qualitative evidence for the core claim is real but fragile. read the letter →

arxiv 2412.02641 v1 pith:I4PIWTF7 submitted 2024-12-03 cs.HC

classification cs.HC
keywords semanticsee-throughgoggleslinguisticvirtualrealityAIbiasimagecaptioningtext-to-imagegenerationfirst-personexperiencemediationhuman-in-the-loop
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that reducing a live camera view to a single sentence and then redrawing it as an image—before it reaches the eye—lets a person keep perceiving and acting in the physical world while feeling, from the inside, what AI mediation does to reality. It reports building such a real-time prototype, checking that the caption-then-redraw loop preserves some semantic content while dropping local visual detail, and running a workshop where nine participants wore the goggles and described their experience. The authors' central argument is that this makes AI bias and information loss a first-person bodily experience rather than an abstract critique, and that the same losses occur in all communication of scenes through language, including memory. The paper also introduces a concept the authors call Linguistic/Semantic Virtual Reality: realities that collapse into the same sentence are equivalent as virtual realities.

What carries the argument

The load-bearing object is the two-step 'writer and painter' loop housed in an HMD: a first AI compresses the camera image into a single sentence (image-to-text adaptation), and a second AI expands that sentence into a fresh image (text-to-image adaptation), with the sentence as the narrow channel through which all information must pass. The loop is what turns the abstract idea of semantic mediation into a first-person experience: information that does not fit in the sentence is lost, and information the models assume is injected back, so the generated view carries the captioning model's salience and the generation model's stereotypes. A human at the end of the loop closes it by acting on the redrawn world, which is how the paper establishes that the goggles are genuinely 'see-through' despite the transformation. The same loop defines the paper's conceptual contribution, Linguistic/Semantic Virtual Reality, by declaring equivalence among realities that produce the same sentence.

What would settle it

If participants who are blind to the hypothesis and wearing inert glasses that show unprocessed camera video report the same themes of semantic loss, averaging, and bias as those wearing the goggles, then the claim that the goggles specifically enable first-person insight into AI mediation would be falsified.

Watch

Extended reading notes

Core claim

The paper proposes Semantic See-through Goggles as an experimental framework and a working prototype. A camera feeds images into an image-to-text model that produces one sentence, and a text-to-image model redraws that sentence in about one second; the wearer sees only the redrawn image, while an external display shows the original image, the sentence, and the redrawn image to bystanders. The authors claim this arrangement subjectively captures the situation in which AI serves as a proxy for our perception of the world. In their quantitative checks, captions of input and output images were more similar to each other than to random pairings across four linguistic metrics, with the effect growing as the metrics became more semantic; visually, color histograms survived better than local features such as edges, and perceptual similarity was only moderate. In the nine-person workshop, participants could walk, reach, pick up and eat an apple, and respond to people, while reporting that the view was stereotyped and beautified—white muscular men, slender white women—and that the feeling of algorithmic bias and fear was no longer merely knowledge. Participants gradually identified the same structure in ordinary verbal communication and in their own memories, and the paper concludes that the experience is not limited to AI but belongs to any intelligence that sees the world under meaning.

Load-bearing premise

The load-bearing premise is that participants' interview responses reflect genuine first-person insight rather than expectations, because the first author conducted the interviews, participants knew the study's purpose, and there was no control or blind comparison.

Editorial extensions

If this is right

  • AI bias becomes a felt, first-person phenomenon: wearers report fear and discomfort at having their view selected, averaged, and interpreted without permission, which can move the debate from abstract critique to bodily understanding.
  • The framework acts as a human-in-the-loop tool for bias inspection, letting users compare input and output on the same semantic layer and challenge the model by performing gestures that push the view toward or away from stereotypes.
  • The issues revealed are general to language: telling someone about a scene compresses and reconstructs it just as the goggles do, so the experience extends to everyday communication and to memory, where the authors note people make memories smaller by putting them into words.
  • The proposed equivalence 'realities that become the same sentence are equivalent' provides a new definition of virtual reality based on linguistic identity rather than retinal or sensory identity.
  • If the method is right, tuning parameters such as caption length and generation steps directly governs the boundary between preserved and lost meaning, making the system a testbed for how semantic compression changes experience.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper leaves implicit is a controlled comparison: the same scene described verbally by a participant without the goggles, or mediated by a human writer, would test which reported effects are specific to AI rather than to any language-mediated perception.
  • The quantitative preservation measurements could be turned into a testable prediction: models whose linguistic similarity scores are higher should produce goggles experiences rated as more 'see-through' and less disorienting, linking first-person reports to objective metrics.
  • The framework could serve as an audit instrument for generative models: by having wearers localize which objects, genders, or races are stereotyped in their own view, biased outputs become traceable to specific caption-generation steps in a way that static image comparisons do not capture.
  • If the equivalence claim is taken seriously, it implies that changing the captioning model changes the ontology of the wearer's experience—realities that are equivalent under one model need not be equivalent under another—so the goggles formalize model-dependent perception.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper proposes 'Semantic See-through Goggles' (SSTG), a wearable video see-through HMD that captures a camera view, converts it in real time to a single sentence via an image-captioning model (BLIP), and regenerates an image from that sentence via a text-to-image model (Latent Consistency Model). The authors build a prototype, report a quantitative study comparing paired versus random image-to-image transformations on eight similarity metrics, collect open-ended questionnaire responses from 24 online participants, and conduct a workshop with nine participants whose interviews are analyzed qualitatively. The central claim is that wearing the goggles yields a first-person, subjective experience of AI's linguistic mediation of perception, including its biases and information losses. The paper also introduces the concept of 'Linguistic Virtual Reality,' in which realities that reduce to the same sentence are treated as equivalent.

Significance. If the experiential claim is validated, the contribution is notable: SSTG offers a new human-in-the-loop instrument for making AI bias and semantic compression experientially concrete, complementing audit-style and visualization-based approaches. The prototype is described in enough detail to be reproduced, the workshop participants report vivid, sometimes critical observations, and the conceptual discussion connects AI ethics, VR theory, and communication studies. The quantitative results, while preliminary, indicate that the pipeline preserves some high-level semantics and color structure while losing local features such as edges and exact spatial layout. However, the paper's main claim about first-person subjective insight is not established by the current evidence, and several internal inconsistencies need to be resolved before the qualitative results can be taken as reliable.

major comments (5)
  1. [Section 6.1 and 6.2.2] The central claim of the paper—that wearing SSTG produces genuine first-person insight into AI's bias and semantic mediation—rests on interviews that are structurally vulnerable to demand characteristics. Participants were told the study's purpose and how the system works before the experience (Section 6.1), and face-to-face interviews were conducted by the first author. The thematic analysis in Section 6.2.2 is a narrative selection of quotes with no coding scheme, no inter-rater reliability, no member checking, and no control condition. Statements such as P1's 'I felt that I was subconsciously aware of algorithmic bias' cannot, under this design, be distinguished from participants echoing the experimenter's framing. This is load-bearing because the abstract and Section 7.2 claim a subjective understanding as the main contribution. The authors should add an independent or blinded interviewer, a control condition (e.g., a human captioner or a different re-depiction modality), and a structured analysis plan for the qualitative data, or explicitly reframe the workshop as an exploratory pilot.
  2. [Sections 4.2, 5.1, 7.1] The caption length parameter is inconsistent across the manuscript. Section 4.2 says captions are limited to 'between 20 and 40 words,' while Sections 5.1 and 7.1 say '20-50 characters.' Since the caption length determines the grain of the linguistic bottleneck and thus the perceptual experience, this discrepancy must be resolved; if the actual constraint was a character count, Section 4.2 is incorrect, and if it was a word count, Sections 5.1 and 7.1 are incorrect. The resolution matters because the parameter defines what information is retained or discarded in the pipeline.
  3. [Sections 5.4 and 6] The manuscript states in Section 6 that the workshop was conducted with 'the same participants as the preliminary experiment,' reporting N=9 with ages 23–35 (mean 25, SD 3.77), while Section 5.4 reports N=24 respondents with ages 22–25 (20 male, 4 female; 23 Japanese, 1 Chinese). These are incompatible descriptions. The mismatch undermines the credibility of both sets of data and must be clarified: are these two separate participant pools, or does the text incorrectly describe an overlap? Without this clarification, the qualitative results cannot be linked to the quantitative screening as claimed.
  4. [Sections 5.2 and 5.3] The quantitative evaluation uses paired t-tests on four linguistic and four visual metrics without correction for multiple comparisons, and the only baseline is random pairing of images or sentences. For the visual metrics, the evidence for preservation is weak: SIFT similarity has Cohen's d = 0.19, and LPIPS absolute scores are close between the paired and random conditions (0.64 vs. 0.69 for AlexNet). The paper's conclusion that the transformation 'retains a certain amount of higher-order semantic information' (Section 5.2) would be more convincing with confidence intervals, a multiple-comparison correction, and an additional control condition such as direct image-to-image translation without a text bottleneck, which would isolate the effect of the linguistic mediation.
  5. [Section 7.3] The definition of Linguistic Virtual Reality relies on an equivalence relation ('realities that reduce to the same sentence are equivalent') without specifying what counts as 'the same sentence'—e.g., exact string identity, paraphrase, translation, or semantic similarity—or how the equivalence classes are constructed. The tree/tower example is illustrative but not operationalized, and the relation between this proposal and existing notions of semantic similarity in NLP is not discussed. As a conceptual contribution, the notion needs more formal grounding to be usable by other researchers.
minor comments (6)
  1. [Throughout] The manuscript contains numerous ACM template placeholders, including 'Conference acronym ’XX,' 'Do Not Use This Code' in the CCS Concepts and Keywords sections, and 'Make sure to enter the correct conference title from your rights confirmation email' in the ACM Reference Format. These must be removed and replaced with the journal's actual metadata.
  2. [Section 6.2.1] The heading 'Symboric behaviors' appears to be a typo for 'Symbolic behaviors.'
  3. [Figures 4 and 5] The panels in Figures 4 and 5 use mismatched vertical axis ranges; the authors note this, but it makes direct comparison of distributions between conditions difficult, and the definition of 'Outliers are defined as values in the 99th percentile' is ambiguous.
  4. [Section 7.3] The text refers to items (1)–(6) in Figure 7, but that figure is not included in the submitted manuscript, making the sensory-versus-linguistic equivalence argument difficult to follow.
  5. [References] The reference list has inconsistent formatting; for example, reference [43] lists only the surname 'Reimers, N.' without the full author team, and several conference bibliographies are missing page numbers or DOI information.
  6. [Section 1] The phrase 'At the same time, It also attempts' has an unnecessary capitalization of 'It' and a comma splice; please edit for style.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the empirical comparisons use external datasets and random baselines, and the conceptual VR claim is an explicit definition rather than a derived prediction.

full rationale

The derivation chain is self-contained. The framework in Section 3 is an explicit procedural definition of Semantic See-through Goggles, not a result derived from the later studies. Section 5 tests the system's transformation properties against independent baselines: paired input/output captions and images are compared with random pairings using external metrics (TF-IDF, WMD, USE, SBERT, HI, SIFT, LPIPS) on ILSVRC DET images, and no parameter is fitted to the reported outcomes, so the similarity findings are not forced by construction. The workshop evidence in Section 6 is qualitative self-report, which has validity limitations such as demand characteristics, a single interviewer, and no control condition, and the paper itself concedes in Section 7.2.1 that it is an oversimplification to call the experience 'the experience of seeing the world as an AI in general'; however, these are evidentiary weaknesses, not circular reasoning. The proposed Linguistic/Semantic Virtual Reality in Section 7.3 is introduced as a definitional equivalence relation based on 'the identity of "the sentence when we put it into words"', and the workshop outcomes are not presented as a proof of that definition. The only author self-citation is [37] Lived Montage, used in Section 2.2 as related work to broaden the notion of see-through goggles; it is not load-bearing for the central claim. No quoted prediction reduces to its input, and no load-bearing argument rests on a self-citation chain.

Assumptions & free parameters 3 free parameters · 4 assumptions · 1 invented entities

The central claim rests on a small number of free parameters chosen by hand (caption length, inference steps, resolution), on domain assumptions about language and AI being adequate proxies, and on the author-proposed concept of semantic virtuality. No derivation is attempted; the paper is an empirical and conceptual exploration. The least independent assumption is that the specific models and the qualitative self-reports are representative and valid.

free parameters (3)
  • Caption length limit = 20-50 characters (also stated as 20-40 words in Section 4.2)
    Chosen by hand to balance detail and speed, and it defines the boundary of what meaning is preserved.
  • LCM inference steps = 4 steps
    Chosen for real-time speed; the authors note it affects the level of visual detail.
  • Image resolution = 640x640 for prototype, 256x256 for preliminary study
    Standardized for processing, affects the information fed into the captioning model.
assumptions (4)
  • domain assumption Language mediates sensory information and necessarily introduces selection, reduction, imposition, and bias.
    Stated in the Introduction as the motivating premise for the whole project.
  • domain assumption BLIP and LCM are adequate representatives of AI's linguistic mediation of visual perception.
    The implementation uses these two models as the 'writer' and 'painter'; no comparison with other models is made in Section 4.
  • domain assumption Participants' verbal reports are valid evidence of their subjective experience.
    The workshop analysis in Section 6 treats interview quotes as direct evidence without independent coding or controls.
  • ad hoc to paper Realities that reduce to the same sentence are equivalent for the purpose of defining a virtual reality.
    Proposed in Section 7.3 as the definition of Linguistic/Semantic Virtual Reality; it is not derived from external principles.
invented entities (1)
  • Semantic/Linguistic Virtual Reality
    purpose: To define a new equivalence relation, 'same sentence implies same reality', as a basis for a type of VR.
    Defined in Section 7.3 with no falsifiable prediction or measurable consequence outside the paper; it is a philosophical construct.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SEMANTIC SEE-THROUGH GOGGLES: Wearing Linguistic Virtual Reality in (Artificial) Intelligence." pith.science (2026). https://pith.science/paper/I4PIWTF7

@misc{pith2026241202641,
  author       = {Pith},
  title        = {Pith review of: SEMANTIC SEE-THROUGH GOGGLES: Wearing Linguistic Virtual Reality in (Artificial) Intelligence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/I4PIWTF7}},
  note         = {Machine review of arXiv:2412.02641}
}
read the original abstract

When language is utilized as a medium to store and communicate sensory information, there arises a kind of radical virtual reality, namely "the realities that are reduced into the same sentence are virtual/equivalent." In the current era, in which artificial intelligence engages in the linguistic mediation of sensory information, it is imperative to re-examine the various issues pertaining to this potential VR, particularly in relation to bias and (dis)communication. Semantic See-through Goggles represent an experimental framework for glasses through which the view is fully verbalized and re-depicted into the wearer's view. The participants wear the goggles equipped with a camera and head-mounted display (HMD). In real-time, the image captured by the camera is converted by the AI into a single line of text, which is then transformed into an image and presented to the user's eyes. This process enables users to perceive and interact with the real physical world through this redrawn view. We constructed a prototype of these goggles, examined their fundamental characteristics, and then conducted a qualitative analysis of the wearer's experience. This project investigates a methodology for subjectively capturing the situation in which AI serves as a proxy for our perception of the world. At the same time, It also attempts to appropriate some of the energy of today's debate over artificial intelligence for a classical inquiry around the fact that "intelligence can only see the world under meaning."

Figures

Figures reproduced from arXiv: 2412.02641 by the authors.

Figure 1
Figure 1. Top: The framework of Semantic See-through Goggles. Bottom-left: The prototype of Semantic See-though Goggles. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. See-through HMD is a goggle to see what is in front [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 4
Figure 4. Comparison of linguistic similarities using different [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figures from the paper (3 more)
Figure 5
Figure 5. Figure 5: Comparison of visual similarities using different [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: The workshop view. Participants wear Semantic See-through Goggles and observe, walk, and interact with the [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: sketches a comparison between sensory virtual reality in vision (investment diagram method and HMD-like virtual reality) and linguistic virtual reality. (1)-(3) are sensory equivalents of each other, since the image on the retina does not change. However, (4) is not eq…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

58 extracted references · 31 canonical work pages

  1. [1]

    Enhancing image captioning with depth information using a Transformer-based framework

    Aya Mahmoud Ahmed, Mohamed Yousef, K. Hussain, and Yousef B. Mahdy. 2023. Enhancing image captioning with depth information using a Transformer-based framework. ArXiv abs/2308.03767 (2023). https://doi.org/10.48550/arXiv.2308. 03767

  2. [2]

    Phyllis Ang, Bhuwan Dhingra, and Lisa Wu Wills. 2022. Characterizing the efficiency vs. accuracy trade-off for long-context NLP models. arXiv preprint arXiv:2204.07288 (2022)

  3. [3]

    Ronald T Azuma. 1997. A survey of augmented reality. Presence: teleoperators & virtual environments 6, 4 (1997), 355–385

  4. [4]

    Michael Bajura, Henry Fuchs, and Ryutarou Ohbuchi. 1992. Merging virtual objects with the real world: Seeing ultrasound imagery within the patient. ACM SIGGRAPH Computer Graphics 26, 2 (1992), 203–210

  5. [5]

    ImageCaptioner$^2$: Image Captioner for Image Captioning Bias Amplification Assessment

    Eslam Mohamed Bakr, Pengzhan Sun, Erran L. Li, and Mohamed Elhoseiny. 2023. ImageCaptioner2: Image Captioner for Image Captioning Bias Amplification Assessment. ArXiv abs/2304.04874 (2023). https://doi.org/10.48550/arXiv.2304. 04874

  6. [6]

    Seth D Baum, Ben Goertzel, and Ted G Goertzel. 2011. How long until human- level AI? Results from an expert assessment. Technological Forecasting and Social Change 78, 1 (2011), 185–195

  7. [7]

    Shruti Bhargava and David Forsyth. 2019. Exposing and correcting the gender bias in image captioning datasets and models. arXiv preprint arXiv:1912.00578 (2019)

  8. [8]

    Pierre Bourdieu. 1991. Language and symbolic power. Polity (1991)

Show all 58 references
  1. [9]

    Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, et al. 2018. Universal sentence encoder for English. In Proceedings of the 2018 conference on empirical methods in natural language ...

  2. [10]

    Jacob Cohen. 2013. Statistical power analysis for the behavioral sciences. routledge

  3. [11]

    Dabrowski and Ethan V

    James R. Dabrowski and Ethan V. Munson. 2001. Is 100 Milliseconds Too Fast?. In CHI ’01 Extended Abstracts on Human Factors in Computing Systems (Seattle, Washington) (CHI EA ’01). Association for Computing Machinery, New York, NY, USA, 317–318. https://doi.org/10.1145/634067.634255

  4. [12]

    Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. 2022. Cogview2: Faster and better text-to-image generation via hierarchical transformers. Advances in Neural Information Processing Systems 35 (2022), 16890–16902

  5. [13]

    Julia Dressel and Hany Farid. 2018. The accuracy, fairness, and limits of predicting recidivism. Science advances 4, 1 (2018), eaao5580

  6. [15]

    Lianli Gao, Zhao Guo, Hanwang Zhang, Xing Xu, and Heng Tao Shen. 2017. Video Captioning With Attention-Based LSTM and Semantic Consistency. IEEE Transactions on Multimedia 19 (2017), 2045–2055. https://doi.org/10.1109/TMM. 2017.2729019

  7. [16]

    Abid Haleem, Mohd Javaid, Mohd Asim Qadri, Ravi Pratap Singh, and Rajiv Suman. 2022. Artificial intelligence (AI) applications for marketing: A literature- based study. International Journal of Intelligent Networks 3 (2022), 119–132

  8. [17]

    Lienhart, Carolin Kaiser, and René Schallner

    Philipp Harzig, Stephan Brehm, R. Lienhart, Carolin Kaiser, and René Schallner

  9. [18]

    Johann Gottfried Herder. 2019. Treatise on the Origin of Language

  10. [19]

    Grit- senko, Diederik P

    Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, A. Grit- senko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. 2022. Imagen Video: High Definition Video Generation with Diffusion Models. ArXiv abs/2210.02303 (2022). https:/...

  11. [20]

    James S Holmes. 1972. The cross-temporal factor in verse translation. Meta 17, 2 (1972), 102–110

  12. [21]

    He-Yen Hsieh, Jenq-Shiou Leu, and Sheng-An Huang. 2019. Implementing a real- time image captioning service for scene identification using embedded system. Multimedia Tools and Applications 80 (2019), 12525 – 12537. https://doi.org/10. 1007/s11042-020-10292-y

  13. [22]

    Thiruvathukal, and Ming Yin

    Xiao Hu, Haobo Wang, Anirudh Vegesana, Somesh Dube, Kaiwen Yu, Gore Kao, Shuo-Han Chen, Yung-Hsiang Lu, G. Thiruvathukal, and Ming Yin. 2020. Crowdsourcing Detection of Sampling Biases in Image Datasets. Proceedings of The Web Conference 2020 (2020). https://doi.org/10.1145/33...

  14. [23]

    Benjamin Jowett et al. 1892. Charmides. Lysis. Laches. Protagoras. Euthydemus. Cratylus. Phaedrus. Ion. Symposium . Vol. 1. Oxford University Press, American branch

  15. [24]

    Jungo Kasai, Keisuke Sakaguchi, Lavinia Dunagan, Jacob Daniel Morrison, Ro- nan Le Bras, Yejin Choi, and Noah A. Smith. 2021. Transparent Human Evaluation for Image Captioning. ArXiv abs/2111.08940 (2021). https://doi.org/10.18653/v1/ 2022.naacl-main.254

  16. [25]

    Yeonju Kim, Junho Kim, Byung-Kwan Lee, Sebin Shin, and Yong Man Ro. 2023. Mitigating dataset bias in image captioning through clip confounder-free cap- tioning network. In 2023 IEEE International Conference on Image Processing (ICIP) . IEEE, 1720–1724

  17. [26]

    Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger. 2015. From word embeddings to document distances. In International conference on machine learn- ing. PMLR, 957–966

  18. [27]

    Xin, and Stephen Westland

    S Lee, John H. Xin, and Stephen Westland. 2005. Evaluation of Image Similarity by Histogram Intersection. Color Research and Application 30 (2005), 265–274. https://api.semanticscholar.org/CorpusID:54697123

  19. [28]

    André Lefevere. 2016. Translation, rewriting, and the manipulation of literary fame. Routledge

  20. [29]

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation. https://doi.org/10.48550/ARXIV.2201.12086

  21. [30]

    Jun Li, Chenyang Zhang, Wei Zhu, and Yawei Ren. 2024. A Comprehensive Survey of Image Generation Models Based on Deep Learning. Annals of Data Science (2024), 1–30

  22. [31]

    Zhiheng Li and Chenliang Xu. 2021. Discover the Unknown Biased Attribute of an Image Classifier. 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (2021), 14950–14959. https://doi.org/10.1109/ICCV48922.2021.01470

  23. [32]

    Simian Luo, Yiqin Tan, Longbo Huang, Jian Li, and Hang Zhao. 2023. Latent Con- sistency Models: Synthesizing High-Resolution Images with Few-Step Inference. arXiv:2310.04378 [cs.CV]

  24. [33]

    Burak Makav and V. Kılıç. 2019. Smartphone-based Image Captioning for Vi- sually and Hearing Impaired. 2019 11th International Conference on Electrical and Electronics Engineering (ELECO) (2019), 950–953. https://doi.org/10.23919/ ELECO47770.2019.8990395

  25. [34]

    Ludovica Marinucci, Claudia Mazzuca, and Aldo Gangemi. 2023. Exposing implicit biases and stereotypes in human and artificial intelligence: state of the art and challenges with a focus on gender. AI & SOCIETY 38, 2 (2023), 747–761

  26. [35]

    Marshall McLuhan. 2011. The gutenberg galaxy. University of Toronto Press

  27. [36]

    Cristian Muñoz, Sara Zannone, Umar Mohammed, and Adriano Koshiyama. 2023. Uncovering bias in face generation models. arXiv preprint arXiv:2302.11562 (2023)

  28. [37]

    Goki Muramoto. 2020. Lived Montage. https://www.goki-muramoto.com/ lemontagevecu

  29. [38]

    Edwin G Ng, Bo Pang, Piyush Sharma, and Radu Soricut. 2020. Understand- ing guided image captioning performance across domains. arXiv preprint arXiv:2012.02339 (2020)

  30. [39]

    A Michael Noll. 1994. The beginnings of computer art in the United States: A memoir. Leonardo (1994), 39–44

  31. [40]

    Eirini Ntoutsi, Pavlos Fafalios, Ujwal Gadiraju, Vasileios Iosifidis, Wolfgang Nejdl, Maria-Esther Vidal, Salvatore Ruggieri, Franco Turini, Symeon Papadopoulos, Emmanouil Krasanakis, et al . 2020. Bias in data-driven artificial intelligence systems—An introductory survey. Wil...

  32. [41]

    Heikkila, and Shin’ichi Satoh

    Mayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami, Yuta Nakashima, Esa Rahtu, J. Heikkila, and Shin’ichi Satoh. 2023. Toward Verifiable and Reproducible Human Evaluation for Text-to-Image Generation. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)...

  33. [42]

    Ville Paananen, Jonas Oppenlaender, and Aku Visuri. 2023. Using text-to-image generation for architectural design ideation. International Journal of Architectural Computing (2023), 14780771231222783. Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Muramoto et al

  34. [43]

    N Reimers. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT- Networks. arXiv preprint arXiv:1908.10084 (2019)

  35. [44]

    Ehud Reiter. 2018. A structured review of the validity of BLEU. Computational Linguistics 44, 3 (2018), 393–401

  36. [45]

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10684–10695

  37. [46]

    Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al

  38. [47]

    Himanshu Sharma and Devanand Padha. 2023. A comprehensive survey on image captioning: from handcrafted to deep learning-based techniques, a taxonomy and open research issues. Artificial Intelligence Review 56, 11 (2023), 13619–13661

  39. [48]

    George M Stratton. 1896. Some preliminary experiments on vision without inversion of the retinal image. Psychological review 3, 6 (1896), 611

  40. [49]

    Mathews, and Lexing Xie

    Alasdair Tran, A. Mathews, and Lexing Xie. 2020. Transform and Tell: Entity- Aware News Image Captioning. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020), 13032–13042. https://doi.org/10.1109/ CVPR42600.2020.01305

  41. [50]

    Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE confer- ence on computer vision and pattern recognition . 4566–4575

  42. [51]

    Jesse Vig. 2019. Visualizing Attention in Transformer-Based Language Represen- tation Models. ArXiv abs/1904.02679 (2019)

  43. [52]

    Sandra Wachter, Brent Mittelstadt, and Chris Russell. 2021. Why fairness cannot be automated: Bridging the gap between EU non-discrimination law and AI. Computer Law & Security Review 41 (2021), 105567

  44. [53]

    Josiah Wang, Fei Yan, Ahmet Aker, and Robert Gaizauskas. 2014. A Poodle or a Dog? Evaluating Automatic Image Annotation Using Human Descriptions at Different Levels of Granularity. In Proceedings of the Third Workshop on Vision and Language, Anja Belz, Darren Cosker, Frank Kel...

  45. [54]

    Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li. 2023. Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.ArXiv abs/2306.09341 (2023). https://doi.org/10.48550/arXiv.2306.09341

  46. [55]

    Jingbo Zhang, Xiaoyu Li, Ziyu Wan, Can Wang, and Jing Liao. 2024. Text2nerf: Text-driven 3d scene generation with neural radiance fields. IEEE Transactions on Visualization and Computer Graphics (2024)

  47. [56]

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang

  48. [59]

    In Proceedings of the IEEE conference on computer vision and pattern recognition

    The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition . 586–595. Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009

  49. [2015]

    International journal of computer vision 115 (2015), 211–252

    Imagenet large scale visual recognition challenge. International journal of computer vision 115 (2015), 211–252

  50. [2018]

    https: //doi.org/10.1109/MIPR.2018.00035

    Multimodal Image Captioning for Marketing Analysis.2018 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR) (2018), 158–161. https: //doi.org/10.1109/MIPR.2018.00035

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.