Pith. sign in

REVIEW 3 major objections 2 cited by

Scientific poster generation becomes measurable when image models design only the layout and leave blank, labeled figure slots for real source assets.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 05:29 UTC pith:VDUXYZHG

load-bearing objection Solid systems paper: placeholder-first makes scientific-figure provenance auditable and ships real artifacts; the 34 o0 probe is only as strong as a same-family VLM count, but the authors already scope it as pilot and the deterministic half is clean. the 3 major comments →

arxiv 2607.03006 v1 pith:VDUXYZHG submitted 2026-07-03 cs.CV cs.AIcs.SE

PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark

classification cs.CV cs.AIcs.SE
keywords scientific poster generationplaceholder-first contractinstruction followingsource-figure provenancetext-rich image modelsvisual summarizationauditable benchmarkfailure taxonomy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Text-rich image models can compose poster-scale layouts, but scientific communication still needs legible labels, correct aspect ratios, and—above all—no fabricated data plots. This paper introduces a harness that reframes poster generation as a set of auditable instruction-following checks rather than free image synthesis. The model is allowed only visual-summary design (typography, reading path, color, background) and must leave every scientific figure region as an empty labeled placeholder; a deterministic compositor later pastes the paper’s real figures at detected coordinates. Failures become explicit rejections instead of plausible-looking fakes. On a 12-paper pilot the contract drives synthesized scientific figures from 34 to 0 in a small counterfactual probe, surfaces a concrete failure taxonomy, and yields a clear trade-off against a deterministic slide-style baseline: richer visuals and stronger visual preference scores, slightly less retained quiz-style information, and higher runtime.

Core claim

The authors claim that a placeholder-first contract turns open-ended scientific poster generation into a battery of measurable instruction-following tasks. By forcing the image model to leave every data-bearing figure region as a blank labeled box (ID, short label, target aspect ratio) and inserting real source-paper figures only after deterministic detection and geometry checks, the system makes placeholder count and ID accuracy, blankness, aspect-ratio compliance, abstention from synthesized graphics, public-text hygiene, and source-figure provenance into logged pass/fail conditions rather than hidden visual judgments.

What carries the argument

The placeholder-first contract: the generative model designs global composition, typography, color, and background but is forbidden to draw any data-bearing content inside figure slots; a deterministic replacement engine detects those slots and pastes real source assets at locked coordinates and aspect ratios, with multi-stage QA logging every rejection.

Load-bearing premise

The method treats vision-language model critics, placeholder detectors, and preference judges—often from the same model family used for generation—as reliable enough proxies for the layout contracts and for visual quality on a small pilot set.

What would settle it

Repeat the three-paper no-placeholder counterfactual on a larger budget-matched paper set with independent multi-judge or human audits: if unconstrained generation no longer produces countable synthesized scientific figures, or if the explicit placeholder contract fails to keep figure slots blank, correctly labeled, and free of data-like decoration, the central architectural claim does not hold.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Accepted scientific figure regions on generated posters can be audited for source provenance instead of being treated as possible hallucinations.
  • Failures of aspect ratio, blankness, ID labeling, and public-text hygiene become explicit rejection categories rather than silent visual defects.
  • Declarative domain profiles and density presets can encode field-specific narrative spines and forbidden decorations without rewriting the replacement engine.
  • The same pipeline can serve as a one-page visual reading interface that preserves figure provenance for arXiv-scale papers.
  • Released prompts, manifests, and audit scripts become a reusable evaluation component for text-rich image models on scientific layout contracts.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the contract generalizes, similar blank-slot plus deterministic insertion patterns could audit other knowledge-intensive layouts (slides, review figures, grant visuals) where free synthesis is scientifically unacceptable.
  • The observed visual-richness versus information-density trade-off suggests that specialist dense modes will need separate factuality metrics, not only aesthetic VLM preference.
  • Self-preference risk in same-family VLM judges implies that future harnesses will need external or multi-model raters before preference scores can be treated as more than pilot regime signals.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. PosterHarness reframes scientific poster generation as an auditable instruction-following problem for text-rich image models. The core mechanism is a placeholder-first contract: the image model designs typography, reading path, color, and background, but every scientific figure region must be an empty labeled placeholder (ID, short label, aspect ratio); a deterministic compositor then inserts real source-paper figures at detected coordinates. This is intended to make placeholder count/ID accuracy, blankness, aspect-ratio compliance, abstention from synthesized scientific graphics, public-text hygiene, and source-figure provenance measurable, with failures logged as explicit rejections. The paper instantiates the harness on a 12-paper pilot (6 HEP, 6 AI/ML-adjacent), reports a three-paper counterfactual prompt probe (VLM-counted synthesized figures 34 o0 under the explicit contract), a run-level failure taxonomy from manifests (geometry, placeholder QA, template critic, public text), and a trade-off comparison with Paper2Poster (higher resolution, lower white-canvas fraction, stronger single-VLM preference for PosterHarness; slightly higher PosterQuiz-style information and lower runtime for Paper2Poster). Claims are scoped as pilot/regime characterization, not superiority; code, prompts, manifests, and audit scripts are released.

Significance. If the architectural split holds under stronger measurement, the paper contributes a reusable evaluation component for scientific layout contracts rather than another unconstrained poster generator. The placeholder-first separation, deterministic replacement, run manifests, and failure taxonomy are concrete engineering contributions that make compliance failures visible instead of burying them in plausible output. The open release of paired posters, prompts, and audit scripts is a genuine strength for a systems paper in this area. The work is timely given recent text-rich image models and the documented unreliability of model-synthesized scientific figures. Significance is currently limited by pilot scale and heavy reliance on same-family VLM judges for the soft contracts and preference metrics, but the framing as an auditable harness rather than a production authoring claim is appropriate and useful.

major comments (3)
  1. §5.4 and Table 7: the load-bearing causal claim that the placeholder contract drives synthesized scientific figures from 34 to 0 rests on a single VLM auditor counting “synth. figs” and “data-like decor.” under four prompt conditions on only three papers. The same model family supplies the template critic thresholds (§3.5) and the blinded preference wins (20/24 in Table 4). §5.7 and Appendix E already flag single-judge self-preference risk and unfinished multi-judge / critic-disabled ablations. Without an independent second auditor (different VLM family or human count of residual data-bearing decoration) or a larger budget-matched no-placeholder set, the 34 o0 number and the visual-preference trade-off remain only weakly independent of the generation stack. This is the central measurement gap for the architectural claim.
  2. Table 4 and §5.1–5.2: the consolidated 12-paper comparison is carefully scoped as regime characterization, but several free parameters (critic thresholds 0.72/0.65/0.65/0.75, geometry IoU 0.06 and center-distance 0.22, max_figures=8, detection confidence 0.15) are not ablated against acceptance rate or failure categories. Appendix E lists critic-disabled and domain-gain ablations as not yet run. At minimum, report sensitivity of acceptance and failure taxonomy to the critic thresholds and the figure-count budget, or clearly demote any claim that the multi-stage QA chain is necessary rather than merely present in the pilot configuration.
  3. §3.1 R1 and §3.3 source-insertion invariant: source-figure provenance is correctly narrowed to accepted figure regions conditional on extraction, assignment, and detection, and the paper states it is not full poster factuality. However, the main results still lead with “source-figure provenance proxy 12/12” for both systems (Table 4) without a crop-level similarity or assignment-error audit (reserved for future work). For the pilot claim to be load-bearing, either add a simple deterministic check that each pasted region matches the assigned source asset (hash/SSIM on crops) or further qualify the proxy language in the abstract and Table 4 so readers do not over-read provenance as end-to-end figure correctness.

Circularity Check

0 steps flagged

No derivation circularity: empirical systems paper with external metrics and self-flagged VLM-judge limits, not a prediction that reduces to its inputs.

full rationale

PosterHarness is an engineering/systems paper, not a first-principles derivation. Its central claim is architectural (placeholder-first split + deterministic insertion) and is supported by (i) deterministic geometry/containment/provenance checks, (ii) a small prompt-condition probe counting synthesized figures, (iii) run-manifest failure taxonomy, and (iv) multi-metric comparison to Paper2Poster (resolution, white-canvas fraction, PosterQuiz-style scores, runtime). None of these reduce by construction to fitted parameters or to a self-citation uniqueness chain. There is no self-definitional loop (X defined as Y then used to predict Y), no fitted input renamed as prediction, no load-bearing uniqueness theorem imported from the same authors, and no ansatz smuggled via self-citation. Shared-family VLM critics/preference judges are a validity threat the paper itself flags (§5.7, Appendix E) and reports as pilot/single-judge evidence—not a circular derivation. Score 0 is the honest finding under the analyzer rules.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 3 invented entities

Load-bearing content is mostly engineering contracts and hand-chosen QA thresholds rather than physical free parameters. The central claim rests on domain assumptions about scientific communication (figures must be source-grounded; models must not invent data-bearing plots) and on operational thresholds that gate acceptance. Invented entities are the harness constructs themselves; they have independent operational handles via released manifests and geometry checks.

free parameters (4)
  • Template critic score thresholds = 0.72 / 0.65 / 0.65 / 0.75
    Overall 0.72, artistry 0.65, information density 0.65, placeholder-contract 0.75 are hand-chosen gates that control regeneration and acceptance.
  • Placeholder geometry rejection thresholds = IoU 0.06; center-distance 0.22; 20% AR tolerance
    IoU < 0.06 and center-distance fraction > 0.22 reject detections; 20% aspect-ratio tolerance and square band 0.9–1.1 are chosen cutoffs.
  • max_figures layout budget = 8
    Default cap of eight placeholders is an empirical stability choice, not a derived optimum.
  • Detection confidence minimum = 0.15
    Minimum detection confidence 0.15 for placeholder boxes is a hand-set filter.
axioms (4)
  • domain assumption Scientific figure regions on accepted posters must be populated by real source-paper assets, not model-synthesized plots.
    Stated as R1 and the source-insertion invariant; foundational to the entire harness design.
  • domain assumption Text-rich image models are comparatively suited to global composition/typography but untrustworthy for data-bearing scientific figures.
    Motivated by cited evaluations of native image models on scientific illustrations; drives the architectural split.
  • ad hoc to paper Declarative domain profiles (narrative spines, figure hierarchy, decorative prohibitions) can encode field-specific poster conventions without code changes.
    Profiles are manually specified defaults, not empirically validated universal story structures; paper notes they are design defaults.
  • ad hoc to paper Deterministic geometry and containment checks plus VLM audits are sufficient to surface contract failures as explicit rejections.
    Operational assumption of the multi-stage QA chain; reliability of VLM judges is only partially stress-tested.
invented entities (3)
  • Placeholder-first contract independent evidence
    purpose: Force blank labeled figure slots so source figures can be inserted deterministically and compliance can be audited.
    Core mechanism of the paper; independent operational evidence via the three-paper probe and geometry/QA logs.
  • PosterHarness multi-stage QA chain and run manifests independent evidence
    purpose: Log retries, failure categories, and provenance so accepted posters remain auditable.
    Engineering construct enabling the benchmark claim; evidence is the released manifests and failure taxonomy.
  • Declarative domain profiles and density presets (e.g., hep dense) no independent evidence
    purpose: Adapt narrative spine, text priorities, and decoration bans across disciplines without pipeline code changes.
    New configuration layer; quantitative gain over generic prompting is explicitly left to future ablation.

pith-pipeline@v1.1.0-grok45 · 23636 in / 3399 out tokens · 34755 ms · 2026-07-12T05:29:40.322067+00:00 · methodology

0 comments
read the original abstract

Text-rich image models can now design poster-scale layouts, but we lack ways to measure whether they honor scientific communication contracts: legible labels, prescribed aspect ratios, and -- above all -- abstaining from fabricated scientific figures. We present POSTERHARNESS, an auditable harness reframing poster generation as measurable instruction-following tasks, with a pilot benchmark and failure taxonomy. POSTERHARNESS uses a placeholder-first contract to separate two jobs models otherwise conflate. The model performs visual-summary design: typography, reading path, color, and background -- but never draws data-bearing figures. Every figure region must be an empty labeled placeholder; a deterministic compositor inserts real source-paper figures at detected coordinates. This makes properties measurable: placeholder count and ID accuracy, blankness, aspect-ratio compliance, abstention from synthesized graphics, public-text hygiene, and source-figure provenance -- with failures logged as explicit rejections, not hidden in plausible-looking output. We instantiate the harness on 12 papers (6 HEP, 6 AI/ML-adjacent) and report three findings. (i) A counterfactual probe shows the placeholder contract drives VLM-counted synthesized figures from 34 to 0 across three papers. (ii) A failure taxonomy identifies blocking contracts: placeholder geometry, placeholder QA, template critic, and public text. (iii) Comparison with Paper2Poster shows a trade-off: PosterHarness yields higher-resolution artifacts, lower white-canvas fraction, and stronger VLM visual preference; the deterministic baseline retains slightly more PosterQuiz-style information and runs faster. We report this as regime characterization, not a superiority claim. All artifacts, prompts, manifests, and audit scripts are released as a reusable evaluation component.

Figures

Figures reproduced from arXiv: 2607.03006 by Dawei Fu, Linrui Chen, Qiang Li, Ruobing Jiang, Tianyi Yang, Youpeng Wu, Zijian Wang, Zixun Kou.

Figure 1
Figure 1. Figure 1: Overview of the POSTERHARNESS pipeline. Inputs are first converted into text and figure assets, then routed through domain-aware LLM planning. The upper half produces and audits a placeholder template: the image model may design the visual layout but must leave labeled blank figure slots. The lower half consumes the detected slot geometry: deterministic insertion places real source figures, followed by con… view at source ↗
Figure 2
Figure 2. Figure 2: Placeholder template (left) and final poster after source-figure insertion (right). [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Micro-repair example: a localized header defect is corrected via targeted image edit. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Representative visual result from the no-placeholder probe on the SSWW HEP paper. This figure isolates [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Representative side-by-side comparisons from the selected benchmark. Paper2Poster produces clean [PITH_FULL_IMAGE:figures/full_fig_p015_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

    cs.CV 2026-07 conditional novelty 6.0

    A five-skill agent pipeline with one shared paper extractor and hard render gates produces editable posters, videos, and bilingual blogs, leading the Paper2Poster benchmark on aesthetics.

  2. ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

    cs.CV 2026-07 conditional novelty 5.0

    A five-skill agent pipeline generates an editable poster, video deck, and bilingual blog from a paper PDF, binds them in an interactive viewer, and reports poster scores above the authors' own under two VLM judges.

Reference graph

Works this paper leans on

29 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    Paper2poster: Towards multimodal poster automation from scientific papers.arXiv preprint arXiv:2505.21497, 2025

    Wei Pang, Kevin Qinghong Lin, Xiangru Jian, Xi He, and Philip Torr. Paper2poster: Towards multimodal poster automation from scientific papers.arXiv preprint arXiv:2505.21497, 2025

  2. [2]

    Autogen: Enabling next-gen llm applications via multi-agent conversation.arXiv preprint arXiv:2308.08155, 2023

    Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang. Autogen: Enabling next-gen llm applications via multi-agent conversation.arXiv preprint arXiv:2308.08155, 2023

  3. [3]

    Metagpt: Meta programming for a multi-agent collaborative framework.arXiv preprint arXiv:2308.00352, 2023

    Sirui Hong, Mingchen Zhuge, Jiaqi Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Jinlin Wang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and J¨urgen Schmid- huber. Metagpt: Meta programming for a multi-agent collaborative framework.arXiv preprint arXiv:2308.00352, 2023

  4. [4]

    CAMEL: Communicative agents for ”mind” exploration of large language model society

    Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. CAMEL: Communicative agents for ”mind” exploration of large language model society. InAdvances in Neural Information Processing Systems, 2023

  5. [5]

    Textdiffuser: Diffusion models as text painters

    Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui, Qifeng Chen, and Furu Wei. Textdiffuser: Diffusion models as text painters. InAdvances in Neural Information Processing Systems, 2023

  6. [6]

    Anytext: Multilingual visual text generation and editing.arXiv preprint arXiv:2311.03054, 2023

    Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng, and Xuansong Xie. Anytext: Multilingual visual text generation and editing.arXiv preprint arXiv:2311.03054, 2023

  7. [7]

    Glyphcontrol: Glyph conditional control for visual text generation

    Yukang Yang, Dongnan Gui, Yuhui Yuan, Weicong Liang, Haisong Ding, Han Hu, and Kai Chen. Glyphcontrol: Glyph conditional control for visual text generation. InAdvances in Neural Information Processing Systems, 2023

  8. [8]

    Textdiffuser-2: Unleashing the power of language models for text rendering

    Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui, Qifeng Chen, and Furu Wei. Textdiffuser-2: Unleashing the power of language models for text rendering. InEuropean Conference on Computer Vision (ECCV), pages 386–402, 2024. doi: 10.1007/978-3-031-72652-1 23

  9. [9]

    Type-R: Automatically retouching typos for text-to-image generation

    Wataru Shimoda, Naoto Inoue, Daichi Haraguchi, Hayato Mitani, Seiichi Uchida, and Kota Yamaguchi. Type-R: Automatically retouching typos for text-to-image generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2745–2754, 2025. doi: 10.1109/CVPR52734.2025. 00262

  10. [10]

    Preliminary explorations with GPT-4o(mni) native image generation.arXiv preprint arXiv:2505.05501, 2025

    Pu Cao, Feng Zhou, Junyi Ji, Qingye Kong, Zhixiang Lv, Mingjian Zhang, Xuekun Zhao, Siqi Wu, Yinghui Lin, Qing Song, and Lu Yang. Preliminary explorations with GPT-4o(mni) native image generation.arXiv preprint arXiv:2505.05501, 2025

  11. [11]

    Introducing 4o image generation, 2025

    OpenAI. Introducing 4o image generation, 2025. Product release. Published March 25, 2025. Accessed June 17, 2026.https://openai.com/index/introducing-4o-image-generation/

  12. [12]

    ChatGPT Images 2.0 system card, 2026

    OpenAI. ChatGPT Images 2.0 system card, 2026. System card. Published April 21, 2026. Accessed June 17, 2026.https://deploymentsafety.openai.com/chatgpt-images-2-0

  13. [13]

    Image generation, 2026

    OpenAI. Image generation, 2026. API documentation. Accessed June 17, 2026. https://developers.openai. com/api/docs/guides/image-generation

  14. [14]

    BizGenEval: A systematic benchmark for commercial visual content generation.arXiv preprint arXiv:2603.25732, 2026

    Yan Li, Zezi Zeng, Ziwei Zhou, Xin Gao, Muzhao Tian, Yifan Yang, Mingxi Cheng, Qi Dai, Yuqing Yang, Lili Qiu, Zhendong Wang, Zhengyuan Yang, Xue Yang, Lijuan Wang, Ji Li, and Chong Luo. BizGenEval: A systematic benchmark for commercial visual content generation.arXiv preprint arXiv:2603.25732, 2026

  15. [15]

    DreamPoster: A unified framework for image-conditioned generative poster design.arXiv preprint arXiv:2507.04218, 2025

    Xiwei Hu, Haokun Chen, Zhongqi Qi, Hui Zhang, Dexiang Hong, Jie Shao, and Xinglong Wu. DreamPoster: A unified framework for image-conditioned generative poster design.arXiv preprint arXiv:2507.04218, 2025

  16. [16]

    Wenxin Tang, Jingyu Xiao, Yanpei Gong, Fengyuan Ran, Tongchuan Xia, Junliang Liu, Man Ho Lam, Wenxuan Wang, and Michael R. Lyu. EfficientPosterGen: Semantic-aware efficient poster generation via token compression and accurate violation detection.arXiv preprint arXiv:2603.00155, 2026

  17. [17]

    PosterIQ: A design perspective benchmark for poster understanding and generation.arXiv preprint arXiv:2603.24078, 2026

    Yuheng Feng, Wen Zhang, Haodong Duan, and Xingxing Zou. PosterIQ: A design perspective benchmark for poster understanding and generation.arXiv preprint arXiv:2603.24078, 2026

  18. [18]

    PosterReward: Unlocking accurate evaluation for high-quality graphic design generation

    Jianyu Lai, Sixiang Chen, Jialin Gao, Hengyu Shi, Zhongying Liu, Fuxiang Zhai, Junfeng Luo, Xiaoming Wei, Lujia Wang, and Lei Zhu. PosterReward: Unlocking accurate evaluation for high-quality graphic design generation. arXiv preprint arXiv:2603.29855, 2026. 18 PosterHarness

  19. [19]

    Adding conditional control to text-to-image diffusion models

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. arXiv preprint arXiv:2302.05543, 2023

  20. [20]

    Prompt-to-prompt image editing with cross attention control.arXiv preprint arXiv:2208.01626, 2022

    Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross attention control.arXiv preprint arXiv:2208.01626, 2022

  21. [21]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-judge with MT-Bench and chatbot arena. InAdvances in Neural Information Processing Systems, 2023

  22. [22]

    SciBERT: A pretrained language model for scientific text.arXiv preprint arXiv:1903.10676, 2019

    Iz Beltagy, Kyle Lo, and Arman Cohan. SciBERT: A pretrained language model for scientific text.arXiv preprint arXiv:1903.10676, 2019

  23. [23]

    PubLayNet: Largest dataset ever for document layout analysis.arXiv preprint arXiv:1908.07836, 2019

    Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes. PubLayNet: Largest dataset ever for document layout analysis.arXiv preprint arXiv:1908.07836, 2019

  24. [24]

    Create image api reference, 2026

    OpenAI. Create image api reference, 2026. API reference. Accessed June 17, 2026. https://developers. openai.com/api/reference/resources/images/methods/generate. 19 PosterHarness A Paper List Table 8: Complete selected benchmark paper list. ID Domain arXiv Short title 01 HEP 1207.7235 CMS Higgs discovery 02 HEP 2206.08956 CMS same-sign WW / heavy Majorana ...

  25. [25]

    Randomize left/right order independently for each rater and paper

    Stimuli: for each paper, show the Paper2Poster and PosterHarness outputs as anonymized A/B images. Randomize left/right order independently for each rater and paper

  26. [26]

    HEP papers should include at least a small number of HEP-trained raters because SR/CR, fit, and limit information are specialist requirements

    Raters: recruit a mixed panel of domain experts and non-expert scientific readers. HEP papers should include at least a small number of HEP-trained raters because SR/CR, fit, and limit information are specialist requirements

  27. [27]

    Preference questions: ask which poster is more likely to attract a viewer, which is easier to skim, which appears more scientifically trustworthy, and which the rater would prefer to use for a poster session

  28. [28]

    For HEP, questions should cover process, luminosity/energy, event selection, fit/result, and supporting figures

    Factual utility questions: ask raters to answer short paper-specific questions from the poster alone, analogous to PosterQuiz-style scoring. For HEP, questions should cover process, luminosity/energy, event selection, fit/result, and supporting figures

  29. [29]

    Ties and unreadable cases should be recorded explicitly rather than forced into binary preference

    Analysis: report win rate with bootstrap confidence intervals over papers, mean Likert scores per criterion, and answer accuracy for the factual questions. Ties and unreadable cases should be recorded explicitly rather than forced into binary preference. This protocol separates visual preference from information transfer, which is important because the cu...