{"id":"8284094f-7343-43ee-b813-f513c3619d6e","arxiv_id":"2608.03357","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Text-to-image models systematically fail to apply an object's intrinsic frame of reference: mean final accuracy drops from 26.5% on camera-view prompts to 15.4% on matched frame-of-reference prompts, and the failure peaks when the anchor's left or right is reversed relative to the image.","lead":"This paper introduces FoR-T2I, a benchmark of 1,200 prompt pairs that asks whether text-to-image models can place objects using an object's own left, right, front, or back instead of the camera view. Across 22 models, accuracy on such frame-of-reference prompts falls by about 42% relative to matched camera-view prompts, and even the best model succeeds only 44.3% of the time.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FoR scoring depends on VLM orientation judgment with only 65% standalone human agreement; unmeasured per-condition evaluator bias could inflate the headline gap.","rationale":"The paper's central contribution is a measured deficit, so the validity of the measurement is the crux. I considered the wording-complexity confound raised by the reader; however, Section 6.2's preserved-mapping row (FoR 60.6 vs Cam 66.0, gap 5.3 points) shows that the same awkward FoR phrasing does not produce a large drop when the mapping is preserved, so wording complexity alone is unlikely to explain the 49.7-point reversed gap. The weaker link is the orientation-dependent FoR scoring. Table 5b's standalone VLM judge agreement of 65.0% is concerning because FoR geometry requires the judge to know which side of the anchor is 'left'; if orientation perception is poor or biased toward the image frame, FoR failures are over-counted and the gap is inflated. The paper does not report agreement split by prompt type or by frame-mapping condition, and the 87.1% overall agreement does not rule out systematic bias. The 74.1% literal-side failure pattern is strong and less dependent on orientation judgment, but it is not the headline number and is not separately human-validated. A human rescoring test on the existing annotation set would settle this. If the gap survives human scoring, I would accept the central claim as well-supported. Given the conditional verdict already placed, I do not change it.","tokens_in":15718,"tokens_out":8808,"duration_ms":83108,"concrete_test":"Use the existing 120-pair human annotations (Section 5.5) to recompute accuracy separately for Cam and FoR with human labels, and compare the human Cam–FoR gap to the automatic gap, both overall and for the reversed-mapping subset. If the human-labeled gap is materially smaller (e.g., more than 5 points) or the reversed-mapping FoR accuracy rises above about 30%, the reported deficit is partly an evaluator artifact. Also report VLM judge orientation-only agreement against humans on FoR prompts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline Cam–FoR gap is measured by an automatic evaluator whose FoR scoring requires judging the anchor's rendered orientation. Table 5b reports that the standalone VLM judge agrees with humans at only 65.0% on final correctness; the full pipeline reaches 87.1%, but orientation judgment is still delegated to Qwen3.6-27B. No per-condition (Cam vs FoR) agreement, no confidence intervals, and no human FoR-only accuracy are reported. If the VLM systematically misreads orientation in exactly the reversed-mapping cases that drive the 49.7-point gap, FoR accuracy would be underestimated and the central 41.8% relative deficit could be inflated. The paper's internal evidence that models place targets on the literal image side (74.1% vs 16.7%) mitigates this, but that analysis is separate from the headline final accuracy and is not human-validated. Because the central claim is a quantitative deficit, this measurement validity issue is load-bearing: an evaluator artifact could create the gap rather than T2I models' reference-frame failure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FoR-T2I, a benchmark of 1,200 matched camera-view (Cam) and frame-of-reference (FoR) prompt pairs across three task levels, and evaluates 22 closed- and open-source text-to-image models. The central empirical claim is that mean final accuracy is 41.8% lower on FoR prompts than on matched Cam prompts (26.5% vs 15.4%), with the best model reaching only 44.3% FoR accuracy. The paper decomposes the deficit by relation type, camera view, and frame mapping, reporting that the gap is largest when an anchor-relative left/right mapping is reversed relative to the image (49.7 points). It also proposes a training-free VLM-gated rewriting method that improves average FoR accuracy from 25.0% to 29.2% under a matched generation budget.","tokens_in":15712,"tokens_out":10314,"duration_ms":92610,"significance":"If the headline measurement is valid, FoR-T2I is a useful controlled instrument for separating viewer-centered from object-centered spatial instruction following, a distinction existing T2I benchmarks do not isolate. The evaluation is broad (22 models), the Cam-FoR pairing is a sound design, and the paper contains several falsifiable checks: the Level 3 'square of L1' prediction is rejected in the data, the preserved-mapping condition shows only a 5.3-point gap versus 49.7 under reversal, and the literal image-side placement analysis (74.1% vs 16.7%) provides a mechanism for the wrong-frame failure that is largely independent of the orientation judge. These features make the central claim credible and non-circular. The main risk is measurement validity of the automatic evaluator's orientation judgment, which the authors validate only at the aggregate level.","major_comments":[{"comment":"The automatic evaluator delegates orientation judgment to Qwen3.6-27B, whose standalone agreement with human final-correctness judgments is 65.0%; the full pipeline reaches 87.1%, but no per-condition agreement for Cam vs FoR or for the preserved/remapped/reversed mapping groups is reported. Because FoR scoring requires judging the anchor's rendered orientation, a condition-specific evaluator bias could inflate the 41.8% relative deficit and the 49.7-point reversal gap in §6.2. The literal image-side placement analysis in §6.3 (74.1% vs 16.7%) is good mitigating evidence, but it is not a substitute for per-condition evaluator validation. Please report per-condition human-evaluator agreement and, if feasible, human FoR-only orientation accuracy on the 120-pair validation set, or demonstrate that the main gaps persist under human scoring on that subset.","section":"§5.5, Table 5b"},{"comment":"The construction protocol states that paired prompts 'must differ only in the frame in which the direction word is resolved,' but the displayed FoR prompts are considerably longer and grammatically unnatural (e.g., 'The baristas is behind girl, opposite girl's own facing direction'). While the preserved-mapping control in §6.2 (5.3-point gap) is strong evidence that wording alone does not explain the reversal effect, the paper should explicitly quantify template length and lexical complexity across conditions and state the preserved-mapping result as a control for this confound.","section":"§3.2, Figure 3"},{"comment":"The claim that L3 FoR failure 'costs more than two independent conversions' rests on comparing observed L3 accuracy (10.5%) with the square of L1 accuracy (15.0%). However, the same independence model already overpredicts on Cam prompts (37.5% predicted vs 33.7% observed), so the FoR-specific excess is only about 0.7 points when measured as an absolute shortfall relative to the control condition. Please report the shortfall relative to the Cam baseline (or use a ratio/relative measure) before concluding the excess is specific to FoR.","section":"§5.2, Level 3 analysis"},{"comment":"The paper does not state whether the benchmark data (the 1,200 prompt pairs, the layout engine, and the evaluation code) will be released or where. As the contribution is a benchmark, this is essential for reproducibility and for the field to build on it; please include a clear data/code availability statement and a release plan.","section":"Benchmark availability"}],"minor_comments":[{"comment":"The label 'Avg.' is repeated for the geometry block, the orientation block, and the final strict-accuracy columns; the caption should state explicitly which averages are macro-averages over L1-L3 and which columns define final accuracy.","section":"Table 3"},{"comment":"There are several typos in displayed prompts: 'choopstics' in Figure 1, and 'A baristas is present' and 'The baristas is behind girl' in Figure 3; please proofread all example prompts.","section":"Figures 1 and 3"},{"comment":"The reference list contains two entries with identical titles and author lists for 'Wu et al. 2025a' and 'Wu et al. 2025b'; verify that these correspond to distinct models and cite the correct technical report for Qwen-Image-2512.","section":"References"},{"comment":"The model list refers to 'Seedream 5.0 Pro' but the tables use 'Seedream 5.0'; please unify the naming throughout.","section":"Model naming"},{"comment":"The text promises that 'Appendix reports our benchmark's statistics,' but no appendix is present in the manuscript; please include the appendix or remove the pointer.","section":"Section 3.2"},{"comment":"The text states that a language model produces K rewrites but never specifies K or the sampling parameters; please report these for reproducibility.","section":"Section 4"},{"comment":"The main results and mitigation results are point estimates without confidence intervals or significance tests; for the central 41.8% relative-deficit claim, at least standard errors on the macro-averaged accuracy would strengthen the paper.","section":"Tables 2 and 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript cites 2026 model versions and arXiv dates, which is unusual; the editor may wish to verify the accessibility of the cited models and the intended timeline. The ERNIE-Image model is from the authors' own team, which is not itself a problem, but its results should be interpreted with awareness of that association."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper introduces the first controlled benchmark that isolates frame of reference in text-to-image generation, and it reports a large, consistent deficit across 22 models, with the sharpest failure when an anchor-relative left/right maps to the opposite image side (18.8% vs 68.5% on matched camera-view prompts). I think the qualitative result—models default to the camera frame when the anchor's orientation disagrees—is credible and well evidenced. The headline magnitudes, though, are softer than the abstract implies.\n\nWhat's genuinely new: FoR-T2I is the first benchmark with matched Cam-FoR prompt pairs, three task levels including chained frame conversions, and a component split separating target placement from anchor orientation. That split is useful: the FoR deficit shows up almost entirely in geometry accuracy, not orientation accuracy, which argues against a trivial 'models can't draw oriented objects' explanation. The preserved/reversed mapping control in Section 6.2 is the paper's best evidence—when the two frames agree, the gap shrinks to 5.3 points; when they reverse, it jumps to 49.7. That pattern is hard to explain without real reference-frame confusion.\n\nThe soft spots, in rough order of importance. First, no code, benchmark, or appendix is shipped, so exact magnitudes can't be independently checked; for a benchmark paper that's a serious omission. Second, there are no confidence intervals or significance tests anywhere. The headline 41.8% relative drop is a mean over models, but the mitigation gain is 4.2 points on 120 prompts—that could be noise. Third, the FoR prompts are visibly longer and less natural than their Cam counterparts; the paper claims they 'differ only in the frame,' but the examples show syntax and detail also differ. The preserved-mapping control removes part of that confound, not all. Fourth, the evaluator's orientation judgment is delegated to a VLM whose standalone agreement with humans is only 65%, and per-condition (Cam vs FoR) agreement isn't reported. The stress-test worry that VLM bias could invent the gap isn't confirmed—Section 6.3's literal-side placement evidence (74.1% vs 16.7%) suggests models truly ignore orientation—but the missing per-condition validation is a real gap. Finally, Qwen is used for pool construction, evaluation, and the mitigation gate, a modest same-family coupling. I'd also flag the Level 3 'excess cost specific to FoR' claim as slightly overstated: the Cam condition shows a similar absolute shortfall against the independence model.\n\nTaken together, the central claim holds as a qualitative finding, but the quantitative magnitudes should be treated as provisional. The paper deserves peer review, and a serious referee should ask for release of the benchmark and code, per-condition evaluator agreement, and an explicit check on the wording confound. If those are addressed, this could become a standard reference for spatial reasoning in T2I evaluation.","headline":"A genuinely new T2I evaluation axis with a large, credible deficit, but the exact magnitudes need released artifacts, CIs, and a cleaner prompt contrast before I'd trust the 41.8% headline.","tokens_in":16517,"tokens_out":5338,"would_cite":true,"duration_ms":43248,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Text-to-image models fail when spatial directions refer to an object's own frame of reference, with mean accuracy 41.8% lower than on matched camera-view prompts.","keywords":["frame of reference","text-to-image generation","spatial reasoning","layout generation","benchmark","prompt rewriting","vision-language evaluation","object orientation"],"falsifier":"Re-run the 1,200 layouts with FoR prompts rephrased in short, natural language while keeping the anchor orientation unchanged; if the gap shrinks to near zero, the deficit is largely a phrasing artifact. Alternatively, if any model scores near its Cam accuracy on the reversed-mapping subset, the claim that all models fail reversal would be false.","tokens_in":15316,"feed_emoji":"🖼️","tokens_out":8368,"duration_ms":62448,"temperature":0.7,"pith_summary":"The paper asks whether text-to-image models can honor an explicitly specified frame of reference when it disagrees with the camera view. To answer it, the authors build FoR-T2I, a benchmark of 1,200 matched prompt pairs: each pair describes the same layout once from the camera view and once through an oriented anchor object's own left, right, front, or back. Across 22 closed- and open-source models, mean final accuracy is 41.8% lower on FoR prompts, and even the strongest model reaches only 44.3% FoR accuracy. The failure concentrates in target placement rather than anchor rendering, and it is sharpest when an anchor-relative left or right maps to the opposite image side, where FoR accuracy collapses to 18.8% against 68.5% on the paired Cam prompts. The paper also shows that a vision-language-model-gated rewriting strategy improves FoR accuracy from 25.0% to 29.2% under the same generation budget.","feed_headline":"T2I models miss object-centered directions by 41.8%","feed_subtitle":"New 1,200-pair benchmark shows every tested model drops when 'left' means the object's left, not the image's left.","key_machinery":"The load-bearing object is FoR-T2I's matched prompt-pair construction: each layout deterministically yields one Cam prompt and one FoR prompt that differ only in the frame in which the direction word is resolved. A second mechanism is the frame-mapping trichotomy used in the analysis, which classifies anchor-relative left and right relations as preserved, axis/depth remapped, or reversed relative to the image; the reversed cell is where the FoR deficit becomes a collapse. The automatic evaluator separately scores geometry (target placement), orientation (anchor facing), and their conjunction, which lets the paper attribute the gap to placement rather than orientation.","core_discovery":"The paper's central claim is that current text-to-image models systematically resolve directional language in the camera's frame and fail to draw from an object's intrinsic frame of reference when the two disagree. This is established by a controlled contrast: the same abstract layout is verbalized as a camera-view prompt and as an object-frame prompt, so any accuracy difference isolates frame-conversion difficulty. The paper reports that every one of 22 models scores lower on FoR prompts; the deficit is greatest for left and right relations under reversal, where the Cam–FoR gap reaches 49.7 points, and component analysis shows the loss appears in target geometry rather than orientation rendering. A secondary claim is that a training-free, vision-language-model-gated prompt-rewriting method can recover a few points of accuracy, but most of the gap remains unresolved.","pith_inferences":["An untested extension would be to probe whether the same reversal failure appears in video generation or embodied instruction following, where object-relative 'left' is operationally critical.","The 18.8% versus 68.5% reversal finding implies a cheap diagnostic for future models: testing only reversed left/right mappings may predict most of the FoR deficit.","Because the benchmark's FoR phrasing is templated and less natural than everyday spatial language, natural rephrasing could change the measured gap; if it does, part of the deficit is a register-matching problem rather than a purely spatial one.","A testable prediction is that fine-tuning on reversed-mapping examples with orientation supervision would generalize to unseen anchors and reduce the wrong-frame error pattern."],"forward_implications":["If the claim holds, strong image-frame spatial control does not transfer to object-centered descriptions; the two abilities separate empirically in every model tested.","Benchmarks should report frame-of-reference accuracy separately, because aggregate layout accuracy hides the 49.7-point reversal failure.","Since the failure survives in the best closed-source models, training-free prompt rewriting alone will not close the gap; orientation-aware representations or training objectives are needed.","The vision-language-model-gated rewriting result suggests visual feedback can serve as a cheap partial mitigation in deployed systems without updating the generation model."],"supporting_citations":[{"why":"Supplies the intrinsic-versus-relative frame-of-reference taxonomy that defines the FoR condition.","marker":"(Levinson 2003)"},{"why":"VISOR is the prior spatial-relations benchmark that FoR-T2I contrasts with on object-relative directions.","marker":"(Gokhale et al. 2022)"},{"why":"GenEval contributes object vocabulary and serves as the composition-benchmark comparison.","marker":"(Ghosh, Hajishirzi, and Schmidt 2023)"},{"why":"T2I-CompBench++ is a source of object categories and a compositional baseline.","marker":"(Huang et al. 2025)"},{"why":"SpatialGenEval covers object-relative scenes but lacks the controlled Cam–FoR contrast FoR-T2I provides.","marker":"(Wang et al. 2026b)"},{"why":"FoREST evaluates frame-of-reference interpretation in language models and motivates separating it in generation.","marker":"(Premsri and Kordjamshidi 2025)"},{"why":"SAM3 grounds objects in the automatic evaluator so that geometry can be scored independently.","marker":"(Carion et al. 2025)"},{"why":"Depth Anything 3 supplies depth estimates the evaluator uses to judge spatial relations.","marker":"(Lin et al. 2025)"}],"fun_headline_variants":["T2I models lose 41.8% accuracy when 'left' is object-relative","Object-frame prompts cut T2I accuracy by 41.8% on average","When 'left' means the object's left, T2I models fail 41.8%","FoR-T2I benchmark: 22 models all drop when direction is object-relative"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The Cam–FoR gap is interpreted as frame-conversion difficulty because the paired prompts are meant to differ only in the frame of the direction word, so wording complexity would confound the comparison.","fun_headline_variants_meta":{"raw":{"variants":["T2I models lose 41.8% accuracy when 'left' is object-relative","Object-frame prompts cut T2I accuracy by 41.8% on average","When 'left' means the object's left, T2I models fail 41.8%","FoR-T2I benchmark: 22 models all drop when direction is object-relative"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000961,"raw_usage":{"total_tokens":4118,"prompt_tokens":993,"completion_tokens":3125,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":3029}},"tokens_in":609,"tokens_out":3125,"duration_ms":21097,"temperature":1.0,"reasoning_tokens":3029,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:54:15.000300+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the 1,200 layouts with FoR prompts rephrased in short, natural language while keeping the anchor orientation unchanged; if the gap shrinks to near zero, the deficit is largely a phrasing artifact. Alternatively, if any model scores near its Cam accuracy on the reversed-mapping subset, the claim that all models fail reversal would be false.","supporting_citations":[{"cited_title":", title =","cited_arxiv_id":null,"evidence_quote":"Supplies the intrinsic-versus-relative frame-of-reference taxonomy that defines the FoR condition."}],"review_version":2}