Pith. sign in

Paper Citation Record · LEDGER

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models

As of 7 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 0 inbound Pith citation observations for arXiv:2506.23418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23418 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:48:57.591417Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

85 of 85 outbound references displayed

  • verified exact6
  • verified fuzzy27
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b45c9dc-5db5-4982-9ef1-e50edfe83359 · outbound

This paper cites write newline.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:50.808315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:50.808315Z digest=sha256:250e8c699a08d7e5544549cef09b5fd957935440b43ac33582041b6970bb00d0

Observation 3f71b3e6-8ec6-463b-bbec-afd2add3b53b · outbound

This paper cites A-star: Test-time attention segregation and retention for text-to-image synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models A-star: Test-time attention segregation and retention for text-to-image synthesis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:50.867916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:50.867916Z digest=sha256:98ca19474d1e1600e2ed023107005654bebcf49b25f23996fa3a06308c3bf82c

Observation 07799072-ad50-41b4-b658-3d6ceaf1db9d · outbound

This paper cites Kandinsky 3.0 Technical Report.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Kandinsky 3.0 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:50.976891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:50.976891Z digest=sha256:5f4803d8bcdde58bb4d4a5ee053b2668c06460c33b6c5fa465298346c802c051

Observation f2a69bfd-0dca-4dd9-8e21-23820b2719b6 · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Finite-time analysis of the multiarmed bandit problem

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.033654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.033654Z digest=sha256:9a2a94ba00c2a40f6d967bf049880f82be8b4674b2ad0a5806e7de7480fb24b5

Observation 6a8c774a-95af-48c2-849f-65c0d9e7309f · outbound

This paper cites Spatext: Spatio-textual representation for controllable image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Spatext: Spatio-textual representation for controllable image generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.080754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.080754Z digest=sha256:880de5d4a68005348477b915554a0740197668c83934f54d525017c9f388861f

Observation 4b624120-e17e-4a45-9d74-32299bfb1238 · outbound

This paper cites Cc3d: Layout-conditioned generation of compositional 3d scenes.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Cc3d: Layout-conditioned generation of compositional 3d scenes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.127391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.127391Z digest=sha256:d858e80f0b76926f9b99d0e496ffc8fd5f391dfe94e4d6264fb9652e25837e6f

Observation 2e40a2e8-8777-44e5-873e-7d0bd3629d80 · outbound

This paper cites Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models, 2023.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.167595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.167595Z digest=sha256:010629ec7f7456ca41ca5d3e2d2a3e248bb72fef50161c30a7fa63abbacabaff

Observation c3b64dc5-5646-4eef-900f-5df72f601870 · outbound

This paper cites MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.252650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.252650Z digest=sha256:b8f621a7131c9bd489de8ac414684af05b9bdd651995aa24a25ed471fc094642

Observation 41a8f068-b396-4a9b-b4e4-133192444c86 · outbound

This paper cites Improving image generation with better captions.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Improving image generation with better captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.298620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.298620Z digest=sha256:337cb78ddd7c18dccc54b4094d024e4bf17f2ee9f153d99422520325f784ff6f

Observation 4fd1b3e2-497d-4ced-ae61-9af9773af035 · outbound

This paper cites Make It Count: Text-to-Image Generation with an Accurate Number of Objects.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Make It Count: Text-to-Image Generation with an Accurate Number of Objects

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.348576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.348576Z digest=sha256:dd56c9e8859de98ed52a6458b2affd2682b583901193250684e2facc86b5cf95

Observation 26a059a3-eb12-4ace-b2ed-612a9e13c0e9 · outbound

This paper cites Conceptual 12M : Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Conceptual 12M : Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.878208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:51.432547Z digest=sha256:d21eeda870d6510773d8b1d1701514a60cdcfcec1b76521484e2180c8b8c8171

Observation b3a0d950-88b9-4ea0-9329-aa5f0ab4ffb8 · outbound

This paper cites Getting it Right: Improving Spatial Consistency in Text-to-Image Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Getting it Right: Improving Spatial Consistency in Text-to-Image Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.518955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.518955Z digest=sha256:647edc2f8eef871071353ddc29123e2680dbf911e1099336f649e6cfbda6d194

Observation 4bfb679c-f849-4704-bf91-e714a875ec45 · outbound

This paper cites Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.719627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:51.601008Z digest=sha256:64369ba40b6785d6b6ff48a9ceb17137ebf9a2e4e93ddaae9b8a1faf8ca2a58d

Observation e1d998d1-f294-4f63-905f-d4b06ff9f1f6 · outbound

This paper cites SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.657856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.657856Z digest=sha256:0f310b8fa99cacec49b000ab1283ad9d82caa215465b3a5ef8a0b73cf9b81ad3

Observation aaae8dd0-4517-47b7-bf31-a63a6426e88d · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.736340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.736340Z digest=sha256:a35e2d2e73e74a0bee2c0cedd4614fdad729956e2300d5c1f64dd6695b7d5724

Observation 307c10fb-40cd-45ef-8cf9-05c055f67e7e · outbound

This paper cites Training-free layout control with cross-attention guidance.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Training-free layout control with cross-attention guidance

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.564234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:51.783311Z digest=sha256:027578e96918651dbf91e67e55c8e7576382eeec14ed6dbe0fec77101b5ff1c4

Observation f60b0b08-79b3-4092-81a7-67085a3a6130 · outbound

This paper cites LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.832817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.832817Z digest=sha256:2c965216533627a351c7a9bdd25ea2fbb7de2879e58a9746173770177590f887

Observation 4a510a3f-4052-400d-b187-98e87f161fd8 · outbound

This paper cites Dall·e mini, 7 2021.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Dall·e mini, 7 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.406246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:51.883431Z digest=sha256:44405f2ed0a49a194c61a85a3c2b5457a4f3f383a430e0a42c09eaaa3d0792b0

Observation 9765d374-7cba-47c7-84b2-6d72e643d53c · outbound

This paper cites CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.971405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.971405Z digest=sha256:e5c89359c63836d9effb676d0d17b7d790febab4edd42125b145b49e373684f9

Observation c7458bd6-c8d0-4088-bc99-df1b62ad21c2 · outbound

This paper cites Training-free structured diffusion guidance for compositional text-to-image synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Training-free structured diffusion guidance for compositional text-to-image synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.187547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:52.059608Z digest=sha256:2f5d48d023211b6fe63377433060f6b1b3906c888ef82442867cddae2ed7f1be

Observation 6e88afe7-aa0c-466e-b77f-134e9213febc · outbound

This paper cites Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.182184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.182184Z digest=sha256:c1419eebe991508fe4e233695eb558a83ae49b0c7c355e938839c0c4159af17c

Observation f2b3f176-fa36-40c8-99f9-e08ba899c89e · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.298002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.298002Z digest=sha256:45c2c81cf7211cb563ddc571e8524c39ce826a3cf4304656a3b191655da66f2a

Observation 54e5f73d-cf88-4460-bc27-fa90af3275fe · outbound

This paper cites Benchmarking Spatial Relationships in Text-to-Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.391265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.391265Z digest=sha256:60e5b07c9fa3bc4696a1721091e09be291d8d89273e7374303120f71d56c9097

Observation a97cb007-cd59-49c9-8682-85279d24b011 · outbound

This paper cites Gradient Guidance for Diffusion Models: An Optimization Perspective.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Gradient Guidance for Diffusion Models: An Optimization Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.464757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.464757Z digest=sha256:7ddf1108aaecca21c356511d96a34a0850b19d79e2578883d7f91b176c278a40

Observation 56e3b960-01db-4269-88bc-a7be5e2209ab · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.605567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.605567Z digest=sha256:b520fe4aeab0d67eb76e9eeadea0c2950403fa103f028d1e69c6801d654ea9df

Observation a5d12599-6650-44dc-a625-5dbe2c18cc58 · outbound

This paper cites GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.737177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.737177Z digest=sha256:86b0764bc4521103bf84f5faab013de30612c909af2d5348d940a514ed0a8e32

Observation a7b727f7-47fc-4c9b-83c7-40f10f905237 · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Probability inequalities for sums of bounded random variables

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.842219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.842219Z digest=sha256:77a5f36de86ad60b7606456084ce7d45b4ea450560bdaf7451b4a6474158ed11

Observation b6111a05-8052-4a32-8d30-9ecd28e2c8e2 · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Probability inequalities for sums of bounded random variables

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.928167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.928167Z digest=sha256:5d92ab59ee388705aa1d8cbafd88a73d9ab44b6881a1cf4887a85359bbf86088

Observation f00c05be-3b62-4de4-a843-a95554a680d8 · outbound

This paper cites Token merging for training-free semantic binding in text-to-image synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Token merging for training-free semantic binding in text-to-image synthesis

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.036753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.000087Z digest=sha256:b44f6ab2af12e8614cd2b7799aaf98b4205634d36c18591c96d3db0344a5d1f3

Observation 79b60aa4-8021-40ee-a135-487356241c7b · outbound

This paper cites A Multi-Armed Bandit Approach to Online Selection and Evaluation of Generative Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models A Multi-Armed Bandit Approach to Online Selection and Evaluation of Generative Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:58.751381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.058729Z digest=sha256:7635528625aee23b23caf5319e4be8f5b4050e4c03fb130d73b1cf5f0a488853

Observation 68e853c2-b619-4f9d-95ab-a9eda965bb51 · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.829715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.116358Z digest=sha256:7fe1d97217c86fd8002faca7f8b94c772ce1fee3c98f9212ee84cf3a0ed05838

Observation 8206daf3-8e7a-48dc-9aa8-fead9a8a5ba7 · outbound

This paper cites An information-theoretic evaluation of generative models in learning multi-modal distributions.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models An information-theoretic evaluation of generative models in learning multi-modal distributions

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.654824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.228368Z digest=sha256:3d33699f86e01541f9f75e8831aba8f00653ee83082d0a4a1d9e639d37b3f940

Observation 07d48034-0798-4fe7-9a7d-d13cc77c49dc · outbound

This paper cites Rethinking FID: Towards a Better Evaluation Metric for Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Rethinking FID: Towards a Better Evaluation Metric for Image Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:53.316391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:53.316391Z digest=sha256:4d2ab8e43773d7efda3f56fad4ef43d307b497fb9cd707800ae36340e320fd24

Observation f2b20c63-ef67-4c6f-8e17-fdad4f6363f6 · outbound

This paper cites What's ``up'' with vision-language models? investigating their struggle with spatial reasoning.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models What's ``up'' with vision-language models? investigating their struggle with spatial reasoning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.501071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.352085Z digest=sha256:fea4fa8365cb6d705c435dce46ac7be8e9b6222f87359768a9ea5fb8e2b1525a

Observation 568f43b0-6aeb-4c70-a4c0-43df7e04bb1f · outbound

This paper cites If at First You Don't Succeed, Try, Try Again: Faithful Diffusion-based Text-to-Image Generation by Selection.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models If at First You Don't Succeed, Try, Try Again: Faithful Diffusion-based Text-to-Image Generation by Selection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:53.437835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:53.437835Z digest=sha256:ce67e900ad3ca58fbbaefd8cb8c268f50928591bfd98124de08f91fa80edd179

Observation 71edc826-1f08-423d-bc26-8dfe123c5cfb · outbound

This paper cites BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:58.598720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.533737Z digest=sha256:b5e55904e0120ce3ce8fa00f63433034d0be1721f6cb9fde67be888b1bdc94d4

Observation 65b24792-db47-47ac-a5cb-977279bd3622 · outbound

This paper cites Dense text-to-image generation with attention modulation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Dense text-to-image generation with attention modulation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.338882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.586132Z digest=sha256:d38294a1ccc296376100fd0798f5f45862e169627d6ef648e2a147b6adf3dc2f

Observation 60788508-c8c4-4c60-b0e9-270c434757eb · outbound

This paper cites Segment Anything.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Segment Anything

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:53.722447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:53.722447Z digest=sha256:c3d3b294f0fcf6e363c4296ba506b69c87d76f3ac54aee8740c17775d321a904

Observation 632e5ff8-76bb-4f4e-89de-f5cc318033e6 · outbound

This paper cites Improved Precision and Recall Metric for Assessing Generative Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Improved Precision and Recall Metric for Assessing Generative Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:53.852853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:53.852853Z digest=sha256:56f7343113487383f1d737124dacbce03fe2b50feb984bc1eba7735f1e1b3fcf

Observation 938ff7f5-8e90-47a5-b9f5-08632237ec65 · outbound

This paper cites Laion-coco 600m.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Laion-coco 600m

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.158280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.993200Z digest=sha256:b661c7482542e10b1909a03e1e47221866add79c5f9c0fcd28b993b399047077

Observation cc4b0226-9d4c-407a-8619-11b898df9d86 · outbound

This paper cites GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:48:58.472999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:54.063558Z digest=sha256:b29ff209ab4c1e984fedd79d443ab0f8642c2e3a3303ba51f2ab153d137fac66

Observation 83dd898d-37de-4d8f-b1b2-ea746ca26f09 · outbound

This paper cites GLIGEN: Open-Set Grounded Text-to-Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GLIGEN: Open-Set Grounded Text-to-Image Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.226577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.226577Z digest=sha256:ec35c8fb36b9b02fc1c6a709a7bfb482e48ccdf61072b7c5b1b8ee0d2601f82a

Observation 0981c3cf-42c0-4f5a-ad88-a0fa44bdb9bb · outbound

This paper cites Divide & bind your attention for improved generative semantic nursing.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Divide & bind your attention for improved generative semantic nursing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.947376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:54.383202Z digest=sha256:2a5da985e0ef912c798dcfeab0ae549cfcfcff6d4f19141eeafa58eba6dcf22c

Observation 163ad252-7a8e-46fe-966d-63458b240b4e · outbound

This paper cites LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.541171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.541171Z digest=sha256:b9ea144be9d541a1052fef499ea459d9e867f050cd0429a0cd0ed719b4d09a60

Observation 1b6393a3-ac2e-4b9f-b839-c7024428db99 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Microsoft COCO: Common Objects in Context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.670059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.670059Z digest=sha256:d22c8c52bcae0774a4d930db6a27006384940cbce1b875a53019a7562b987a07

Observation babd993d-63a6-4fb9-89a9-c3b2f7b3f4cd · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.778142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.778142Z digest=sha256:e5f3aafd24b6dd4be57735b969bf4ccd5921970108f65a2762842c43054a399b

Observation 0c9d4880-0781-4eaa-a364-e680ac9a6c13 · outbound

This paper cites Compositional Visual Generation with Composable Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Compositional Visual Generation with Composable Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.880447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.880447Z digest=sha256:ab804f9fabe0b46f1c8ba9795fe1088d892c67f13b7e11abbb21fe8c5b269f2b

Observation ced64cf3-11ee-4234-afc1-7abfb9ed26c6 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.783156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:54.923631Z digest=sha256:3cfeae47054199f84301be1682c132db2f7e9d5a7344da96adb7509df0f40abb

Observation 9fd69e75-d069-47bf-9d8d-e0c8a10bb6cf · outbound

This paper cites Correcting diffusion generation through resampling.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Correcting diffusion generation through resampling

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.584642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:55.042044Z digest=sha256:ff35e64afc47ec706cb9e40aa2b0cc239313a424beccd3aac06eb5228a0e574b

Observation 9090d239-61cb-4e85-98b5-3771c792f67e · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.131271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.131271Z digest=sha256:9312e4c9da6d61977b9dd56c105a6741a7925d56b1e1093fde90b51a5be960f4

Observation 1a67cb10-e9b4-462b-8869-d8d9ce9ea2b3 · outbound

This paper cites Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:58.305767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:55.186979Z digest=sha256:b2a8bcef510986b09e154d58f35b7eeeb482572e118e7ea4fd3936f80dddf56a

Observation b8d16741-9771-454f-b974-e8b3bc0b739c · outbound

This paper cites Scaling open-vocabulary object detection.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Scaling open-vocabulary object detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.393772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:55.224103Z digest=sha256:7a049ca63b7547a6500708d8efe19258ed7904255d139134eae8017256aeb34d

Observation 9970ec12-3997-4701-9dba-82053fae0d7e · outbound

This paper cites Simple Open-Vocabulary Object Detection with Vision Transformers.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.316835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.316835Z digest=sha256:2c31506deb703e3ed79b6d0fc95756e9f2426372e26e690f1ef28798522b77fe

Observation a378cbee-6979-48cf-a0c9-9d1ed0c2c8aa · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.441530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.441530Z digest=sha256:1b7832c9fa2ba6f8dae842ab4fb9d4db940011f774180d81cc03508bdda1238c

Observation 35e4c707-9e52-4348-b0d1-5155491146cf · outbound

This paper cites GPT-4 Technical Report.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GPT-4 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.619085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.619085Z digest=sha256:f01e45539e2d9af2112bae03087a988f143c7560ec8d1e18069acedd9bc1c551

Observation 84c356ae-bf5c-4a8f-b989-9ee6242478a9 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models DINOv2: Learning Robust Visual Features without Supervision

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.723845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.723845Z digest=sha256:e4648b119f66ce096d409b4f49357c1f13fe6868a52d21fe7bd268a10b755bad

Observation 676cee09-5884-4021-9fad-a5e72b17e219 · outbound

This paper cites Grounded Text-to-Image Synthesis with Attention Refocusing.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Grounded Text-to-Image Synthesis with Attention Refocusing

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.795064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.795064Z digest=sha256:90c182e0dc16b3d1eb5508427bd9a585f290fb251a7385264dd40b2ff69d156c

Observation 721246fa-07dd-4586-a22c-97860a7034c4 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.869024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.869024Z digest=sha256:05a8d0bedcd32ef4a2ea15d2fbe7c738e21493820e1322929e4dc92e0c8387fd

Observation 3a181d03-2143-46e2-898a-5e4a2f6f696a · outbound

This paper cites Learning transferable visual models from natural language supervision.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Learning transferable visual models from natural language supervision

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.192900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:55.965615Z digest=sha256:8a6182491a4da44c023b9a41b4ea0d109227cbc2c210b47f49ab5ae89ebd9bfc

Observation e6aace94-dd9d-4064-93eb-a33d6b4d5ae1 · outbound

This paper cites Zero-shot text-to-image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Zero-shot text-to-image generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.928478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.023411Z digest=sha256:acc3f5e2c953f7e026ad06f80054e4bbad87c44b4cc8cb74054ea8ab3f09be48

Observation 26b51eb6-89dd-43c0-83c6-b8ba3c63a021 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.058839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.058839Z digest=sha256:9f66097932aaa75a3cbe6521ef3a6c55376ef8bf0bf35b58ed87ecfccd455b42

Observation f2edd81a-8cec-4252-a975-0b63c4d33ac0 · outbound

This paper cites Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.717211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.114052Z digest=sha256:39e644469296a3d0cf9daafb9351df5a8acfd92aac8d8cd4f73fe4d06306aa4a

Observation 367c2b0b-cd2b-4eb0-9ee8-07b77b7c9b72 · outbound

This paper cites Arkhipkin, Igor Pavlov, Ilya Ryabov, Angelina Kuts, Alexander Panchenko, Andrey Kuznetsov, and Denis Dimitrov.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Arkhipkin, Igor Pavlov, Ilya Ryabov, Angelina Kuts, Alexander Panchenko, Andrey Kuznetsov, and Denis Dimitrov

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.511252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.170409Z digest=sha256:d048e0d630ed84021fd7ff536bbe60f877fff958a68f74d04d559d215097162e

Observation 112b0fc5-524c-4449-a5eb-e9e146ef55fd · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.243827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.243827Z digest=sha256:2cc0737b57ab2533289c04fc621235bf457547037fe51e99dc2a1fd896b637bd

Observation 9bf64bec-9478-44dd-9edb-b24615ee4229 · outbound

This paper cites Be more diverse than the most diverse: Optimal mixtures of generative models via mixture- UCB bandit algorithms.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Be more diverse than the most diverse: Optimal mixtures of generative models via mixture- UCB bandit algorithms

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.335950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.338961Z digest=sha256:a63e39a70734e60f95e48b936d75fbd214a2515f41fbc617ade03ac4589b397d

Observation dc1a2e9c-df3d-4cad-9d28-43c19116b4c3 · outbound

This paper cites Some aspects of the sequential design of experiments.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Some aspects of the sequential design of experiments

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.381683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.381683Z digest=sha256:ff10bb35e4363854fd4524f0ea2f0200b856ca49960e76e88a23d0b8f9aeee2b

Observation dce2eed6-d5f8-4927-b29d-0b19ead8005f · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models High-resolution image synthesis with latent diffusion models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.449019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.449019Z digest=sha256:17dfcb4dfadc87c8293302f072c9344895a944fa4d7d9364cc21ebc75ecd7864

Observation e97464c5-dbec-4752-bd3f-816cbb486125 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.178187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.503848Z digest=sha256:4c30f02f28293bc9231dc2ca1860fa90c456bc170845ea8054beb5300e3f47e1

Observation 1878b31b-1707-4e95-b649-8c5be5da182b · outbound

This paper cites Predicated diffusion: Predicate logic-based attention guidance for text-to-image diffusion models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Predicated diffusion: Predicate logic-based attention guidance for text-to-image diffusion models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.979688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.540544Z digest=sha256:60f84d6c91ea8b3400cacc1adca86f9f417f8f2db45aa0552f8a7d5b7960b0f3

Observation e6c76451-9de6-4185-a098-92aa33429390 · outbound

This paper cites InstanceDiffusion: Instance-level Control for Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models InstanceDiffusion: Instance-level Control for Image Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.611116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.611116Z digest=sha256:0fabaf8d1c85c1b5fc6c020ea9c9a61e8a4045f1e24fd82536bac4874066e77d

Observation 6eea2c92-3000-4d91-a3bb-cca1487973e0 · outbound

This paper cites Wolfe and Robert V.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Wolfe and Robert V

Reference 71

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:48:58.107706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.663174Z digest=sha256:168daf72c0668aa9ef544bee521fb06bbbafeefb949ce7b3fbf4593e0dd50033

Observation aa2d77d7-1af9-455d-87fa-796a078c9105 · outbound

This paper cites R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.745122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.745122Z digest=sha256:2dafda60ab6d9d5972e59eb861d0fd6d64630065ff826c56177f681a234d3898

Observation 1d63d2f5-54b2-4226-bed7-c60f7208a5cd · outbound

This paper cites BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:57.931564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.816954Z digest=sha256:61bf2ba665ea1c9c7ab85269cca92545abdb47a17a3c4bcaf5160b8dc31ccc0e

Observation 8ba51f86-cc4a-4f76-a2b8-a0ccd79790fb · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Imagereward: Learning and evaluating human preferences for text-to-image generation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.834159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.852398Z digest=sha256:21503ddf975eb15f4cf2d4e869a7c2393ed0d7f8877b6fc9b2168b6003778700

Observation ff379478-dc49-45a9-9c2d-0c34e5ca44a1 · outbound

This paper cites Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.923723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.923723Z digest=sha256:4d31bace29b0349c9182f61dfffc82c7c99948e50f8e4d6617aa740c7efcf822

Observation 855cada1-bf5a-43e5-8e87-dbdad5b0b6fc · outbound

This paper cites Reco: Region-controlled text-to-image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Reco: Region-controlled text-to-image generation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.681822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.998259Z digest=sha256:9e6ab1cb476f314eff2901be770663e988743e628d0f08e22388aaffeedf49d9

Observation 3c0c1880-0a35-4d0d-a0ed-89c56d17e0ef · outbound

This paper cites Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.035500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.035500Z digest=sha256:d3069b491479feb28a6059e844633bdbb934c0b1327254446a1d7ddfed0169da

Observation 0ecde0f2-5b63-4da1-a25c-e919df18b343 · outbound

This paper cites Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.092979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.092979Z digest=sha256:e4361d1e6d3fbe02ad4bcf7865b419b7f31505668ac01407f34340b7ae91a9af

Observation 6c6e4dad-b9be-4ec0-9c4e-a2e4a8147b8d · outbound

This paper cites Multi-grained vision language pre-training: Aligning texts with visual concepts.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Multi-grained vision language pre-training: Aligning texts with visual concepts

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.553613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:57.225701Z digest=sha256:51ab601eff5e0b45a0b82c63bb7c3f7f7f296934828f103e2ec60e07572e775f

Observation dbe8a91e-ea7b-491b-89ed-248bec555a13 · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Adding Conditional Control to Text-to-Image Diffusion Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.313620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.313620Z digest=sha256:991dd039cee8b880a112545502a7856767aeef028e09c797b601cabb8a74fbe9

Observation 2075568a-7757-4427-bc8b-f412c48e1914 · outbound

This paper cites Realcompo: Balancing realism and compositionality improves text-to-image diffusion models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Realcompo: Balancing realism and compositionality improves text-to-image diffusion models

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.423169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:57.350636Z digest=sha256:d47aca7bc140b227e725d0f854339771746ca782995a4a336e5ee4c5be24bd0f

Observation 640cee54-8bbc-4849-9c19-455d032b15ab · outbound

This paper cites Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.426716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.426716Z digest=sha256:3d9fb782f4cf5ed72d2778e7a19a860b6a2996e52a00d7219c5050f4af8003c1

Observation ac7f8dd9-14e8-4c32-a600-b96082be45b4 · outbound

This paper cites LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:57.770163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:57.484646Z digest=sha256:6c9fe33b531f8bb10c4c587e0aed2425d9ab1220d46fa4961b3a464c95325ea9

Observation 7924c320-10cd-4992-9f9a-459d6d3807d2 · outbound

This paper cites a henb \.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models a henb \

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.258838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T21:48:57.523921Z digest=sha256:2dc1a2497c496dc337134203edf7cb5b7ba47fac3a8b14c4fc224f21b98bd092

Observation 092761a5-b219-4d41-b024-75209b95bd11 · outbound

This paper cites write newline.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models write newline

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.591417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.591417Z digest=sha256:9a2d04abd53d5cc9e21656154b83ddbc7c9f6a10c9ff0e71554ee97a1946a632

Pith citing papers

No inbound Pith citation observations are available.