Pith. sign in

Paper Citation Record · LEDGER

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation

As of 20 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 3 inbound Pith citation observations for arXiv:2411.15435.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15435 v2

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:22:50.307233Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:06:06.096668Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T09:13:30.200834Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy56
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d395cac-7335-4694-8f68-152baff15b89 · outbound

This paper cites GPT-4 Technical Report.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:49.853258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:49.853258Z digest=sha256:47680dd889ae73807422ee9ae6085b3a15c3a83c6f1bb678d475d8f7db8afbe3

Observation d0dd42a4-52c1-4ac9-996b-5a1af3c25aee · outbound

This paper cites Gemini 1.5 flash pricing.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Gemini 1.5 flash pricing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.915610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.859373Z digest=sha256:555333fe8a1fba726c8abc609a2a14b71d8fcfa11c00c50522f4349bd650130c

Observation db369205-f837-4810-8add-7138ffe2407f · outbound

This paper cites Specifying object attributes and relations in interactive scene generation.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Specifying object attributes and relations in interactive scene generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.892400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.876489Z digest=sha256:6e29f8449cc5cf44a10a80d0a97a9e1cc719a2b9ca7240b4f894c42d8cc37173

Observation ac46bc0d-02f3-4f30-8928-0d97554340e3 · outbound

This paper cites Large scale GAN training for high fidelity natural image synthesis.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Large scale GAN training for high fidelity natural image synthesis

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.872432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.885599Z digest=sha256:b0b8e6f0c02d42785cc770e290bc7048087db2205a40586d871758f9f9214a6c

Observation 47e8884d-6f8f-42ff-988e-bcb2685db457 · outbound

This paper cites Language models are few-shot learners.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Language models are few-shot learners

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.854416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.891443Z digest=sha256:3d6fbd345b92f0fd277bbd0e99c8973b76169cbd957d01389a45fc65c47bb3bd

Observation 9550aa2c-cb24-46a1-912e-94f69e2df895 · outbound

This paper cites Coco- stuff: Thing and stuff classes in context.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Coco- stuff: Thing and stuff classes in context

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.835154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.897070Z digest=sha256:1becb390b8dbd6799168dfd92aee46f42200b6ea06b268f794307d9a4b15e432

Observation 8f3b8c7d-4318-4561-b2f2-fee816931df1 · outbound

This paper cites Getting it right: Improving spatial con- sistency in text-to-image models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Getting it right: Improving spatial con- sistency in text-to-image models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.815199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.902598Z digest=sha256:24e31e48c1c87b3d0ffbad03d3253fd66f8b0fb7bbd1f7c5a8a9dcf2f6e820f5

Observation 094f9ada-3ce8-4797-998d-382482199aea · outbound

This paper cites Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.792666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.908599Z digest=sha256:af2145b78295b5d975e3516a6963c44236415c7e8a2c169912b6286d0e854d37

Observation 817cde7a-3001-458e-bda3-3c41d49d8d12 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:49.914732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:49.914732Z digest=sha256:ec5c02de3dbcf51b1772c535054b09377637b304194c5a89b076ac004e4d9836

Observation 44bbf2c9-04df-4258-9752-47f211591eb1 · outbound

This paper cites UNITER: universal image-text representation learning.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation UNITER: universal image-text representation learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.769824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.921314Z digest=sha256:969d3036d26c4474862b0b4f4e87c25adce1ba6b07adca2cad34345fec70ac4a

Observation f04802dc-e165-406b-9962-15b264764033 · outbound

This paper cites Expanding scene graph boundaries: Fully open-vocabulary scene graph generation via visual-concept alignment and retention.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Expanding scene graph boundaries: Fully open-vocabulary scene graph generation via visual-concept alignment and retention

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.750337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.926559Z digest=sha256:da92293b61650b7c4bb7c4915fd873744b4029af718936d1d0114b2de573a59f

Observation cb9a5967-0754-4d7d-a7a5-5a5da9298178 · outbound

This paper cites Baldridge, Roopal Garg, Peter Anderson, Ranjay Krishna, Mohit Bansal, Jordi Pont- Tuset, and Su Wang.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Baldridge, Roopal Garg, Peter Anderson, Ranjay Krishna, Mohit Bansal, Jordi Pont- Tuset, and Su Wang

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.730080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.931595Z digest=sha256:dfbbc4b9639983464d13aaecfb20a22aeba2f8fc3f4c5ab4eb2b4e01e6e55338

Observation 612579c4-1f22-4c54-81d3-4752ee162960 · outbound

This paper cites Text-to-image diffusion mod- els are zero shot classifiers.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Text-to-image diffusion mod- els are zero shot classifiers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.709711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.936906Z digest=sha256:230ebf7f4f0101c76650914c35b4d5687d025915df3c06642d09d9f9b5ac031f

Observation d6a46c81-3f01-4835-84b2-988573a9e13f · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.683897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.942270Z digest=sha256:fff1258321ea3b8f74538f4f16c4c087084a74ae99485d88cfb0d7fd3a170ab8

Observation 75f4d74a-e93e-4cc5-9df8-5f1db2655733 · outbound

This paper cites Re- inforcement learning for fine-tuning text-to-image diffusion models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Re- inforcement learning for fine-tuning text-to-image diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.661131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.947326Z digest=sha256:a0b7ab1554e77039f71cfd6e9def970384f5f5b4e4c1f4d7bcd5a31cd6adfde5

Observation 5ef2c247-d263-4b35-b3fb-33d2473bb7b4 · outbound

This paper cites Scenegenie: Scene graph guided diffusion models for image synthesis.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Scenegenie: Scene graph guided diffusion models for image synthesis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.641869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.952460Z digest=sha256:59e93131fec9df83ebd898b5696991546c725b4aafa9206731d83b5802ccc018

Observation 9347571b-0ac8-4b36-84af-e642b6f01903 · outbound

This paper cites Akula, Pradyumna Narayana, Sugato Basu, Xin Eric Wang, and William Yang Wang.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Akula, Pradyumna Narayana, Sugato Basu, Xin Eric Wang, and William Yang Wang

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.620697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.957432Z digest=sha256:43486732d68386a7b8ac773b6c4bfa789c375fbdfd26611ff4d5a49647ec7060

Observation 5fc276be-d5b9-47b6-923a-ce8295fc6df8 · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Layoutgpt: Compositional visual plan- ning and generation with large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.601436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.962781Z digest=sha256:95687c7325cb4ee98254a9b75c9bc36e29b9a43101f9adbe96e1151b3b542e88

Observation aa13cbfa-3350-4be3-afdf-cc92ea9b8e0e · outbound

This paper cites Generative adversarial nets.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Generative adversarial nets

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.575600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.968547Z digest=sha256:4914d0556af4260a18a7875099968a990b271914dd918f5e66a5c5ae4711fbd3

Observation 828d2225-26d5-4f76-a9c4-45789995b0c6 · outbound

This paper cites Generative adversarial networks.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Generative adversarial networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:49.973749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:49.973749Z digest=sha256:694675f63b3cd557650fa2abfd4cbb66f828dbe35ccfc0c68748138d8588b15a

Observation 11231f11-932d-4432-a1f5-6aeaa76f3c83 · outbound

This paper cites CLIPScore: A reference-free evaluation metric for image captioning.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation CLIPScore: A reference-free evaluation metric for image captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:49.978789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:49.978789Z digest=sha256:121c01d07cf062c6dbad0bf0c3b79ec642f1144899826e09486c71de6377a816

Observation a8fd5e22-e12c-4692-8db6-0d69caa95800 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.511672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.983999Z digest=sha256:b368841ec39f46f8c70212850bbd5ac0738391f2e8d7c5c5145a06356e7538ff

Observation 477cc678-7c39-476d-863d-f3f1240a8839 · outbound

This paper cites Denoising dif- fusion probabilistic models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Denoising dif- fusion probabilistic models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:49.990190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:49.990190Z digest=sha256:2cae64eb24d79904b2b2aaba6b05189723aac3b9ec0769805af12ee9d8c67c83

Observation a8de767e-cb97-442d-b8c7-020bcd0feb7e · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.475855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:49.995748Z digest=sha256:1668c2752a72f80c0396e609d17ddec5f8c433005c1936f4871da76e34e554dd

Observation afe98a1b-91cd-4b07-8651-8c4b4b171411 · outbound

This paper cites Image-to-image translation with conditional adver- sarial networks.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Image-to-image translation with conditional adver- sarial networks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.454322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.001069Z digest=sha256:39e632cb43b597050e8ead193c07eb62c1b2423af72972c6b9badcc45401d67e

Observation cbada310-ad53-4f6e-87ae-c262512c1e94 · outbound

This paper cites Image 9 retrieval using scene graphs.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Image 9 retrieval using scene graphs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.433488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.006411Z digest=sha256:17fb358581293880d14c59f1fa8e762f5e78b43be6dea5b3a61673ef0470b957

Observation 9842cf47-36d4-4c4d-8409-6f3320f2d2d1 · outbound

This paper cites Image gener- ation from scene graphs.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Image gener- ation from scene graphs

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.413144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.012053Z digest=sha256:cc3af0c76feb3ee5ec6b188d1cf2154c6a7cd30b7e926bce492e9f7494c1b075

Observation 5db871c5-a176-4a2e-b74a-475b4b965b5f · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation A style-based generator architecture for generative adversarial networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.018142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.018142Z digest=sha256:6e23315de4139cfae821b2e24ddc3070195e771b80bfd27e08aaa5b5ca9c45ec

Observation 1ca4b1a6-2b0e-48aa-905e-bc5dac7a3c68 · outbound

This paper cites Imagic: Text-based real image editing with diffusion models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Imagic: Text-based real image editing with diffusion models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.023111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.023111Z digest=sha256:701884ccb4cb965c22c770fc5a62c2540fede265c0fb5d48edc7f82f6c35bd31

Observation 891122af-47fd-4ab8-9c7a-ac93e3ad333f · outbound

This paper cites Kingma and Max Welling.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Kingma and Max Welling

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.360983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.028078Z digest=sha256:9e6a7744b4ea09d61416ea9e30cfd3e8abc541de9a1dac00283ad42e3b0e37dd

Observation c2221a2d-eee4-4dc5-be9d-3ade3b1824d9 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.341011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.033731Z digest=sha256:d68a550b8a859b8b5a8adaaded259511d8b330ad98683e024a816ca43aba767f

Observation f0bb6deb-5265-4e44-a9dc-44e0f5d754c8 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.316285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.038859Z digest=sha256:ea50d2ac7017c57d81ee265fc6b7a44567e8f11ae1414a29151463f6dd47d36d

Observation 6b0f4928-dc5a-4d26-a2a0-391d85ed5b79 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Aligning Text-to-Image Models using Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.044466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.044466Z digest=sha256:3b41afdbbf4693c3102a439450cf1ef2b656f939477c48b5fb9c75315e58386e

Observation 40db3215-75b6-4301-a1e5-4b01cfe5ec1f · outbound

This paper cites mplug: Effective and efficient vision-language learning by cross-modal skip-connections.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation mplug: Effective and efficient vision-language learning by cross-modal skip-connections

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.278469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.050811Z digest=sha256:24d8e981d8d9bc8352f40d1fe0fea15cc8df386a1564f462f91ac1385ff4e775

Observation 1b2d84e7-35f8-4a41-927e-940d11edb04d · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.056257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.056257Z digest=sha256:56344fcc5b2d12850fde04b1d556ce5cfedfc4c6299ca47ca5c774fe460dd991

Observation 54bc075f-ead1-4c8a-aec8-d6619f52008a · outbound

This paper cites an unresolved cited work.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:22:51.242759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.061605Z digest=sha256:59d72e45eb1ec0e0fbf2e3b4e3619234c4f03286607343f24553f4feae8e51ef

Observation 9186c36c-2fae-4bd0-9723-b4ceadeba588 · outbound

This paper cites LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.066615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.066615Z digest=sha256:60d0cffa67b2fbc1878c0d590b026304599bc105d70243e507cd733d9925231e

Observation 80b32004-625d-4fcf-9035-babc41bb4424 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.072102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.072102Z digest=sha256:810d1468c4d7af1ee89744459d049e58249675a8960ce71c6eb3e6d9769a6991

Observation 645b5a7e-8c42-48bf-b911-b501d0cb2891 · outbound

This paper cites Visual instruction tuning.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Visual instruction tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.213253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.076904Z digest=sha256:10f639dbda2eee2eddf4e3dd7a9881739f203b0a9b2ac08471c4d60c4293a8e8

Observation fd975203-434b-4026-a4f5-1a48f49cfbab · outbound

This paper cites R3CD: Scene graph to image gener- ation with relation-aware compositional contrastive control diffusion.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation R3CD: Scene graph to image gener- ation with relation-aware compositional contrastive control diffusion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.195540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.081812Z digest=sha256:163dd23d0114d3614fba84b114a281f5a7a685740eb0d27beed07761af0a5796

Observation 76a9b296-6bf1-4b63-aa6b-f253424ad0fc · outbound

This paper cites Draw Like an Artist: Complex Scene Generation with Diffusion Model via Composition, Painting, and Retouching.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Draw Like an Artist: Complex Scene Generation with Diffusion Model via Composition, Painting, and Retouching

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.086737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.086737Z digest=sha256:9598c21b447e77193fa7f8a40dde11d0fefe98b76adccd094b958ddb7ce9347b

Observation 43dc9169-11c3-4859-8cb0-653b0aecb3b4 · outbound

This paper cites Compositional visual generation with composable diffusion models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Compositional visual generation with composable diffusion models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.176037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.092391Z digest=sha256:23b4d0db59e7c45999f250b7a5d1719dfc90477b2f486856db51a8717e8ab658

Observation 071e682d-d8b9-4b32-abb5-a25f76363170 · outbound

This paper cites Tf-icon: Diffusion-based training-free cross-domain image composi- tion.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Tf-icon: Diffusion-based training-free cross-domain image composi- tion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.155750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.098025Z digest=sha256:08dfe5403a8f0d58a39805ecb327042f0ad37a685254cecdb4125ec994951049

Observation 435c9c68-8a5d-4fa3-8d11-2946a0ee1303 · outbound

This paper cites Scene graph parser.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Scene graph parser

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.135886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.103514Z digest=sha256:dc0a79535881a4670d5091d534e1cd7347595c5111b95e90cbc7001cb97152fa

Observation 201c4037-0f26-4748-a973-bf41f266fcf3 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.108580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.108580Z digest=sha256:457c552e9f8eaecf0d9f2b3b2a579e69e49e5f2f15a66836c70ebcbcb89b8d96

Observation f3ca1d1a-a422-4d75-9996-cc390bc81eb2 · outbound

This paper cites Gpt api pricing.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Gpt api pricing

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.114054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.114171Z digest=sha256:20feea8f2c92907f21a99a0a954b305dd758317900ace7f343915b2acbde4c54

Observation 34e2a39f-f32a-48bd-b7fb-5fbaf7052bc9 · outbound

This paper cites Benchmark for compositional text-to- image synthesis.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Benchmark for compositional text-to- image synthesis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.094793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.122789Z digest=sha256:3532d80fc27a669d574d7e36b81a791e60f3134b37f7063df8104f31b25e765c

Observation 788f72d3-e265-4640-a4f8-e46906d149f3 · outbound

This paper cites SDXL: improving latent diffusion mod- els for high-resolution image synthesis.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation SDXL: improving latent diffusion mod- els for high-resolution image synthesis

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.077132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.128465Z digest=sha256:d28410f6e0e4fda3d7d3aee9f927f063ac4c2ff1438f1d31eec87425b2cc08a9

Observation c50d8cfc-2f60-4835-9ce0-78964044ddd9 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Learn- ing transferable visual models from natural language super- vision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.133538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.133538Z digest=sha256:a1faae25decde022286c379f296ac8ed2f47c484d999992af12712a1a5e9ca87

Observation 165f0b10-7561-456f-bdb9-946ab4d901f2 · outbound

This paper cites Zero-shot text-to-image generation.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Zero-shot text-to-image generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.044949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.139020Z digest=sha256:7471859f47009b62bd879149a48f5cca9521e2c307e693bd0bf591a8be660265

Observation 211af1ec-9754-44f1-9e28-7743f6a3147e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.146461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.146461Z digest=sha256:029a44fe96460660b22a2bae9bcbacf7e9169d2f7b5cb8d7d138e16795d6ca92

Observation e731a66b-1006-45d4-91bb-4a9d47b7d767 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation High-resolution image syn- thesis with latent diffusion models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.025935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.152678Z digest=sha256:db5d81babba36261edb36ce96b0237c231dfeab7fcf4398fa00fa1fe68f363cf

Observation 2aea27dd-606b-4556-8446-a9b914108511 · outbound

This paper cites Improved techniques for training gans.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Improved techniques for training gans

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:51.005459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.158087Z digest=sha256:86e6c75f5f07f2f736cc5e38331334ab0772af9001ad9a4595ac993b34fa7ecd

Observation d36b5c4d-0cea-4840-842c-2c7bf9465afb · outbound

This paper cites Laion-5b: An open large-scale dataset for train- ing next generation image-text models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Laion-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.985278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.164052Z digest=sha256:f5d8152c3dcc0e9a1a6a581e8b03ae90e997650e2b0c364ae879fa281ef3b21f

Observation 18d90abe-414a-49c7-bdc6-580d689d05dc · outbound

This paper cites Objects365: 10 A large-scale, high-quality dataset for object detection.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Objects365: 10 A large-scale, high-quality dataset for object detection

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.965174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.169483Z digest=sha256:f10b9fd423b716438413094037b3454ec2ebaf8bf235cb1ad3f0d60cf5383af9

Observation 80737f01-b7c2-4430-84f0-7074f1f5f0c0 · outbound

This paper cites SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.175158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.175158Z digest=sha256:847bf07cc37f4b431acdcc0ff0ce7712773c0234b405bbf72e2315b7a68f2994

Observation 3ac9e22a-be16-4331-bffb-fc2ff272f825 · outbound

This paper cites Rethinking the inception ar- chitecture for computer vision.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Rethinking the inception ar- chitecture for computer vision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.181120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.181120Z digest=sha256:2355d08e8da7f80033ebaedb1afbf4414126f05d8ffac8d2b387d3f81c68699c

Observation 8af626e1-1917-4551-b8cb-38277a91b073 · outbound

This paper cites Learning to compose dynamic tree structures for visual contexts.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Learning to compose dynamic tree structures for visual contexts

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.933113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.188095Z digest=sha256:477ff4ee286193cb794caf93dbcdd648f76c7d1924800eec35e3de7d489cb364

Observation 7b1aea14-ba7e-4d14-8711-64f961ef2f7f · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.914373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.193298Z digest=sha256:fe55e2911f7b0473161905186cc989315e28c380ec28a320802e5e977d67b634

Observation 633cea7e-7a5a-4436-aaad-b96d49d0a874 · outbound

This paper cites Diffusers: State-of-the-art diffu- sion models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Diffusers: State-of-the-art diffu- sion models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.895099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.199030Z digest=sha256:d48b024f5a91da8edd0301f119977fea51e3dfea3152cfe2353c8d69431209ba

Observation 218a99d6-4598-4392-acd4-7dd24a3feb44 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.205044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.205044Z digest=sha256:522d34af90fdb42a4dc7221d28f1f7cf4c34bc28562a7ebc8844af3860e0fd81

Observation 0267cac8-3b74-4b48-880c-3b62d4eaca0a · outbound

This paper cites Scene graph disentanglement and composition for generalizable complex image generation.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Scene graph disentanglement and composition for generalizable complex image generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.873590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.209894Z digest=sha256:7f926df47fbb495e50094f509be2dbf78236d223b108fd4a85d25a7548881ab8

Observation b59c4094-44b9-4910-901d-c181c385588d · outbound

This paper cites Self-correcting llm-controlled diffusion models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Self-correcting llm-controlled diffusion models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.853369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.214934Z digest=sha256:b4d6a7260902ddf4e175f183dc331f81fc560d97eee1a3ab08884a8852653e14

Observation d16e17d7-7833-4624-9160-dded966a5633 · outbound

This paper cites Choy, and Li Fei-Fei.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Choy, and Li Fei-Fei

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.835257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.220248Z digest=sha256:e14d70d740c2d8a70e734b0356e7aa2b157861831ddedf2653df1e3f1244a457

Observation 3bf0f42e-82f2-4a98-be05-3141c1863bb6 · outbound

This paper cites Attngan: Fine- grained text to image generation with attentional generative adversarial networks.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Attngan: Fine- grained text to image generation with attentional generative adversarial networks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.816696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.225492Z digest=sha256:38210c0e5562e6db1ba429f39dd9c8d44b26778068380f6bafa319ad18d932d5

Observation f213f97f-67f3-49d6-b64e-d171aa481721 · outbound

This paper cites Diffusion-Based Scene Graph to Image Generation with Masked Contrastive Pre-Training.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Diffusion-Based Scene Graph to Image Generation with Masked Contrastive Pre-Training

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.230992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.230992Z digest=sha256:dc995c99dc68b435b9393090f2eaf02c168af90f7a4b7deb0958572b2a7340a6

Observation 0fa2d131-30c6-40f7-a918-267aa4b99e62 · outbound

This paper cites Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.796514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.236114Z digest=sha256:b9a4c0ebe1cb089935946385b9d3c0a1e8ef493a3e3c9f5a7e21a084b181d63d

Observation bb3b278c-c863-4bf3-90ac-21ce056b99ed · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.241932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.241932Z digest=sha256:6a0231d66ffc96c0c30f4480a1ace03bef3cf90248500de4cfb4e6d545751319

Observation 39cf9e5a-7cdb-45ac-82e3-afec88e37b97 · outbound

This paper cites Linguistic structures as weak supervision for visual scene graph generation.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Linguistic structures as weak supervision for visual scene graph generation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.778764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.247304Z digest=sha256:002bf90917b79f419ec251e4ea77a9d98cc282bff911bb18b0b1a092155333ef

Observation b4f6e5cb-4b5a-4a08-8972-6279ef702621 · outbound

This paper cites Neural motifs: Scene graph parsing with global con- text.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Neural motifs: Scene graph parsing with global con- text

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.761619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.252656Z digest=sha256:ef59fac6b00153821f720a41c876d7cf7985f817896cc4605c4eb04b07e5d61d

Observation e69df88c-7713-4bb3-ad9a-f837e4e00072 · outbound

This paper cites Stack- gan: Text to photo-realistic image synthesis with stacked generative adversarial networks.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Stack- gan: Text to photo-realistic image synthesis with stacked generative adversarial networks

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.742910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.257940Z digest=sha256:98d936d9555eb45e87110f130ab8f11dd04ee69464b73b1b30b07fc7e7fac7c2

Observation 50998f56-9a54-456b-a571-6c68b24acaf6 · outbound

This paper cites Learning to generate language- supervised and open-vocabulary scene graph using pre- trained visual-semantic space.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Learning to generate language- supervised and open-vocabulary scene graph using pre- trained visual-semantic space

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.724123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.262772Z digest=sha256:696a2e72ff9020bf93eb610222a8cd97ffe7e360a7c71c93c829e6d90a9a5469

Observation d7084c0b-2690-4979-9039-c1f01a0ceeec · outbound

This paper cites Unpaired image-to-image translation using cycle- consistent adversarial networks.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Unpaired image-to-image translation using cycle- consistent adversarial networks

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T14:22:50.267900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:22:50.267900Z digest=sha256:b54a6b4d3a5a8126af60e31dda79e7c28542d46de1b370341dc338ee1290b0c1

Observation 389cd348-50c2-45ab-a069-b181721edeb9 · outbound

This paper cites If no explicit background object exists, infer a suitable one from the present objects and relationships.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation If no explicit background object exists, infer a suitable one from the present objects and relationships

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.669449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.280771Z digest=sha256:981e9ea44401dfe4641877d87f221a1bb6f91474c1c7a5bf51521fb3cf61bb66

Observation 50aab08a-3578-4ede-a8e4-8365b170cc86 · outbound

This paper cites When describing their placement, utilize the ‘relationships’ to accurately depict their positions relative to each other and the scene.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation When describing their placement, utilize the ‘relationships’ to accurately depict their positions relative to each other and the scene

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.651573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.288606Z digest=sha256:48da8188c91ed1f5fd1fb23fbe984811f333184ef244d7536b739d990a6c5339

Observation 740ddb27-cbcd-43ea-9d47-484da98c5b45 · outbound

This paper cites Avoid simply stating the relationship; instead, showcase it through the objects’ placement, appearance, or actions.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Avoid simply stating the relationship; instead, showcase it through the objects’ placement, appearance, or actions

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.633607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.295315Z digest=sha256:9f1871958fbd5b9e01846a07fa9b7a166e83208deeed97c929abc6921f54b068

Observation bbad1afc-ac3c-4525-9fa1-108a3a76f339 · outbound

This paper cites Use evocative lan- guage that captures the essence of the scene and guides the diffusion model.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation Use evocative lan- guage that captures the essence of the scene and guides the diffusion model

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.614133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.301259Z digest=sha256:c576b048ace4c133cb5e584318b612adecfbf0921530999c446d22ca18792288

Observation 01360bca-a87e-4bb2-90e8-eab9f0e84f0e · outbound

This paper cites description.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation description

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:22:50.593476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.307233Z digest=sha256:06d051da2b10b5dfd8747332e8d000469c863bd2c8b275184ddb28932464a658

Observation dce6bf83-47d3-4e18-b17d-1f2264ddffaa · outbound

This paper cites person” kick- ing a “sports ball.

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation person” kick- ing a “sports ball

Reference 2017

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T14:22:50.690291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T14:22:50.274906Z digest=sha256:3f19be179cc290539986fb49c3fb41e673c9f28a9607988437615bafb6f8ff0a

Pith citing papers

Observation 9212e703-b1af-4a6b-bfc1-e00ae40ceed4 · inbound

From Data to Modeling: Fully Open-vocabulary Scene Graph Generation cites this paper.

From Data to Modeling: Fully Open-vocabulary Scene Graph Generation What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:06.096668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:06.096668Z digest=sha256:13ac599038ec11d6e939810e218ba46683d72ae622e6bfcfb3ac315de55d63b2

Observation ed9f02f3-6e9c-4840-9651-8859e65981d8 · inbound

SciFig: Towards Automating Editable Figure Generation for Scientific Papers cites this paper.

SciFig: Towards Automating Editable Figure Generation for Scientific Papers What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T12:07:38.423497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:07:38.423497Z digest=sha256:c2a1f23c16cd19a4c51bba10f39b8d104e5015a12f2a9cea72f5aac28ea88df2

Observation 06886df0-dade-4d35-bf16-42797322a63f · inbound

Integrating Graphs, Large Language Models, and Agents: Reasoning and Retrieval cites this paper.

Integrating Graphs, Large Language Models, and Agents: Reasoning and Retrieval What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:13:30.202840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T09:08:48.468564Z digest=sha256:af81aaccf0bff536f6f19e245157b867ca7c079bd0d7fc6ab0e3f4f556b1b1cc