Pith. sign in

Paper Citation Record · LEDGER

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation

As of 14 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2412.05818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05818 v2

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:25:15.110746Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:30:27.103280Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:30:27.496049Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact0
  • verified fuzzy42
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cea0948c-e204-4389-98f3-0ab474309b98 · outbound

This paper cites A general theoretical paradigm to un- derstand learning from human preferences.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation A general theoretical paradigm to un- derstand learning from human preferences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.425670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.425670Z digest=sha256:7b03b87a452d3e4f33ffc272ebce1a8bfb1fa5dd44d4761d51012d908170f856

Observation 87972546-4366-430f-99ce-a603d9a9cab7 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.432116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.432116Z digest=sha256:c185e1e6c057af79bae6a38e8679c3fec1ccc46d56ca82024d3ad956aea034e6

Observation 3e2308b4-26cd-4692-a7b1-0ab937bf4918 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.438022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.438022Z digest=sha256:1dca771b9ed7881fb3f560a15838b8b6ead7c2215f5516c726ab6d282bef6609

Observation 0c51e86d-9037-43a5-b399-b44baeeca17f · outbound

This paper cites Improving image generation with better captions.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Improving image generation with better captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.444537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.444537Z digest=sha256:15cf807074dfc1f88aa4d2ea330c949c999f907559e9f9c626e5cc3b79c795ac

Observation 9df390e6-4d00-4133-b67d-81db4ee16d9b · outbound

This paper cites Training diffusion models with reinforce- ment learning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Training diffusion models with reinforce- ment learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.450209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.450209Z digest=sha256:c6b331d93fd8632edf6cf215b213701428fd9a4d88782829818a8c8de81201dc

Observation e9e42c48-eed0-4469-9c66-9a18b80a00d9 · outbound

This paper cites Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.456093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.456093Z digest=sha256:1ece11470641ca24eddb741a931a23ae6e34eb336303483b0e280bb009c578df

Observation b83a1fc6-8c7b-4ed9-bc02-a5b7bcb64cef · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.461932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.461932Z digest=sha256:2acea7fc949c585385e9b7a484671feca7b6250cbde84ebb5ebdc04892d3928c

Observation e8d922ed-9047-4648-8215-7797aa739e7b · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.467979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.467979Z digest=sha256:00d81bf487484d99a183c21eba8da91423cf1015ee738f01ebcf496de8aae13f

Observation b939c50c-136c-4707-bb3e-2aceab3c850a · outbound

This paper cites Multimodal dialog systems with dual knowledge-enhanced generative pretrained language model.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Multimodal dialog systems with dual knowledge-enhanced generative pretrained language model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.474582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.474582Z digest=sha256:b6a2153652dbdd3eb48da2dbe07177de81971018033dc5afcecdba37dba3b6cf

Observation 5c756234-837b-4fc9-ae05-f335eea6c755 · outbound

This paper cites Ultrafeedback: Boosting language mod- els with scaled ai feedback.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Ultrafeedback: Boosting language mod- els with scaled ai feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.479995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.479995Z digest=sha256:2519af5a84cf954e204f13f431db35e135dc7f13264eeba7c5a1ba75f82ef23b

Observation e2b9dcb9-7d2e-47b7-aea4-1fbbb6f6809e · outbound

This paper cites How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.485181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.485181Z digest=sha256:267165f91d0cfe59b1c273e7792b42e47b47f47d4737911ed22c3e9235782348

Observation 570ae0a7-16f7-49e8-a123-ccc5b30618cc · outbound

This paper cites A Survey on In-context Learning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation A Survey on In-context Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.490817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.490817Z digest=sha256:404cc3486c56ac59f9c8784dd2c3f18e9e0eb28e3ba80c75ac4684d1c83d50eb

Observation 86bce015-8c03-4ab8-bc54-ea9998cb64d2 · outbound

This paper cites Dreamllm: Synergistic multimodal com- prehension and creation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Dreamllm: Synergistic multimodal com- prehension and creation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.320857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.496137Z digest=sha256:83ee27dff0d7d0f4b43266dee2ee94a3a8948b9fc8fcacfd860cd0c86e80765e

Observation 0e54758c-4964-4bd5-a531-493f70c7c7a9 · outbound

This paper cites The Llama 3 Herd of Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.501287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.501287Z digest=sha256:9be2ed9c10341e5738259ff89c4b57bd9d2fad638bbdae8a3ed7da990bc4a167

Observation 9b6f0ada-8bab-4cb2-a5b4-672a6c967600 · outbound

This paper cites Re- inforcement learning for fine-tuning text-to-image diffusion models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Re- inforcement learning for fine-tuning text-to-image diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.301343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.506597Z digest=sha256:b8e735eae500d2ee6e7bcd7f60d2289be523282aeb0460e81030ae445b43e798

Observation 8ac07f17-bf7b-44c3-b003-1f728254dd55 · outbound

This paper cites Training- free structured diffusion guidance for compositional text-to- image synthesis.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Training- free structured diffusion guidance for compositional text-to- image synthesis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.281347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.511744Z digest=sha256:1133073bc17c4b18d46af8605eb5ef553a3c30ebf4708d7af6578086528fa50e

Observation 32ed574c-a15e-4127-b44e-fb49d7ed90a8 · outbound

This paper cites LayoutGPT: Compositional Visual Planning and Generation with Large Language Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation LayoutGPT: Compositional Visual Planning and Generation with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.516822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.516822Z digest=sha256:63f48a6000d0065be63c7f284365b6e63f894cbfb7b7d3b38ae481690c400beb

Observation c1380e6d-aa3c-4550-8c78-1ec7694fda77 · outbound

This paper cites Dropout as a bayesian approximation: Representing model uncertainty in deep learning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Dropout as a bayesian approximation: Representing model uncertainty in deep learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.262716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.522862Z digest=sha256:e28051cf00d49ade73624b296e060d6e484ddb05a6a580fd3660013855d08192

Observation 76b2b88d-6473-4fb6-9d22-063fcbf954f3 · outbound

This paper cites Making llama see and draw with seed tokenizer.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Making llama see and draw with seed tokenizer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.242557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.529239Z digest=sha256:0ecf9eb1b4fa841f41f2ac179d0d239baeaa9e98f18d84b246b536ebab7e1e9f

Observation 43f55e14-52cb-43d9-87de-c429d821c81b · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.534962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.534962Z digest=sha256:06c2ee4e5ef6b405460b37560df79a484b149397dd15872e0d11ead5ef0e8bce

Observation 02dd1084-32f7-4518-8fcb-7253bb27af22 · outbound

This paper cites Pela: Learning parameter-efficient models with low-rank ap- proximation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Pela: Learning parameter-efficient models with low-rank ap- proximation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.222788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.541064Z digest=sha256:b2a0183e6e1ed7b79220773d82bb0f6f43615a80efce054e3c8cdc76e2d1042d

Observation 37d22593-872a-4da7-a0e7-fa0b0031ee45 · outbound

This paper cites Hearst, Susan T Dumais, Edgar Osuna, John Platt, and Bernhard Scholkopf.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Hearst, Susan T Dumais, Edgar Osuna, John Platt, and Bernhard Scholkopf

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.202807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.546799Z digest=sha256:d85028d0302e5e3bb62beae05ee3f00af2d81cb45253340603293a5ed65d52f0

Observation b6104bc7-3929-4b4f-b642-bbec0342d082 · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Clipscore: A reference-free evaluation met- ric for image captioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.552087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.552087Z digest=sha256:4b12223c2238ed98948edc9fd425454c10f3408c603a75e559936b3ac3c0d19f

Observation ef7b8ac8-23ab-45b7-8521-232b896f4edf · outbound

This paper cites spaCy: Industrial-strength Natural Lan- guage Processing in Python.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation spaCy: Industrial-strength Natural Lan- guage Processing in Python

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.162098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.557758Z digest=sha256:b1ab3e5dbf0ac97bd76d18451caefe087b1b1d9e3016e6b7e805c80a2f8a405d

Observation ca859cb2-f157-4de5-a242-4b22b7b76119 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.563093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.563093Z digest=sha256:4b32e2adc3bc663f72ebcf9d103722234c63355fdbad99ca61e249dc467a84fc

Observation 8ad3bfda-995a-42fd-8007-f4b48db7f005 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.569982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.569982Z digest=sha256:07aad81967b03662563d9600b33220dc53317da50acb840808825127062cda56

Observation 8c13b01b-5833-48ac-aca2-b64f90424c1c · outbound

This paper cites Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.140222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.575294Z digest=sha256:bd9a8b67ad864c4dbe6c49469d7364cda2c9d2511bc5f165255a3ad504a0681f

Observation ba8b844a-14c1-4197-9215-addeb8761065 · outbound

This paper cites T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.109749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.580288Z digest=sha256:b0a50a59c8ac22ebfe58584248123457acd15fb7366baf261fd25be6ce846f36

Observation 89755587-7531-4663-8ead-3d3171a34ae1 · outbound

This paper cites Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.085843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.584846Z digest=sha256:05daa87038c9a31163e287b37d3e42991d85863099f86e7840d36f0503e78452

Observation be89228f-fb0b-45fc-b0d1-a2e725cbf520 · outbound

This paper cites Human-centric Dialog Training via Offline Reinforcement Learning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Human-centric Dialog Training via Offline Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.589831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.589831Z digest=sha256:ca23b761439617b1df46b7e8f2b1f9ba5edb9d95352e48be15cbf84e4eddbcfe

Observation c8dac8bf-9d68-40f7-9015-351ea247033e · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.059195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.596060Z digest=sha256:a3a900ac192289df70b0d1f0ecc797a07d740944e29fa5c17dc126f79f0fae71

Observation 408d493a-9ba9-43af-bed1-d49fb29ef99a · outbound

This paper cites Rlaif vs.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Rlaif vs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.039698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.601314Z digest=sha256:715af75dc65e8ae6a79d52e9fae8f12a601adca931bc315d949951c1c15bbfab

Observation 2bfce24b-305f-4c3b-914c-5e927ee7b661 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Aligning Text-to-Image Models using Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.606432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.606432Z digest=sha256:42a4a9bb997ebcbcaa0b0c0c2b4450232f5c9f1a9173d0b9cb3dd35bd5a477a1

Observation 35d94c09-4032-4761-942c-a42bd99ed973 · outbound

This paper cites Invariant grounding for video question answering.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Invariant grounding for video question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:17.017086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.760767Z digest=sha256:9449dc36ee1474f6cd3ef0eb35746422f00143b86623c0b4317b158dd3880dc8

Observation 5228c811-4ab1-4013-a631-82f86b8ad6aa · outbound

This paper cites Transformer-empowered invariant grounding for video question answering.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Transformer-empowered invariant grounding for video question answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.997434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.766878Z digest=sha256:06600ff2b5e65aa8ad95fded6a2aad1cd76a517639136aebb5b0415fc0d21ede

Observation 70999c1c-a7c6-4b16-b553-49a1a8147b13 · outbound

This paper cites Attribute-driven disentangled representation learning for multimodal recommendation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Attribute-driven disentangled representation learning for multimodal recommendation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.975428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.773886Z digest=sha256:b994e411c35251d6208e478bd3fb0a420696a7467ff9a743c765c9358ce03964

Observation e06eca76-f3a3-42bc-ad0c-6911f2a7fef2 · outbound

This paper cites LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.780369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.780369Z digest=sha256:80d05529ddddd92eb68ac845b5a24331bdbff50053de6742327eeed07366ea92

Observation 320e6226-04a0-46c8-b2f5-c38b42ce22c9 · outbound

This paper cites Llm-grounded video diffusion models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Llm-grounded video diffusion models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.952412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.786456Z digest=sha256:351829c22e708f6cb91ee4a903ff7aeb79d8e681651ee0e9571e69baddf7fdca

Observation 0e70dd17-b9c3-4039-93f3-576100664f06 · outbound

This paper cites Visual Instruction Tuning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Visual Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.791940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.791940Z digest=sha256:3a16169c7d5f0f4b242a7a6ce77774ead23b91e5eab79fc52d12c995fe115e9e

Observation dd9190e3-a68e-4306-8cb9-74b5bb6d8b48 · outbound

This paper cites Ipo: Interior-point policy optimization under constraints.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Ipo: Interior-point policy optimization under constraints

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.931447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.798321Z digest=sha256:683ecbf3747ff238fcef408fe01df31c56b5be131e4e5a2cd401a080a2fbb7c9

Observation ae637458-13c3-4aad-809b-7d8353dda7f6 · outbound

This paper cites Cheap and quick: Efficient vision- language instruction tuning for large language models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Cheap and quick: Efficient vision- language instruction tuning for large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.906814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.805016Z digest=sha256:a6928f8c51654dfbd3c4507d938de5c7e0e3e303251e4354c656d1383e76e43e

Observation de3e33b1-a0d0-4d4b-a91b-a0e450e21476 · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Compositional chain-of-thought prompting for large multimodal models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.880847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.810931Z digest=sha256:eb4a1831bfcfa07bda4d5aeb6d924fdbe4c56adbd0af9abb4f6347135b50ed27

Observation 7c854fc5-b474-4022-902e-04a7e97c823c · outbound

This paper cites Training language models to follow instructions with human feedback.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Training language models to follow instructions with human feedback

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.857710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.816283Z digest=sha256:c56e7cf3bf5ae38a15b71c1c3521ccd3253e5d13cc14ca88421750d6363d3cd0

Observation c90f62e0-ffb5-49f2-ad00-1f67506be9ac · outbound

This paper cites Lan- guage model self-improvement by reinforcement learning contemplation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Lan- guage model self-improvement by reinforcement learning contemplation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.830206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.822176Z digest=sha256:f775c68b039d90901d02b183f766fbee93d28c65b2d2967c65da2c0b68a2e0cb

Observation a10cdbb2-3f62-43ad-b1a6-358bb7a8d34e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.828251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.828251Z digest=sha256:0e932a207cbb45c43762c1e4f2bcb00d120df3a210bae2f1128dafc5f68a3dd7

Observation 8b8f96f5-f778-4f94-a4aa-efc8f256dfda · outbound

This paper cites Diffusiongpt: Llm-driven text-to-image generation system.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Diffusiongpt: Llm-driven text-to-image generation system

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.835904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.835904Z digest=sha256:2726654e0d6afef2c634f6e30b73e7a55fad0d600b45067c5a692a40de0b01e4

Observation 1c91c364-838c-42e7-9248-30a6077d8814 · outbound

This paper cites 3d-immc: Incomplete multi-modal 3d shape clustering via cross mapping and dual adaptive fu- sion.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation 3d-immc: Incomplete multi-modal 3d shape clustering via cross mapping and dual adaptive fu- sion

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.801841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.843001Z digest=sha256:a752ccc0ba2828679360d33f4dc6a36d37000a386b1b17389d966c1b94a7cee2

Observation cc049737-d86c-4483-beab-17e02817bdc0 · outbound

This paper cites Dynamic modality interaction modeling for image-text retrieval.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Dynamic modality interaction modeling for image-text retrieval

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.780318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.848502Z digest=sha256:0722f339d2eaf9258bbdadf9453e8eb4f2ff809f8979aae46d9cf7666e63d078

Observation d33920c6-87e8-4169-a282-5973f0c44eae · outbound

This paper cites Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.755881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.854363Z digest=sha256:0c3bbe709acbfec52171b7ab7a9d4823a125038166af788692c79e088d2081a4

Observation 887a7e9c-0f78-47e2-85d2-337131800f72 · outbound

This paper cites TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.862417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.862417Z digest=sha256:0864c75a5f1712393023695cf7afe5c68c2ca5613b841168b2e3a4cf74f600ff

Observation 84421468-b731-4c91-a58d-75318e869631 · outbound

This paper cites Discriminative probing and tuning for text-to-image generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Discriminative probing and tuning for text-to-image generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.716643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.869638Z digest=sha256:8f43ebb73f2f3fae6217d9f14ab1cc10289f5cecfe661586c292aa6247e897e1

Observation 9eecad53-b581-4051-a8a3-f31dffde9fcd · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Learning transferable visual models from natural language supervi- sion

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.876127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.876127Z digest=sha256:bfa98b575811a4770001e64e2b9e331978d7ceb08eba42d0e32c8e7b70a4d7e1

Observation c18a1184-669e-414b-a192-a2b12e67b21d · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Direct preference optimization: Your language model is secretly a reward model

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.676752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.881529Z digest=sha256:77b0215988e6bde368ab36966db8a82a12b163de7a12187ea86007aeaa9c5ed0

Observation 35a3e8a0-81de-42ef-8dbe-c9c8a2588edd · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.888536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.888536Z digest=sha256:71466b6e852de17be3d8a2972218aa0d9f77ceee3b2079b5a58b2cae6eac715e

Observation 0f48b242-ce8c-4d77-b59a-fc39a8394717 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.895419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.895419Z digest=sha256:c7152709c54abe7c184e63b37ff46bc950d8d7906b79d221a2bc4a7efa3bec5e

Observation a9bda355-e586-4b4a-81a4-b96b7f080455 · outbound

This paper cites Can generative multimodal models count to ten? In Proceedings of the Annual Meeting of the Cognitive Science Society, 2024.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Can generative multimodal models count to ten? In Proceedings of the Annual Meeting of the Cognitive Science Society, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.641581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.901138Z digest=sha256:7b83b735a325786d30d3400bc2754bb4a060a4feb8aa120e87d20d648539f0be

Observation 539b5c37-2b5b-44a7-ae54-0c49d9ba829e · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.908071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.908071Z digest=sha256:77ff8c2003e4fa7f0ca2d010685bc2b4b9d7eb723e3fb10c45ccdd0203ec908d

Observation f7c48f5c-b5ff-4968-b58b-94b54d96ec4b · outbound

This paper cites Proximal Policy Optimization Algorithms.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Proximal Policy Optimization Algorithms

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.916590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.916590Z digest=sha256:e18f0091bcf1ba167e1bd6971ae7649b83fea0d7125726494f4bc52749b164b6

Observation a4fc7d2c-1932-49f0-a8fc-d4745571505c · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.925715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.925715Z digest=sha256:9fedad89dd8e8a763e7d357afdaf51e3fa59cf0de7c5a0abcc80a35faf292075

Observation 2e42345c-659a-491c-86d0-5f8356bd7620 · outbound

This paper cites Kernel methods for pattern analysis.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Kernel methods for pattern analysis

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.602990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.936104Z digest=sha256:2f684e630d1c34b41ed33773557498420efd5eaf204388a66f6f7f87468a2a43

Observation fb1c32aa-617f-4a1b-bdb7-e33bea58efda · outbound

This paper cites Learning to summarize with human feed- back.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Learning to summarize with human feed- back

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.574025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.942597Z digest=sha256:b20ba810391d68e7cd3e416ca018e03e923bdc95c9f817889d947b363058383a

Observation 6db52ec4-4bc8-4b8d-9d1b-80ca89fab095 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.948321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.948321Z digest=sha256:2f097ad902f3e857667920090814f7b82b0bf8e4f909075330f2aa0a4f19afb5

Observation 4a2f98ab-2a47-461a-af6e-df5fc9dad34f · outbound

This paper cites Emu: Generative pretraining in multimodality.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Emu: Generative pretraining in multimodality

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.550116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.954167Z digest=sha256:3ba6bbf8e7edde716afe21bf595ad52b4bc861dae7b5ce7be9eefe2acae7284e

Observation 4a06d198-7078-4c46-90fa-f3e6db573399 · outbound

This paper cites Generative multimodal mod- els are in-context learners.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Generative multimodal mod- els are in-context learners

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.525803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.960207Z digest=sha256:1288390c15a109c85b8968a8060b1ac31c32f30cd8c69dc49db81c6c005f139f

Observation da1636db-9755-48d1-b737-1fc8490a3eba · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation LLaMA: Open and Efficient Foundation Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.965603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.965603Z digest=sha256:a83830d78fcf2d4cda87cd00f527a764ee001d0bd10c62aa221912bb71c30473

Observation 27cde872-daf7-4a47-8a95-a8ecee165677 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.972615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.972615Z digest=sha256:cb203cd248b43ff47fa5419e973cb436bce7d85cf9bf4fd9765b5fff1f692127

Observation e784b1d0-b4e5-4819-a765-59301aadd65e · outbound

This paper cites Diffusion model align- ment using direct preference optimization.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Diffusion model align- ment using direct preference optimization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.476027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:14.977961Z digest=sha256:a2a223c66d843bd8cd0bb022ae369f9094253b1df54b00d518c0da2934154240

Observation 7f86175a-2f06-488b-af03-f0d5a78f6480 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Emu3: Next-Token Prediction is All You Need

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.984003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.984003Z digest=sha256:07163bf95389d228c196bbdfeae50e9a049c01b1fda4d265cbadc1b0871316f6

Observation 386c3901-3ef0-45c6-8a8e-f23fb05974a4 · outbound

This paper cites RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.989517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.989517Z digest=sha256:3a7055eb9b5fed9d529618e6bf5b5d793c3b13608ccd984a472afe9011e0e9e4

Observation 23e8be3f-bce4-4c19-8a68-135c0d0f84dd · outbound

This paper cites Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:14.995830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:14.995830Z digest=sha256:cb2a0fba269964748501febf3fb37ca1fa13d6ecdcb0ac2805f92bdbea98bd8e

Observation f4379731-fbbd-4494-8e0f-4dcbafa00c41 · outbound

This paper cites Comprehensive linguistic-visual composition network for image retrieval.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Comprehensive linguistic-visual composition network for image retrieval

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.454831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:15.001428Z digest=sha256:abdcc1317e6476f4a867a2cf65eccf761a8b4f39c81783a2f38e3a201c04fd8d

Observation 8156953c-c6eb-45bc-a587-202660fd0e40 · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Next-gpt: Any-to-any multimodal llm

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.430022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:15.007045Z digest=sha256:112bde56d6d84a91ddec06928cef02401ab1ef423bf4f2c8c8fb56401121ae26

Observation b6b00d45-594d-4e44-aca3-e58920d56492 · outbound

This paper cites Human Preference Score: Better Aligning Text-to-Image Models with Human Preference.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Human Preference Score: Better Aligning Text-to-Image Models with Human Preference

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.012326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.012326Z digest=sha256:184533a98eef23c473e23fac2e117dfa603ffa11e896e7b4ebbb3082c6d3c0dd

Observation 899eeccf-28a4-4e52-9f08-01e9cb958c8c · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.018504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.018504Z digest=sha256:61648f023ef285b670400039b420160778dfd80006bd7a92445d649679f90a1c

Observation 18dcb2bf-e914-4e33-a6a3-57e8434978b2 · outbound

This paper cites Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.408048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:15.025317Z digest=sha256:e757bfa2c8a56231465e24aa5b7eed7f725fed697e4e7cebee2ffeae45126ad8

Observation 187bde08-e143-4317-ac54-438a1d50ef05 · outbound

This paper cites Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.031298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.031298Z digest=sha256:5d694847b3b3173c1b809eb334e9b1bdcabd2b8d49775002608c8313f1b89664

Observation 81c91a9c-c72f-4739-b8da-d163f2aba32e · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.036973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.036973Z digest=sha256:1b5411170dad99ed80e6a7aec4eb217f113054f7693017386b4ad97d89e7c2b3

Observation 01a12e6a-dc3f-46e9-aa50-b0f1a6c0d1ca · outbound

This paper cites Self-Rewarding Language Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Self-Rewarding Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.043167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.043167Z digest=sha256:fdfd96a7a0e9ea028ba2f680cbef64a7e5f712dd7ff711f33f7ab0be9817cbd6

Observation ce8d0df0-8e3a-43b9-ace0-8fc071093344 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.050001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.050001Z digest=sha256:2d3b5fcac990d7b081f7c9bf2ed977c55afd21cb2e74a01aa393fa290f78dd5b

Observation 32ef74f3-9fd0-437e-b0d4-9ec92225f51a · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Fine-Tuning Language Models from Human Preferences

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T20:25:15.058528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:25:15.058528Z digest=sha256:5cd8f72fdb434f1603d38d3392e2328e882812fae81e92ded9cab910a82b7e29

Observation 2ffb9301-f40b-4731-beeb-4e9a3fc017c8 · outbound

This paper cites Attributes such as color, shape, texture, and 2D/3D spatial relations are also incorpo- rated.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Attributes such as color, shape, texture, and 2D/3D spatial relations are also incorpo- rated

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.374352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:15.064281Z digest=sha256:11a1ac080c6c946a1fbc1a6f418315b236fbb6b399460370f1a24b244518dab4

Observation be913525-b847-4757-8445-144e73a0bfb1 · outbound

This paper cites These atomic concepts are then transformed into simple yes-or-no questions.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation These atomic concepts are then transformed into simple yes-or-no questions

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.347768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:15.072767Z digest=sha256:9823985eaafc1402740069b4a581448deeba31dd6a95ffd7bbd70eaa991e9208

Observation b76f1ddb-96aa-4c1f-9a24-66df37816450 · outbound

This paper cites log σ − ¯σβ 2 LX i=1 ∥hw i − µi∥2 2 − βC + β 2¯σ LX i=1 ∥hw i − µref i ∥2 2 + βC − β 2¯σ LX i=1 ∥hl i − µi∥2 2 − βC + β 2¯σ LX i=1 ∥hl i − µref i ∥2 2 + βC ! # = −E(x,zw ,zl)∼D.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation log σ − ¯σβ 2 LX i=1 ∥hw i − µi∥2 2 − βC + β 2¯σ LX i=1 ∥hw i − µref i ∥2 2 + βC − β 2¯σ LX i=1 ∥hl i − µi∥2 2 − βC + β 2¯σ LX i=1 ∥hl i − µref i ∥2 2 + βC ! # = −E(x,zw ,zl)∼D

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.323140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:15.080018Z digest=sha256:1c0c4393c7a56f1e068bc52a3dce3030275bac96b921ebae280746bc8569023b

Observation ef46e847-313a-4933-8d7b-ac3157509c1f · outbound

This paper cites For SEED-LLaMA, the LLM backbone of DreamLLM is optimized for 1k steps, with a learning rate of 5 × 10−5, 100 warm-up steps, and a cosine learning rate scheduler.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation For SEED-LLaMA, the LLM backbone of DreamLLM is optimized for 1k steps, with a learning rate of 5 × 10−5, 100 warm-up steps, and a cosine learning rate scheduler

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.294856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:15.086924Z digest=sha256:4379c5a7114d38aa042ff6be867d0f5ebcdd29ef317c872bd6e91d38f683039b

Observation b059f2be-e759-4159-b1db-cc5fd264105a · outbound

This paper cites an unresolved cited work.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:25:16.274373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:15.092981Z digest=sha256:a820f67c5646daf4346a7a251688c829c6552f47b06188068f2cf2d92e36a2bf

Observation a0e4efb8-59b7-445c-b30c-c3f8d2abb4f5 · outbound

This paper cites First Half.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation First Half

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.254438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:15.100341Z digest=sha256:6ce085b05079544acd0ed9dd9473d618d9850f5797376fedc9c9001afb8c9cd1

Observation 625205fb-3c24-4363-8277-4bc54bb41e6f · outbound

This paper cites 14 - 24” means the rejected data points are sampled from rank-14 to rank-24 which is a hard range, while “20 - 30.

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation 14 - 24” means the rejected data points are sampled from rank-14 to rank-24 which is a hard range, while “20 - 30

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:25:16.235360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:25:15.110746Z digest=sha256:5a2086b05fe51680de3e9498539d00093c19110ba7b2b7ab944c52f2ada41db4

Pith citing papers

Observation a6bdf839-3534-404d-a26d-b0076853d7a6 · inbound

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards cites this paper.

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:30:27.503745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:30:27.103280Z digest=sha256:6a96e042be98c32e1a60d4094d77e6e6b1d48fe0dbbea0b4ca294133d854d73c