Pith. sign in

Paper Citation Record · LEDGER

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation

As of 7 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2507.04151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04151 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:01:50.856287Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36ea33ba-720d-4b0d-a522-6a82941ee6aa · outbound

This paper cites Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.603688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.603688Z digest=sha256:3c705594bcf44ea6b5a824c9e7267ed01d86bfbd6a6f7acb238cf9b19f3995d7

Observation c0237661-a292-42e9-b8b0-f7c521d28f7a · outbound

This paper cites Score: Story coherence and retrieval enhancement for ai narratives,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Score: Story coherence and retrieval enhancement for ai narratives,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.681445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.681445Z digest=sha256:28e49f223c6c3d645987a754eef55d9e88ffd7c64a52ec642aa7aaa67bccac59

Observation 712b153c-1902-48dc-9661-2acf37a83993 · outbound

This paper cites Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.736481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.736481Z digest=sha256:b5eabdbfbbe6411c31d261a4f458f0af52c75c589f0743222d9ee3ad52925339

Observation 471daa47-9b5e-439c-9265-31438a124170 · outbound

This paper cites Diffusion models beat gans on image synthesis,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Diffusion models beat gans on image synthesis,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:52.022091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:48.799947Z digest=sha256:ddc2a58ca7e852c9d0d017e80e78b360f727d30e9ec3c872fde2aa532851310a

Observation c1544d88-8c49-4e86-89c5-bce1d492e2cb · outbound

This paper cites Triple sequence generative adversarial nets for unsupervised image captioning,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Triple sequence generative adversarial nets for unsupervised image captioning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:52.012625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:48.852793Z digest=sha256:a6ff4b874d1607e7303c9bb7aec2f19551b267f946ea58fb46574ce143dc3392

Observation 0380e3cc-e4dd-4c90-8ed4-29e8044e2766 · outbound

This paper cites Visual in-context learning for large vision-language models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Visual in-context learning for large vision-language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.915151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.915151Z digest=sha256:ab8e7c5eb34e34c749419df89733a79a84f949f714ca3fe2b2277b70a979696f

Observation 5f51352c-1325-4fa7-8205-16168498195e · outbound

This paper cites Weak to strong generalization for large language models with multi-capabilities,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Weak to strong generalization for large language models with multi-capabilities,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:48.968028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:48.968028Z digest=sha256:b00e2006de2245a5355aef9532fd56f339f14b0d39d5e62101d774694d83701a

Observation 7c2e7e97-51ae-48d3-89ca-cb45602dea6d · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.027115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.027115Z digest=sha256:4e7a626f0eb0309c7a7518c37f866bf412ff092c0c10308b0cfef41888bc5007

Observation e769d58f-c0b2-436b-b2a5-4884d2b33751 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.072704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.072704Z digest=sha256:d039cf7599268addd9c5f343f604de146f6a33c3d01debef78a7c95eab2cca35

Observation a64bd966-0714-446e-baf7-df3e2de89e81 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.124936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.124936Z digest=sha256:af8f49500cd5cab9764e988d8fd27129dbc1369b75902716b99a894afd5d0929

Observation 3e983ef3-f2c4-4479-84e6-ed3e42306a8d · outbound

This paper cites Microsoft COCO: common objects in context,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Microsoft COCO: common objects in context,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.182156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.182156Z digest=sha256:e7d99041989170282f9b1a19db7f97738d6f1b595a123de289cd114ebad81408

Observation b51647f4-05da-472d-8aa9-a3e55bbafaea · outbound

This paper cites Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Beyond empathy: Integrating diagnostic and therapeutic reasoning with large language models for mental health counseling,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.238458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.238458Z digest=sha256:7f249acac2f1f6678778cb9e775fe8b2ec2e628ad4fad76ce5e2a441ba279d26

Observation c4bb82b5-5d46-4de5-a61e-31fa8521f3f0 · outbound

This paper cites Multimodal event transformer for image-guided story ending generation,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Multimodal event transformer for image-guided story ending generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.987595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:49.291914Z digest=sha256:b6ea99ed23b61f2d924d976e762519ae1e9eb23aecae4d20d7daeadf9a31f587

Observation 52b8e76b-5685-4a07-9ed3-d7699c55c7d8 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.979779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:49.346212Z digest=sha256:c4c19e968068f8ed886ac46d977596d7e3b7f123cbe8ff4e2de54e3437ca9421

Observation 303dce78-eb93-4e1d-af3f-46553e5732e9 · outbound

This paper cites Language models with image descriptors are strong few- shot video-language learners,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Language models with image descriptors are strong few- shot video-language learners,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.415102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.415102Z digest=sha256:a7705b275979a25fd87f53aefc17748ffbd212563ae46b7810e9becca5eb3215

Observation e2b9ee77-6052-4fa3-a5ca-aa9cdc7d36a8 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Learning transferable visual models from natural language supervision,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.477464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.477464Z digest=sha256:09a8b2901e217e02cc84b9c920c851817a7769bbde21dd8d4f89229e898ec3aa

Observation f6b2fe0d-0c95-403f-b6f8-29463105be68 · outbound

This paper cites BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.537759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.537759Z digest=sha256:d34ea80acc02b6f053abbf3418345075f2b5b12ad0725f104f3ca0b68c8b2ae4

Observation 11c69b06-b6ac-46c8-b741-6f2949ee603e · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Flamingo: a visual language model for few-shot learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.950631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:49.665137Z digest=sha256:a599ec76176a3ad98cff3fd571a8cca586f8373b4457866fe57ab59245381881

Observation c8d1758d-90c1-4e67-b954-f884343bde56 · outbound

This paper cites Cross-lingual transfer of large language model by visually- derived supervision toward low-resource languages,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Cross-lingual transfer of large language model by visually- derived supervision toward low-resource languages,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.933684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:49.777370Z digest=sha256:b8c2a583db584c331d790d87b0c69c0db9826965a740fe0fbae5b7ec129a8e7d

Observation bbb1fff5-f82a-4315-a287-ce1fe51623f6 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Coca: Contrastive captioners are image-text foundation models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.857465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.857465Z digest=sha256:1b5280a2f06266d8d55e7e54bb3ddfe38d0fa8ac0673160b8c873b1dd0beb374

Observation 8a8f4c49-5753-4a90-8d2b-349c24d96839 · outbound

This paper cites Tx-llava: Large language and vision assistant for temporal changes in chest x-rays,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Tx-llava: Large language and vision assistant for temporal changes in chest x-rays,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.932790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.932790Z digest=sha256:03386ac5333ae612de255939d5c65c5b50c1e587068222d5b96f9168dc27698c

Observation c0f02167-b11d-450b-84ee-efd249df8742 · outbound

This paper cites MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.990147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.990147Z digest=sha256:324d7b59819687cc84e3292a8729441d2d32d641fb909c908c0f2ae7076871d8

Observation 862c957f-faa2-484a-8d51-e71cc023b4db · outbound

This paper cites GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.919344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:50.130327Z digest=sha256:b0df74364d2de5f684eaf4faf33d2077e8168c99a6b26b43a613fc8e0140b1aa

Observation 297a636c-188f-4afa-b695-eec47deb6d44 · outbound

This paper cites Taming transformers for high-resolution image synthesis,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Taming transformers for high-resolution image synthesis,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.910644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:50.236512Z digest=sha256:091d14a99bdbeea5c0e5cfd8a60a85cc6defa9c88f8d481d7f3339e3665a131d

Observation 58fbf2b1-988a-45fb-aa34-34272b0e1b4e · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:50.429219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:50.429219Z digest=sha256:d57c2b1786aa34d88812de95511a1bda391bb940c1e3c96c0159e6f731f38f88

Observation c9bfb984-33e0-4f54-a912-407aa9600c49 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.891777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:50.532097Z digest=sha256:c0e0c478430d2be99298e772eb1e97aec8b1af66fc9254e99e10f5047a4dc74b

Observation 23f68944-0b40-4f1c-8bc8-6956db206c4c · outbound

This paper cites Self-rewarding large vision- language models for optimizing prompts in text-to-image generation,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Self-rewarding large vision- language models for optimizing prompts in text-to-image generation,

Reference 27

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:01:51.120901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:50.623402Z digest=sha256:c44723f18f53197f7cc1bc2e43ec08931849611239c48bfbcf491802c82def2b

Observation a03b9272-b4da-457a-93c2-b024dea37138 · outbound

This paper cites Evolvedirector: Approaching advanced text-to-image generation with large vision- language models,.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Evolvedirector: Approaching advanced text-to-image generation with large vision- language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.618270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:50.782997Z digest=sha256:32b47cafc311b5fcbeea32b47da80568df8fcc57a34d148e55c6275fa0e79212

Observation d773c6ae-7ce0-4acc-90e3-fc520c889a19 · outbound

This paper cites Improving Compositional Text-to-image Generation with Large Vision-Language Models.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Improving Compositional Text-to-image Generation with Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:50.856287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:50.856287Z digest=sha256:bd3ad4dbba3fe0004d71b6e9e47fe769ad56252f3c38ac28d14fc950942bce23

Observation ce90422c-0cd8-4b94-aee9-40bbe469c126 · outbound

This paper cites 12 888–12 900.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation 12 888–12 900

Reference 162

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:49.603500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:49.603500Z digest=sha256:41bf1e29bcb5f47e83f03d1289d4f88725decfbdb555699d2b64b31a4ac54cc9

Observation 36bc1f5d-8b4d-4268-81a7-3680fabc8de5 · outbound

This paper cites 12 873–12 883.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation 12 873–12 883

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.901437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:50.329073Z digest=sha256:0659f70ceccf909a7f540c851d2761f6276bd6091dd03c7e4e38db0c51628b18

Observation e42b8339-0647-416f-ab6f-96e8422d85ea · outbound

This paper cites Available: http://papers.nips.cc/paper\ files/paper/2022/ hash/960a172bc7fbf0177ccccbb411a7d800-Abstract-Conference.html 11.

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation Available: http://papers.nips.cc/paper\ files/paper/2022/ hash/960a172bc7fbf0177ccccbb411a7d800-Abstract-Conference.html 11

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:01:51.942532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:01:49.714980Z digest=sha256:22c021900c1a7d2e37e40d5b1fca9ee647e873fd57a28bbfb1bf0189aec760a5

Pith citing papers

No inbound Pith citation observations are available.