Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:26.661269Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2504.12018.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:26.661269Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a039c636-3c55-4eaa-80e4-ffcb035b111d · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e60fa8f0-7b8c-483a-8b78-62f47d8d10ec · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Kandinsky 3.0 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1da81de3-5e3b-410c-8e34-777f46280c87 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6942d79-9386-4c24-b0e9-c94d1b18d996 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e51052a3-f2ff-4b79-8db3-b23f4fd2cf92 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 14ca790b-6790-4ff4-bdd8-0f5c54d2038e · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d4418fb-e54f-4c40-8b10-cfdbdf92f72d · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82dc28c0-d4dc-4fbd-800d-7edc8d30ce76 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44858ea2-b495-49c0-b23d-38c854cc6d35 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Dreamina
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c5fdb2b8-f012-482b-af9b-96683ee950a4 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Scaling recti- fied flow transformers for high-resolution image synthesis
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a77dea49-59fe-459e-b1c2-2fc6a05cd78c · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73ddecb0-e4a0-4d62-b9b0-dec03dc27092 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b129cfa-67e8-4506-b066-9bffb096fd3a · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching NTIRE 2025 challenge on text to image generation model quality assess- ment
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0efaf066-a7b3-4984-84e1-8003036adc20 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b509637-dc17-4219-ba9b-4f04119afc3d · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6203e11-aa58-4904-a891-662c76648e7d · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Midjourney
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 20e112cf-ee31-475f-9882-5693aa838f24 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Lora: Low-rank adaptation of large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7aa9f23b-e078-4099-8f62-4668b2dcc5c2 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 228484db-f87f-4904-932b-b0163dae7e69 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Pick-a-pic: An open dataset of user preferences for text-to-image generation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 41b50943-5093-4153-86d8-35780eaca9f1 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Evaluating and improving composi- tional text-to-visual generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 94a28a47-d2fd-4763-b475-f139e210616e · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3044d9d-2f82-4d73-9e24-07b074d0c87c · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac6b9e6-7cf5-4f49-9e6d-cb4c928b463f · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 700c6fc0-ddbf-46e2-ac51-a1c60f1bdd0a · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Visual instruction tuning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f90556e0-aeb3-4d94-91a9-d8b7d5583717 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Improved baselines with visual instruction tuning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76aa0313-43c2-454b-aa84-1bd9d64fe4b7 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Llavanext: Improved reasoning, ocr, and world knowledge, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6487a08c-b4e9-4dc1-bfc0-013f5d676c53 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cef6bf6-12ac-4b35-b56b-21fb16ccee8a · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191beef7-3212-43e2-93c6-bb505d01d53f · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Learn- ing transferable visual models from natural language super- vision
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cb6d91b-0910-43c7-9e2a-97883407daf7 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24505366-057f-4bdc-974c-7527ec0aa346 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching High-resolution image syn- thesis with latent diffusion models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 677aa1d1-1f91-4a58-b64d-8021cd5b4c79 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Adversarial diffusion distillation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cc23c617-bf0b-4ca4-9609-ba611575f1ed · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 83dae9d7-bcc8-4359-a459-81d9aa2ac634 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Gemini: A Family of Highly Capable Multimodal Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d12f04cc-a6a6-4592-ad0a-8fa2761096de · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3eac9ab-a837-4762-bb2a-efdca2f4f1c4 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Chain-of-thought prompting elicits reasoning in large lan- guage models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aa4a7cf8-23cb-4208-9fa9-1ba2a7e0dada · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4557b27a-5f78-45e7-9426-e61b90c5d129 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a63b09-6cd1-4e1e-b70f-df790981cee3 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Q-align: Teaching lmms for visual scoring via discrete text-defined levels
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eac17490-2252-4540-9922-a54025055502 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 309f6589-57dd-473e-895c-105382268df6 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Imagere- ward: Learning and evaluating human preferences for text- to-image generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eada856-94fd-4d4a-b983-3d8988269077 · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching What you see is what you read? improving text- image alignment evaluation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6d3eace7-8ac9-421e-8c84-af0da819edef · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching mplug- owl3: Towards long image-sequence understanding in multi- modal large language models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation de44943b-669d-4577-a7f3-33a215fc197d · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching Swift:a scal- able lightweight infrastructure for fine-tuning, 2024
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83488eb1-ff86-4677-a0fe-16a21ea732ff · outbound
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.