Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:42:47.923687Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2501.00917.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:42:47.923687Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9d5dcc33-ebaa-4faa-a274-ae49644a8e58 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Learning transferable visual models from na tural language supervision,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e59bc6d-8dad-4484-b063-ceca84274ef6 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Fla mingo: a visual language model for few-shot learning,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6f6057e7-6a52-4df8-ae77-6a10ae1c2804 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c70eaa7e-504a-4f3f-8c2a-a913f2faba0b · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 839626a2-47b2-4584-8e5f-aee86d62c9c1 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Textd iffuser: Diffusion models as text painters,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2842a8d5-f723-4814-a821-3dd074fbe122 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Muse: Text-To-Image Generation via Masked Generative Transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d478ed2-71ab-4f85-8e57-07ab9d33de9c · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models STAR: Scale-wise Text-conditioned AutoRegressive image generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aad78ac-da31-432d-9ade-f8e89ecaf145 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9de7c66-9b72-4071-95e3-3028d3157f0b · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models An analysis of the ingredients for learning interpretable symbolic regression models with human-in-the-loop and genetic prog ramming,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d0ff0900-95db-490d-bbb6-8994ba843467 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Improving cross-modal alignment f or text- guided image inpainting,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee5d1221-7ff3-4ab9-939b-13f0fb837801 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Zero-shot text-to-image generation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cd633819-6bf3-46c3-a9c4-4a84c31ee2ed · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Towards language-driven video inpainting via multimoda l large language models,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7df5d13b-9bd0-4922-82e1-58ada63d4361 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Prompt Expansion for Adaptive Text-to-Image Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bd31eed-2ce6-4aca-b6db-8362b87569cd · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Training-free consistent text-to-image gener ation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b92ebe12-0373-493b-baff-f96f1d5f8d6a · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Style-aware contrastive learning for multi-style image captioning,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eedea3d6-b472-4ad6-86aa-c758ef8bb01f · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Multimodal event transformer for image-guided st ory ending generation,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4f700a97-5b4f-44d2-919d-8013fd81e193 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Triple sequence generati ve adversarial nets for unsupervised image captioning,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccda3c91-5b5b-4ba9-b9f4-0efd3f4651b2 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Sketch storytelling,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70acb0cc-a645-4cc5-9b12-9da7a0068e37 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 84865179-7f60-4189-a0b0-b616a27b54b6 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4647b3ae-dcd7-4ef3-89b5-b507a7c7b778 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Visual in-context l earning for large vision-language models,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 050b7709-0670-4258-b108-1e6a93eb91af · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b6ca13-b271-4ac6-8df8-52467eeb0a85 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6672227c-d77f-4391-9712-fa755c0eb995 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14e2375b-3758-4254-9be4-f3b03b6a1106 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models A Survey on Benchmarks of Multimodal Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4261da68-09ad-43cd-84e4-f7848eafff97 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Ex ploring the frontier of vision-language models: A survey of current met hodologies and future directions,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bc47a9f-7762-4e47-aefd-26d990754358 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b2e9a579-3d1b-43c4-9baf-103d1b168642 · outbound
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Less is more: Vision representation compression for efficient video gene ration with large language models,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.