Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models

As of 13 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2501.00917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00917 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:42:47.923687Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d5dcc33-ebaa-4faa-a274-ae49644a8e58 · outbound

This paper cites Learning transferable visual models from na tural language supervision,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Learning transferable visual models from na tural language supervision,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.800949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.800949Z digest=sha256:6f46930bfb6f56123bc0448db19226c4bb7387f8779267a6eb219908cbf070ce

Observation 3e59bc6d-8dad-4484-b063-ceca84274ef6 · outbound

This paper cites Fla mingo: a visual language model for few-shot learning,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Fla mingo: a visual language model for few-shot learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.429136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.806179Z digest=sha256:20cf0cb91d50cae4d2dffa20bda41eaca3db0e1a3818673d6019b459277cf275

Observation 6f6057e7-6a52-4df8-ae77-6a10ae1c2804 · outbound

This paper cites Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.810775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.810775Z digest=sha256:f771832249cfbc144a14265658c468c07ecca440385b9749ce2d687cf6021105

Observation c70eaa7e-504a-4f3f-8c2a-a913f2faba0b · outbound

This paper cites ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:42:48.232622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.816050Z digest=sha256:403df905c20889a15f7880b9826252ce04bf551c52c8d9ac95ebad7a932b7691

Observation 839626a2-47b2-4584-8e5f-aee86d62c9c1 · outbound

This paper cites Textd iffuser: Diffusion models as text painters,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Textd iffuser: Diffusion models as text painters,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.413065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.820974Z digest=sha256:1cbf49a0aa70d0996b6378073b7e77a3d0c26920cd44c26686ec5c4f829008b1

Observation 2842a8d5-f723-4814-a821-3dd074fbe122 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.825552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.825552Z digest=sha256:82e9eb129a305e715f8bf1d1d8122d11408379e9ac8eca7039299a0d5aec17b4

Observation 7d478ed2-71ab-4f85-8e57-07ab9d33de9c · outbound

This paper cites STAR: Scale-wise Text-conditioned AutoRegressive image generation.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models STAR: Scale-wise Text-conditioned AutoRegressive image generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.830666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.830666Z digest=sha256:f56411995ff9c273ea2ed1b8eb4520a452deae3177a3412f5c928aacf2f82aa2

Observation 3aad78ac-da31-432d-9ade-f8e89ecaf145 · outbound

This paper cites Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.835262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.835262Z digest=sha256:828ad6503704b4edf642eeebe528aeb107bf09ad4370a740ccceb61ed3b658aa

Observation f9de7c66-9b72-4071-95e3-3028d3157f0b · outbound

This paper cites An analysis of the ingredients for learning interpretable symbolic regression models with human-in-the-loop and genetic prog ramming,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models An analysis of the ingredients for learning interpretable symbolic regression models with human-in-the-loop and genetic prog ramming,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.397720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.839951Z digest=sha256:ed2536d02c976260206f6cc10cf3b9525e059bbfc32408ae6429605376d31c3f

Observation d0ff0900-95db-490d-bbb6-8994ba843467 · outbound

This paper cites Improving cross-modal alignment f or text- guided image inpainting,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Improving cross-modal alignment f or text- guided image inpainting,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.843949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.843949Z digest=sha256:0a1bc48b76d2fe7e1eeef582ccfa743eaf160d4e914f6d0fc77ba4d628882582

Observation ee5d1221-7ff3-4ab9-939b-13f0fb837801 · outbound

This paper cites Zero-shot text-to-image generation,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Zero-shot text-to-image generation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.373659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.847927Z digest=sha256:60d1e2d775c2909c5922b897c354b965ae30c37482399c74d5da2d37e98a9a1f

Observation cd633819-6bf3-46c3-a9c4-4a84c31ee2ed · outbound

This paper cites Towards language-driven video inpainting via multimoda l large language models,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Towards language-driven video inpainting via multimoda l large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.359278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.852061Z digest=sha256:5b0d4543d4eeadfeaea555264442408b0c815c58d3939f77c2d344b3af6e2c83

Observation 7df5d13b-9bd0-4922-82e1-58ada63d4361 · outbound

This paper cites Prompt Expansion for Adaptive Text-to-Image Generation.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Prompt Expansion for Adaptive Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.856445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.856445Z digest=sha256:64053d7b649bdcf328603997412086a193537debb1ef4804ee157fbaf33c4557

Observation 5bd31eed-2ce6-4aca-b6db-8362b87569cd · outbound

This paper cites Training-free consistent text-to-image gener ation,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Training-free consistent text-to-image gener ation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.344419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.861291Z digest=sha256:ba59a98ef1c6fa426a735da5dc9e54b568b3419ab43222ff9191cf6e8ee46ce5

Observation b92ebe12-0373-493b-baff-f96f1d5f8d6a · outbound

This paper cites Style-aware contrastive learning for multi-style image captioning,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Style-aware contrastive learning for multi-style image captioning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.865685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.865685Z digest=sha256:c58f5f4e2035df170100ec5f122ddc10b761d5b3cb4129c25da92dfedf3d7ae5

Observation eedea3d6-b472-4ad6-86aa-c758ef8bb01f · outbound

This paper cites Multimodal event transformer for image-guided st ory ending generation,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Multimodal event transformer for image-guided st ory ending generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.320620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.869872Z digest=sha256:dd116173163edcdbdf2870536a3712270767c27d0f03298031f04f4c99e6137b

Observation 4f700a97-5b4f-44d2-919d-8013fd81e193 · outbound

This paper cites Triple sequence generati ve adversarial nets for unsupervised image captioning,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Triple sequence generati ve adversarial nets for unsupervised image captioning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.874406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.874406Z digest=sha256:771c7ea282661b51a79aeea2646273361ba18dc2b2917fd9083d00fb376c34a1

Observation ccda3c91-5b5b-4ba9-b9f4-0efd3f4651b2 · outbound

This paper cites Sketch storytelling,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Sketch storytelling,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.878734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.878734Z digest=sha256:819602aafdb46d8388daf0aa93f35dc6f9c4cf0601b4f36f076a6d963b0aea9e

Observation 70acb0cc-a645-4cc5-9b12-9da7a0068e37 · outbound

This paper cites Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:42:48.155107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.882788Z digest=sha256:857e3eacfe43fad4ed0cc6fe85e34b53fc093117acb68839d28ec4a9ef715c39

Observation 84865179-7f60-4189-a0b0-b616a27b54b6 · outbound

This paper cites Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.887679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.887679Z digest=sha256:8a9d1b2a06c4c1eac9ced97e2a71d448f366a4238b6c08eda4e874f34ce44d77

Observation 4647b3ae-dcd7-4ef3-89b5-b507a7c7b778 · outbound

This paper cites Visual in-context l earning for large vision-language models,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Visual in-context l earning for large vision-language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.892120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.892120Z digest=sha256:ee161fb363be0c6f9dc1df5feff71896b9844fb0abc13c4ece25f551688b1f63

Observation 050b7709-0670-4258-b108-1e6a93eb91af · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.896528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.896528Z digest=sha256:d81f219e19602e72659db27673cb4e52d4f721883080c4158b69634cbe143f05

Observation 31b6ca13-b271-4ac6-8df8-52467eeb0a85 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.901252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.901252Z digest=sha256:043d1741f116ed51e648f0f9c5de22db902c2b4d6854827d4cdf307a7a6e2097

Observation 6672227c-d77f-4391-9712-fa755c0eb995 · outbound

This paper cites Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.905883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.905883Z digest=sha256:5c8d8da09955960655a97707417249139b4a25033c1714a040c5b8a797b96714

Observation 14e2375b-3758-4254-9be4-f3b03b6a1106 · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.910468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.910468Z digest=sha256:f255795456d5a5d76361b24e73dc091b9456300fc7ea941adf86ec1b38201a37

Observation 4261da68-09ad-43cd-84e4-f7848eafff97 · outbound

This paper cites Ex ploring the frontier of vision-language models: A survey of current met hodologies and future directions,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Ex ploring the frontier of vision-language models: A survey of current met hodologies and future directions,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.914817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.914817Z digest=sha256:f71e04801ba829d9e7a13c2d99649d60d8e9707ba5d2ef1be9e7af3f0393fdc9

Observation 4bc47a9f-7762-4e47-aefd-26d990754358 · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.276806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.918957Z digest=sha256:fe50abc7b811d46f9c1329815ada79b656b6a6ee5a9b2e682917931792330d50

Observation b2e9a579-3d1b-43c4-9baf-103d1b168642 · outbound

This paper cites Less is more: Vision representation compression for efficient video gene ration with large language models,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Less is more: Vision representation compression for efficient video gene ration with large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.262399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.923687Z digest=sha256:dc22475835dda903f18f272cd7774775e3a0c31b38d87e0eacacb3994ef41b55

Pith citing papers

No inbound Pith citation observations are available.