Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models

As of 14 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2501.00917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00917 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:42:47.923687Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d5dcc33-ebaa-4faa-a274-ae49644a8e58 · outbound

This paper cites Learning transferable visual models from na tural language supervision,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Learning transferable visual models from na tural language supervision,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.800949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.800949Z digest=sha256:125104ce8d31bebe91feb14a60519db104b8ffcae3dc2ca6806e6806de2f7ddb

Observation 3e59bc6d-8dad-4484-b063-ceca84274ef6 · outbound

This paper cites Fla mingo: a visual language model for few-shot learning,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Fla mingo: a visual language model for few-shot learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.429136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.806179Z digest=sha256:d6c9c19d997b0ffc48a1337d83901eac720ac6d8c6c9ffc46e85c855ec3a0709

Observation 6f6057e7-6a52-4df8-ae77-6a10ae1c2804 · outbound

This paper cites Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.810775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.810775Z digest=sha256:fa7831380bbf436094a58ed10a02821afa2bd22ee44cf9a740b06b9c23fee222

Observation c70eaa7e-504a-4f3f-8c2a-a913f2faba0b · outbound

This paper cites ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:42:48.232622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.816050Z digest=sha256:4c155a1c7622dbc1decf85b6e3dd0270a873f8aab0f86cf51dd9b0b38f19dcba

Observation 839626a2-47b2-4584-8e5f-aee86d62c9c1 · outbound

This paper cites Textd iffuser: Diffusion models as text painters,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Textd iffuser: Diffusion models as text painters,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.413065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.820974Z digest=sha256:3f98f9437b2d17abf17621a0eae72c2dd7358da4a8c1688fce6ff11207c2fdb5

Observation 2842a8d5-f723-4814-a821-3dd074fbe122 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.825552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.825552Z digest=sha256:dbefbb9000c15736fed5c22070b768b3f6a6d6dbd17e38c3a580320a9463c5cf

Observation 7d478ed2-71ab-4f85-8e57-07ab9d33de9c · outbound

This paper cites STAR: Scale-wise Text-conditioned AutoRegressive image generation.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models STAR: Scale-wise Text-conditioned AutoRegressive image generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.830666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.830666Z digest=sha256:a7c1bca7acb3ec849c79d7e11629a51c50df9b6c0f43c9e9426bbd3937d553c5

Observation 3aad78ac-da31-432d-9ade-f8e89ecaf145 · outbound

This paper cites Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.835262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.835262Z digest=sha256:5695f4afa4b6ff51e68dc986c740175f8e7e51cf2ba2f1532f805b3a65cb2e64

Observation f9de7c66-9b72-4071-95e3-3028d3157f0b · outbound

This paper cites An analysis of the ingredients for learning interpretable symbolic regression models with human-in-the-loop and genetic prog ramming,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models An analysis of the ingredients for learning interpretable symbolic regression models with human-in-the-loop and genetic prog ramming,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.397720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.839951Z digest=sha256:c1fb52f905163bd7a30e4bf2af89f2a603b4ed6dd4cb8c9b989a0faf1c62db66

Observation d0ff0900-95db-490d-bbb6-8994ba843467 · outbound

This paper cites Improving cross-modal alignment f or text- guided image inpainting,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Improving cross-modal alignment f or text- guided image inpainting,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.843949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.843949Z digest=sha256:5e87d18cfc59f14ea4c4bac006d6ffcd0d368e14d31ff2d1ac55c68fd9127172

Observation ee5d1221-7ff3-4ab9-939b-13f0fb837801 · outbound

This paper cites Zero-shot text-to-image generation,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Zero-shot text-to-image generation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.373659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.847927Z digest=sha256:863e943895f686d7d0a4460fc9d2f10f1550b8b686c3125a0f7f2274a85c0e53

Observation cd633819-6bf3-46c3-a9c4-4a84c31ee2ed · outbound

This paper cites Towards language-driven video inpainting via multimoda l large language models,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Towards language-driven video inpainting via multimoda l large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.359278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.852061Z digest=sha256:b71d24a80e090609321f04813a46a999914c032a50195a1a0276a3a9f2d581e4

Observation 7df5d13b-9bd0-4922-82e1-58ada63d4361 · outbound

This paper cites Prompt Expansion for Adaptive Text-to-Image Generation.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Prompt Expansion for Adaptive Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.856445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.856445Z digest=sha256:0d19a3e32fc90b7e6d816d3e8da6974cbdc6c81f36128cc4fdc15a749a017d5b

Observation 5bd31eed-2ce6-4aca-b6db-8362b87569cd · outbound

This paper cites Training-free consistent text-to-image gener ation,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Training-free consistent text-to-image gener ation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.344419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.861291Z digest=sha256:5da1b7875bc8725a4d8edf06ced7b96a1172f2fd47d0c9d32a60b1a9f6f831e6

Observation b92ebe12-0373-493b-baff-f96f1d5f8d6a · outbound

This paper cites Style-aware contrastive learning for multi-style image captioning,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Style-aware contrastive learning for multi-style image captioning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.865685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.865685Z digest=sha256:c0c78606d4f643a4bbd09564e660e175c403369f6b739e372436affe4fb7d3af

Observation eedea3d6-b472-4ad6-86aa-c758ef8bb01f · outbound

This paper cites Multimodal event transformer for image-guided st ory ending generation,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Multimodal event transformer for image-guided st ory ending generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.320620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.869872Z digest=sha256:2e62c4cc459f0af9a821368e88282807f21cacfb5aeb5cdce2ccdb83b1b05dc2

Observation 4f700a97-5b4f-44d2-919d-8013fd81e193 · outbound

This paper cites Triple sequence generati ve adversarial nets for unsupervised image captioning,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Triple sequence generati ve adversarial nets for unsupervised image captioning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.874406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.874406Z digest=sha256:54a6b1adc41f1af51d8092c6cc6e72960fc4a6fc1c12c508c09726ebfaa8be4a

Observation ccda3c91-5b5b-4ba9-b9f4-0efd3f4651b2 · outbound

This paper cites Sketch storytelling,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Sketch storytelling,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.878734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.878734Z digest=sha256:e1f2e9e6becdd08156c9dbd66a68e51f1ee6329ecb72821c18c04deb2d36c79e

Observation 70acb0cc-a645-4cc5-9b12-9da7a0068e37 · outbound

This paper cites Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Generalist Multimodal AI: A Review of Architectures, Challenges and Opportunities

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:42:48.155107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.882788Z digest=sha256:7f3da994462839118c266de28619a1899eaad382af33096ecd3100c44aa471b1

Observation 84865179-7f60-4189-a0b0-b616a27b54b6 · outbound

This paper cites Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.887679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.887679Z digest=sha256:2a59736be6cc38183564122eb4e33fe7e78eb1828dfabbaa8c16567a4eb53a6c

Observation 4647b3ae-dcd7-4ef3-89b5-b507a7c7b778 · outbound

This paper cites Visual in-context l earning for large vision-language models,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Visual in-context l earning for large vision-language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.892120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.892120Z digest=sha256:8fa2cfa6419385466fa011e95fb84fe58935ffe9d407b86ed5a61fe0be782557

Observation 050b7709-0670-4258-b108-1e6a93eb91af · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.896528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.896528Z digest=sha256:150d669172ebb40dca1d62da9afc4581527c02c69539f0756d9077c6f70b3f14

Observation 31b6ca13-b271-4ac6-8df8-52467eeb0a85 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.901252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.901252Z digest=sha256:791d3852d174b4d092b2e04b2668fcfafd79fa8170502bfdf53f858e854696d9

Observation 6672227c-d77f-4391-9712-fa755c0eb995 · outbound

This paper cites Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.905883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.905883Z digest=sha256:28c6b1cf812127f91338d210968ffd8f48fd79106a4095dcbfff5ae2532fff78

Observation 14e2375b-3758-4254-9be4-f3b03b6a1106 · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.910468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.910468Z digest=sha256:c5ae57bc67c59c903e4683b62356c5c94f8ecba860bf0fa1f900a67071c373dc

Observation 4261da68-09ad-43cd-84e4-f7848eafff97 · outbound

This paper cites Ex ploring the frontier of vision-language models: A survey of current met hodologies and future directions,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Ex ploring the frontier of vision-language models: A survey of current met hodologies and future directions,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:42:47.914817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:42:47.914817Z digest=sha256:e58f6f008520d7daf1b92af56b8c2ea7e13597602ff5df70f0a08378f42cc6a8

Observation 4bc47a9f-7762-4e47-aefd-26d990754358 · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.276806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.918957Z digest=sha256:e0a3e43ba2165aa376de436a4fa48194d7291a6b410c5d68ee961557a80055a8

Observation b2e9a579-3d1b-43c4-9baf-103d1b168642 · outbound

This paper cites Less is more: Vision representation compression for efficient video gene ration with large language models,.

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models Less is more: Vision representation compression for efficient video gene ration with large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:42:48.262399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:42:47.923687Z digest=sha256:5a506b2b3311b9e1052456b6b71f3fa3dea39897244de0aac9227eb6df28633d

Pith citing papers

No inbound Pith citation observations are available.