Pith. sign in

Paper Citation Record · LEDGER

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers

As of 21 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 2 inbound Pith citation observations for arXiv:2412.12571.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12571 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:01:05.400479Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:17:26.477465Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T22:32:52.410596Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a84ee96-abcd-47e3-82d8-7d19b7780ab5 · outbound

This paper cites Group Diffusion Transformers are Unsupervised Multitask Learners.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers Group Diffusion Transformers are Unsupervised Multitask Learners

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.238666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.238666Z digest=sha256:443b1dd9ebab134b0549da50c33c37aea62aa7af07d686523abf289c425eb272

Observation 5c726b17-9b80-44c8-a6fa-22f00b526035 · outbound

This paper cites an unresolved cited work.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.258897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.258897Z digest=sha256:e7c521dbeaea91c0ecc1b8f40e719e05de76d2dadcc527b49cecb571b3eb31e9

Observation fe374278-9321-462b-ab42-dd0632a37744 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.271030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.271030Z digest=sha256:eb13914144a80382d5c39e7416d51ad1a91d864dd664d91042c82b11d1e5a538

Observation 854ad155-1981-40ea-88ae-7326cc74a9e3 · outbound

This paper cites Composer: Creative and Controllable Image Synthesis with Composable Conditions.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers Composer: Creative and Controllable Image Synthesis with Composable Conditions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.277567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.277567Z digest=sha256:2de797a10736003ea9d8839a53317e8a12613b9dcdb483b3690ebf548f24196b

Observation 814f50cb-5e1b-4e0f-a232-6f155149776f · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.285683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.285683Z digest=sha256:e8bfa66a0c1954e542936383792fb355e49155a5fc608ecb50ff3107f23b07bc

Observation c8b0ed51-e531-441f-914c-5fd9f51aba65 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers Making LLaMA SEE and Draw with SEED Tokenizer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.296562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.296562Z digest=sha256:78c1f90198ff46d75177813d7970e1754ca1ef71d89eb189be356cd3ac5f2097

Observation fa193d87-93b3-4823-a3f7-c8015edde9b0 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers Emu3: Next-Token Prediction is All You Need

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.304220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.304220Z digest=sha256:8441b3f21ff8d27a7122dd2c93eac7f6cdc35400d8a6fc4bf872f3b32fb35851

Observation 9aa03976-dae7-4c8c-88e4-7080a193fee9 · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.318473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.318473Z digest=sha256:c5c9216d75ab4987f50a51d93b513e0568bd9ff763f8c85d2a17ec164ec99064

Observation f6accdaa-c2f7-40cd-a9b8-12aa303c8878 · outbound

This paper cites Drag your gan: Interactive point-based manipulation on the generative image manifold.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers Drag your gan: Interactive point-based manipulation on the generative image manifold

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:01:06.186802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:01:05.324294Z digest=sha256:82ae6913ab4b9dfe2c4cc5efd08bf2b43711b8a7a16f4fefdc4cf32e0b71f8fa

Observation 5aa6fa44-5b85-49c5-9a72-0689a37f0a28 · outbound

This paper cites DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image Editing.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image Editing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.330008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.330008Z digest=sha256:3b15c30a2d2f0ca635b550713c3dd33a7d5019799cd298e6260661fdd03fcb22

Observation 53ba41ce-4067-407f-86c3-b0c9aa164a0d · outbound

This paper cites DiffIR: Efficient Diffusion Model for Image Restoration.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers DiffIR: Efficient Diffusion Model for Image Restoration

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.336272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.336272Z digest=sha256:d5035b5ff63d431db349b24811cffbce98daa2a67efb6183568abffbb0e9730a

Observation 946cf26f-ae14-4293-aede-56a306240e6d · outbound

This paper cites StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.345697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.345697Z digest=sha256:c29804ef044e6cff830354bc4c67ab6c7fa2112dcf04acbd8fb8f05b61360010

Observation fb97c29b-927d-4ba9-9dac-fe8f6739f512 · outbound

This paper cites SeedEdit: Align Image Re-Generation to Image Editing.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers SeedEdit: Align Image Re-Generation to Image Editing

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.352829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.352829Z digest=sha256:d1d85385d85d2d333c4c680a0e910dd97568a4ee9fd3f2a0d2336f59aae705de

Observation 722ab829-29db-4689-b7a8-5a95a491cd71 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.360714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.360714Z digest=sha256:4a7a4a5b8a219694420dc00954e42bdb71ccef36bd390234df24bfc47916bbd0

Observation 0e8a87ec-e2f1-44e0-a10a-3edd261c34bb · outbound

This paper cites OmniGen: Unified Image Generation.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers OmniGen: Unified Image Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.366837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.366837Z digest=sha256:6869e3d5a22189cc995397b6e2225270497a25538ee332b47c0db37db9f05ad9

Observation f01c8d40-8e49-4f17-8936-c1a9e66539f9 · outbound

This paper cites Agent AI: Surveying the Horizons of Multimodal Interaction.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers Agent AI: Surveying the Horizons of Multimodal Interaction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.389630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.389630Z digest=sha256:011d82f13dc022f6ab4e38a0e67e0754bda74ec128a50687db8c47b87c3457b7

Observation 1de2283a-021d-4d3e-84ee-d413f1482a7a · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.400479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.400479Z digest=sha256:86b686b65b3be7293f2bc45f20df7b9177330709d48c39e11fdf1bbdfc168a10

Observation 2a018648-578b-442a-a609-dd0775163792 · outbound

This paper cites Language Models are Few-Shot Learners.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers Language Models are Few-Shot Learners

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.374941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.374941Z digest=sha256:00c81499c50af48ee6bf313bf943d1dcda31052a740b16006683cbc38df441f5

Observation f1917e7b-4642-4b4f-a316-505bca95604e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers LLaMA: Open and Efficient Foundation Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.383221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.383221Z digest=sha256:2185e9592f92f50777c2829da7bb0d007329c6be67d93ac480fc1eba1965b091

Observation b6bea0df-6abc-4bdd-a53f-573f0182d51f · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.246653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.246653Z digest=sha256:721513365f36e416dfcb4f24bbdd228d987bd37d47a7c4234d1c1e28a194538b

Observation a658140b-e0cf-479e-9c97-00122ea9fd1c · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.311019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.311019Z digest=sha256:6db99c04047f3dbeb54f9425fb6e1248acba8557690963d0c40255ab9bcdd413

Observation 8d695622-c5d9-4693-b941-2a5044bfef75 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.252633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.252633Z digest=sha256:e1ae8c34d7e56369479e1fca0acee0580f712e318fa0f714a205095bf87e5f58

Observation 8c9bf03c-38eb-4917-bf3c-55ef18225539 · outbound

This paper cites Lvmin Zhang, Anyi Rao, and Maneesh Agrawala.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers Lvmin Zhang, Anyi Rao, and Maneesh Agrawala

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:01:06.208403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:01:05.264832Z digest=sha256:c783a78fb43588af40451e0dd66765930f0059510d73f2fe2b87c6003aca8187

Pith citing papers

Observation 924ccdb7-a76b-4da7-ac5a-19d3b4368e2d · inbound

IA-T2I: Internet-Augmented Text-to-Image Generation cites this paper.

IA-T2I: Internet-Augmented Text-to-Image Generation ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:26.477465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:26.477465Z digest=sha256:193c7e296f24293d8991309d1b0e8681cd328a9685f54dd497a93c5894db9e2f

Observation ce8287be-baa5-48b5-b1f6-dfe390199b9c · inbound

MultiRef: Controllable Image Generation with Multiple Visual References cites this paper.

MultiRef: Controllable Image Generation with Multiple Visual References ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:32:52.416568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T22:32:51.845176Z digest=sha256:adfd162bcf652b3ad1416758e7c8310d547c47f6b088d17d4832d2b36ae6298d