Pith. sign in

Paper Citation Record · LEDGER

Taming Transformers for High-Resolution Image Synthesis

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2012.09841.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2012.09841 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:07:11.354540Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c9e1fa55-7d71-401f-b199-a0f536a99d2c · inbound

High-Resolution Image Synthesis with Latent Diffusion Models cites this paper.

High-Resolution Image Synthesis with Latent Diffusion Models Taming Transformers for High-Resolution Image Synthesis

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:02:10.049564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T22:02:09.899347Z digest=sha256:8901e7dd454cf1e6e97b6d8c72470aa7d73f381dd2d257dfbc66c43672054850

Observation 8688c964-81cb-44ee-bc32-acf0cc50e4bf · inbound

Hierarchical Text-Conditional Image Generation with CLIP Latents cites this paper.

Hierarchical Text-Conditional Image Generation with CLIP Latents Taming Transformers for High-Resolution Image Synthesis

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:55:57.706155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:55:57.612364Z digest=sha256:6ebbfbad2a2101c350150baa3d5672d15fc2e2c2f716c25b331212199a5aa98e

Observation 627d82d3-d55d-4f65-beb1-a76d206ff8b7 · inbound

CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers cites this paper.

CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers Taming Transformers for High-Resolution Image Synthesis

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:24:30.169637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T12:24:30.071822Z digest=sha256:4a76adab78b85cbf147f8a58815c26d8171f7335b44727ea42e05d3389a5f2f7

Observation d992ae0f-c056-43e8-8e4b-a13887ea7fe6 · inbound

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets cites this paper.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Taming Transformers for High-Resolution Image Synthesis

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:58:52.084829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:4d3a1729edeb42bad49bdb33e4c4be7f2d9e324679942aaeb19c5ab419c34bc8

Observation e8f7e586-273a-4e01-97e8-6a4102a7f999 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Taming Transformers for High-Resolution Image Synthesis

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:56:07.899835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:b7ecdd6b9952188cee3fbca03881a7a1e9aa8839825d606f22a1ac07214e72c1

Observation 72e589fa-862b-4016-b3a5-4f6c6c8f6cea · inbound

Exploring the Effectiveness of Deep Features from Domain-Specific Foundation Models in Retinal Image Synthesis cites this paper.

Exploring the Effectiveness of Deep Features from Domain-Specific Foundation Models in Retinal Image Synthesis Taming Transformers for High-Resolution Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:11.354540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:11.354540Z digest=sha256:6223b5572cbb8bdb4a80bddb595f0d9d83dae5669dffa8e85a4f6a0079e03a8f

Observation 39f2225e-a259-4729-818e-3aae54b1f9ce · inbound

Transition Matching: Scalable and Flexible Generative Modeling cites this paper.

Transition Matching: Scalable and Flexible Generative Modeling Taming Transformers for High-Resolution Image Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:28.121890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:46:28.121890Z digest=sha256:8a7098393c18639033244c34172933aa39becc47f215038acae6ccde8ace03cb

Observation fda2a2e6-294f-44e9-9362-250e07a6b7ec · inbound

Latent Sensor Fusion: Multimedia Learning of Physiological Signals for Resource-Constrained Devices cites this paper.

Latent Sensor Fusion: Multimedia Learning of Physiological Signals for Resource-Constrained Devices Taming Transformers for High-Resolution Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:36.667446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:36.667446Z digest=sha256:821b1b3666837af45e6697c168dcf3645d731a0fbdaac567005e92c2920bf5ee

Observation 25999249-f595-49b7-89a0-68e4bd432bcb · inbound

StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization cites this paper.

StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization Taming Transformers for High-Resolution Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T10:47:00.547306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:47:00.547306Z digest=sha256:463b1161eb3242bbf3769a7e2ce9dd2025a2db91b9b0fd047bf74d39e930a691

Observation 898cb73b-271a-4ba7-830e-301c5af7812a · inbound

Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model cites this paper.

Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model Taming Transformers for High-Resolution Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:41:16.150346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:41:16.150346Z digest=sha256:ce7d7e01083e4d827a4d5cad1d6eae880164a2cd7525075d2ccf09557d0d294d

Observation 64746141-28d3-40d5-b6fa-56f1b65a1650 · inbound

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission cites this paper.

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission Taming Transformers for High-Resolution Image Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:29:03.656854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:29:03.656854Z digest=sha256:7070061b1d502f2d95aee71c72b4c92628dd1d87c35be27f2ca1b88bc08f8e60

Observation a5b1424c-630d-4465-8998-9f1c458fb51b · inbound

InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames cites this paper.

InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames Taming Transformers for High-Resolution Image Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:12.652710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:12.652710Z digest=sha256:c9d8ea0e5e6e5611012849dee1cacb3aaaebd441d7bc137927cf0f0bbdcf3b05

Observation 5b12fbaa-9ad2-42d8-b177-119cdc997842 · inbound

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion cites this paper.

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion Taming Transformers for High-Resolution Image Synthesis

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:05:57.417886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:57:24.033068Z digest=sha256:89571f4d9be6703f9dfa634fc88f77b608aca4c00a032090071ba37e2a87c899

Observation 734f5076-98ec-43e1-8c66-86eb436b1fa8 · inbound

CaloArt: Large-Patch x-Prediction Diffusion Transformers for High-Granularity Calorimeter Shower Generation cites this paper.

CaloArt: Large-Patch x-Prediction Diffusion Transformers for High-Granularity Calorimeter Shower Generation Taming Transformers for High-Resolution Image Synthesis

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:42:15.836286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T04:40:25.725622Z digest=sha256:06bc26af5fe484914fe68a6b73b01e70f8d671e34dc7499bf076c36b1b8745c9

Observation 03843ae1-6ea9-475b-a954-43b6d8fe80ca · inbound

Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation cites this paper.

Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation Taming Transformers for High-Resolution Image Synthesis

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:34:38.743266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T12:28:33.099835Z digest=sha256:ed571a20fa8966ab589308192791d13042f60cc219d3888b6bff432090e31dfe

Observation aa44a522-961e-4bbc-942b-d6f2520129dc · inbound

GPIC: A Giant Permissive Image Corpus for Visual Generation cites this paper.

GPIC: A Giant Permissive Image Corpus for Visual Generation Taming Transformers for High-Resolution Image Synthesis

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:43:14.161246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T07:36:21.262064Z digest=sha256:ea1619ce51230aa1c5ab3c0d0e4ce52612d2206b42702142121d3a783e2e8a07

Observation 1dc90202-abb1-4242-bb93-fc2e14f65420 · inbound

Consistency Training while Mitigating Obfuscation via Rate Matching cites this paper.

Consistency Training while Mitigating Obfuscation via Rate Matching Taming Transformers for High-Resolution Image Synthesis

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-06-28T14:32:18.108266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:25:43.147442Z digest=sha256:d59b29ea331f3452d64b31050d433a21304d0e92f5fca9ea31fd4e065b292645

Observation e0631fdc-ce03-4906-ba42-240e4f13806d · inbound

Mapping Whisper Representations to Human ECoG Responses with Interpretable Time-Resolved Neural Encoding cites this paper.

Mapping Whisper Representations to Human ECoG Responses with Interpretable Time-Resolved Neural Encoding Taming Transformers for High-Resolution Image Synthesis

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T11:42:04.140355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:39:30.563754Z digest=sha256:85e0fff8953d6fb3f4f738a6c0681b3b4af3f5c2c205583a7eb230a49f6c77f2

Observation 7aa61616-3ae2-417a-b70b-fa8eb1c68ca9 · inbound

Variational Proximal Policy Optimization cites this paper.

Variational Proximal Policy Optimization Taming Transformers for High-Resolution Image Synthesis

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:57:26.059020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T19:22:23.768249Z digest=sha256:1eae1fa600a9270766cacaf9d33c9978dc147592ba9b3d03deeef7be7f995bd5

Observation 542bbc52-ab6e-49ae-9736-e90bc4695b2c · inbound

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization cites this paper.

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization Taming Transformers for High-Resolution Image Synthesis

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:27:39.708579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:18:47.472178Z digest=sha256:dae69abbc076459d26fb117b7eb023b558f5495c7a37b1932343a353fec4e655

Observation 70464072-1d41-4a12-98c3-acceea611229 · inbound

The Market in the Model: Latent Diffusion as Neural Economy cites this paper.

The Market in the Model: Latent Diffusion as Neural Economy Taming Transformers for High-Resolution Image Synthesis

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:49:24.758638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:05:25.386543Z digest=sha256:83bf16341271950fc55178b3cb217b3d2b1db04033ee160f747d83714afa1d89

Observation 9f2441b9-0330-4bcb-b0a0-5d1c596fbad7 · inbound

Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance cites this paper.

Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Taming Transformers for High-Resolution Image Synthesis

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.090513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T21:03:43.962738Z digest=sha256:2f852c2c01cc16d2725d412fb7ac62374009b16a281f08886911d5f693c3cfb7

Observation d0707cb4-3fa8-4053-8994-3dca0f5a748d · inbound

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing cites this paper.

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Taming Transformers for High-Resolution Image Synthesis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T13:39:00.040780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:39:00.040780Z digest=sha256:300f63103cd23d2e47de05f2503c3d85e0579d449710d2f6ede104cb760771ac

Observation 01433ba0-a80d-4a11-b89c-51a8c8fefd1a · inbound

FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents cites this paper.

FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents Taming Transformers for High-Resolution Image Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:50:35.363661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:50:35.363661Z digest=sha256:dcb7749206226fe585b1fa54910b75ccc93eded17834716106b1d71260e17e4f