Pith. sign in

Paper Citation Record · LEDGER

Taming Transformers for High-Resolution Image Synthesis

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2012.09841.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2012.09841 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:46:28.121890Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c9e1fa55-7d71-401f-b199-a0f536a99d2c · inbound

High-Resolution Image Synthesis with Latent Diffusion Models cites this paper.

High-Resolution Image Synthesis with Latent Diffusion Models Taming Transformers for High-Resolution Image Synthesis

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:02:10.049564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T22:02:09.899347Z digest=sha256:67fecd097c3e68b53ccd4e3a4c111c7c944d278e77e86ef0fdb4bfff739f1b2a

Observation 8688c964-81cb-44ee-bc32-acf0cc50e4bf · inbound

Hierarchical Text-Conditional Image Generation with CLIP Latents cites this paper.

Hierarchical Text-Conditional Image Generation with CLIP Latents Taming Transformers for High-Resolution Image Synthesis

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:55:57.706155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:55:57.612364Z digest=sha256:f6ef7e3efec3df32610595895016dc338c6a591c688ea2ea3581cc76723191e4

Observation 627d82d3-d55d-4f65-beb1-a76d206ff8b7 · inbound

CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers cites this paper.

CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers Taming Transformers for High-Resolution Image Synthesis

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:24:30.169637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T12:24:30.071822Z digest=sha256:5e22200e96af934bb71a391af99a737cbca171bc7d847fe017268b39e679a3de

Observation d992ae0f-c056-43e8-8e4b-a13887ea7fe6 · inbound

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets cites this paper.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Taming Transformers for High-Resolution Image Synthesis

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:58:52.084829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:1efb1c63a0fa95e58bae163622fe284aa7de414ae79e3c4faf2d56bc29afa724

Observation e8f7e586-273a-4e01-97e8-6a4102a7f999 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Taming Transformers for High-Resolution Image Synthesis

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:56:07.899835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:f3e8365c1c7d50aefd9f68d5ee3930eccb021ba84166b02576145fa39dda36e3

Observation 39f2225e-a259-4729-818e-3aae54b1f9ce · inbound

Transition Matching: Scalable and Flexible Generative Modeling cites this paper.

Transition Matching: Scalable and Flexible Generative Modeling Taming Transformers for High-Resolution Image Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:28.121890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:46:28.121890Z digest=sha256:8a7098393c18639033244c34172933aa39becc47f215038acae6ccde8ace03cb

Observation fda2a2e6-294f-44e9-9362-250e07a6b7ec · inbound

Latent Sensor Fusion: Multimedia Learning of Physiological Signals for Resource-Constrained Devices cites this paper.

Latent Sensor Fusion: Multimedia Learning of Physiological Signals for Resource-Constrained Devices Taming Transformers for High-Resolution Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:36.667446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:36.667446Z digest=sha256:821b1b3666837af45e6697c168dcf3645d731a0fbdaac567005e92c2920bf5ee

Observation 25999249-f595-49b7-89a0-68e4bd432bcb · inbound

StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization cites this paper.

StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization Taming Transformers for High-Resolution Image Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T10:47:00.547306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:47:00.547306Z digest=sha256:463b1161eb3242bbf3769a7e2ce9dd2025a2db91b9b0fd047bf74d39e930a691

Observation 898cb73b-271a-4ba7-830e-301c5af7812a · inbound

Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model cites this paper.

Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model Taming Transformers for High-Resolution Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:41:16.150346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:41:16.150346Z digest=sha256:3664616781d6970787d2b4e5a068b8dcefa01f0f68670864a8d561c4f7cb2ab5

Observation 64746141-28d3-40d5-b6fa-56f1b65a1650 · inbound

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission cites this paper.

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission Taming Transformers for High-Resolution Image Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:29:03.656854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:29:03.656854Z digest=sha256:7070061b1d502f2d95aee71c72b4c92628dd1d87c35be27f2ca1b88bc08f8e60

Observation a5b1424c-630d-4465-8998-9f1c458fb51b · inbound

InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames cites this paper.

InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames Taming Transformers for High-Resolution Image Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:12.652710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:12.652710Z digest=sha256:c9d8ea0e5e6e5611012849dee1cacb3aaaebd441d7bc137927cf0f0bbdcf3b05

Observation 5b12fbaa-9ad2-42d8-b177-119cdc997842 · inbound

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion cites this paper.

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion Taming Transformers for High-Resolution Image Synthesis

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:05:57.417886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:57:24.033068Z digest=sha256:bb28aecea64858693828bd24fa8f23d1c63218653424a85fd1a4d4a406577fdd

Observation 734f5076-98ec-43e1-8c66-86eb436b1fa8 · inbound

CaloArt: Large-Patch x-Prediction Diffusion Transformers for High-Granularity Calorimeter Shower Generation cites this paper.

CaloArt: Large-Patch x-Prediction Diffusion Transformers for High-Granularity Calorimeter Shower Generation Taming Transformers for High-Resolution Image Synthesis

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:42:15.836286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T04:40:25.725622Z digest=sha256:27dbb0cb1189587a669e04077c0877d15b5b998e383d3a91f785d40de8e4cea7

Observation 03843ae1-6ea9-475b-a954-43b6d8fe80ca · inbound

Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation cites this paper.

Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation Taming Transformers for High-Resolution Image Synthesis

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:34:38.743266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T12:28:33.099835Z digest=sha256:ef427c8ad3eda1fb00efd6d3b738a82297358691f439ca2637a27ecec58823b4

Observation aa44a522-961e-4bbc-942b-d6f2520129dc · inbound

GPIC: A Giant Permissive Image Corpus for Visual Generation cites this paper.

GPIC: A Giant Permissive Image Corpus for Visual Generation Taming Transformers for High-Resolution Image Synthesis

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:43:14.161246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T07:36:21.262064Z digest=sha256:388da874697d2e3de862354c8bff0004595c1f36772a2399bc0f067af00f6378

Observation 1dc90202-abb1-4242-bb93-fc2e14f65420 · inbound

Consistency Training while Mitigating Obfuscation via Rate Matching cites this paper.

Consistency Training while Mitigating Obfuscation via Rate Matching Taming Transformers for High-Resolution Image Synthesis

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-06-28T14:32:18.108266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:25:43.147442Z digest=sha256:ce29f58ccbafb552228383adee10bcb6769b455a9174eb2e83373271c456a9c1

Observation e0631fdc-ce03-4906-ba42-240e4f13806d · inbound

Mapping Whisper Representations to Human ECoG Responses with Interpretable Time-Resolved Neural Encoding cites this paper.

Mapping Whisper Representations to Human ECoG Responses with Interpretable Time-Resolved Neural Encoding Taming Transformers for High-Resolution Image Synthesis

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T11:42:04.140355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T11:39:30.563754Z digest=sha256:8a4fc8d239e401075016ba22e30cb00102006ac17ceff19fa7068e24f1fc408f

Observation 7aa61616-3ae2-417a-b70b-fa8eb1c68ca9 · inbound

Variational Proximal Policy Optimization cites this paper.

Variational Proximal Policy Optimization Taming Transformers for High-Resolution Image Synthesis

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:57:26.059020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T19:22:23.768249Z digest=sha256:761a26d6ff408d9681cf7fa860929590c7444fb63213119f70710a831877d4fb

Observation 542bbc52-ab6e-49ae-9736-e90bc4695b2c · inbound

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization cites this paper.

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization Taming Transformers for High-Resolution Image Synthesis

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:27:39.708579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T13:18:47.472178Z digest=sha256:fe0d12cb04b204ac6899d0be3b6e5d381c946e22b6b5a5636c76a596c35abd50

Observation 70464072-1d41-4a12-98c3-acceea611229 · inbound

The Market in the Model: Latent Diffusion as Neural Economy cites this paper.

The Market in the Model: Latent Diffusion as Neural Economy Taming Transformers for High-Resolution Image Synthesis

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:49:24.758638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T19:05:25.386543Z digest=sha256:38b89e2898cfe95ea7c9f6c505284fe2166d05e6876e1a0f45bf987fbfcc652c

Observation 9f2441b9-0330-4bcb-b0a0-5d1c596fbad7 · inbound

Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance cites this paper.

Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Taming Transformers for High-Resolution Image Synthesis

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.090513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T21:03:43.962738Z digest=sha256:f9d976cf724d844040d31723e9a6a93105a390691fcd0d2aa26934937df51771

Observation d0707cb4-3fa8-4053-8994-3dca0f5a748d · inbound

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing cites this paper.

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Taming Transformers for High-Resolution Image Synthesis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T13:39:00.040780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:39:00.040780Z digest=sha256:300f63103cd23d2e47de05f2503c3d85e0579d449710d2f6ede104cb760771ac

Observation 01433ba0-a80d-4a11-b89c-51a8c8fefd1a · inbound

FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents cites this paper.

FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents Taming Transformers for High-Resolution Image Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:50:35.363661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:50:35.363661Z digest=sha256:dcb7749206226fe585b1fa54910b75ccc93eded17834716106b1d71260e17e4f