Pith. sign in

Paper Citation Record · LEDGER

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

As of 6 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2606.11096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.11096 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T13:10:14.308216Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact32
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch10

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ab62cca-5d5f-41ad-b3fa-81091aad638d · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:27:40.595343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:98f4dac153ed837529875dd9664b501626106e059ff4a6d977f20d8b3dc227a6

Observation 07ad680f-c33f-452d-b693-145cfcb517bc · outbound

This paper cites Perception encoder: The best visual embeddings are not at the output of the network.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Perception encoder: The best visual embeddings are not at the output of the network

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:079896f365741422ef3e12959cfcf399e512976688b8b8b9c1408e176e828c46

Observation 5749fc2b-6f54-41a4-8636-6289b596f651 · outbound

This paper cites Perception encoder: The best visual embeddings are not at the output of the network.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Perception encoder: The best visual embeddings are not at the output of the network

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:503977e193196ba85bd04af0b1d7d0359a17a0c1d605b47e7149fbbf8ecfc712

Observation 454d1247-f1ab-4ff3-94dd-f056b3f411d1 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Emerging properties in self-supervised vision transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:48a821733bad140b241b76465b533bcc16816b7b0365144f8a863a47ba4cd020

Observation 8beec1cc-4d1c-4b4c-aee1-66d02d65ec5b · outbound

This paper cites Maskgit: Masked generative image transformer.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Maskgit: Masked generative image transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:29fed41aabfa058f0bbe0c5b612268468f447d35b629e3f3e167f153ec2e88b2

Observation 4d778a25-db20-48c7-afae-87c9e024a33f · outbound

This paper cites Enhancing Vision Foundation Models via Multimodal Continual Pre-Training.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Enhancing Vision Foundation Models via Multimodal Continual Pre-Training

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-21T02:21:17.004501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:fac474ee757adca82197f839fe43f8b887f9e84a64e2b911717cf5372f45467f

Observation 22f0a495-3c75-4f05-baaf-050fe6b32d74 · outbound

This paper cites Vision transformers need registers.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Vision transformers need registers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:190104c2282424b0b9bc306f7745456fea6a67256dfb83fbaa4148d5e3bfc739

Observation b015d138-e0cc-42db-9987-83281921b62b · outbound

This paper cites author Dong, W.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder author Dong, W

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T13:10:55.714993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:fa3567be07379f88e9bad194b95d512ced734926aadfc908f8648b80f6df1303

Observation 23545e46-2959-463f-8052-510b0feae8c6 · outbound

This paper cites arXiv preprint arXiv:2511.23386 (2025) 4, 7, 9.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder arXiv preprint arXiv:2511.23386 (2025) 4, 7, 9

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:37:40.273079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:2bb952a0296260c285ce9d89acbb9658eec3b9fbc5d8f88bc836082b98ddff38

Observation 705bcac2-f4b6-4c43-8c5e-81cbf877b30d · outbound

This paper cites Taming transformers for high-resolution image synthesis.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Taming transformers for high-resolution image synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:b3adbe77722e5aee8d80d4b3312c95eddeac547a160bb73150e2eadab8738288

Observation 96559e00-7569-4b47-a837-4d40fda3ee64 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.215118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:78d918cc950f025749ab44736f7e4434974a3cc41360f33f5e084d8b71fec3f9

Observation 2720aa33-4bf1-4df2-9ede-0fe84d596420 · outbound

This paper cites One layer is enough: Adapting pretrained visual encoders for image generation.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder One layer is enough: Adapting pretrained visual encoders for image generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.222807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:d77aabee73fc8cc49a275fcb2f4d5b84398d3415c348fed0d8f0bdcec1e74152

Observation f0612a31-cdc8-4585-8221-891bffb55bc5 · outbound

This paper cites Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:27:40.600866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:a09547f018fd864a0494c1caf30c51255e2bfa20df906ab73e5968503e78f439

Observation e38956fe-f930-4997-938f-1ba8d2f9deaf · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:bec9277cdfc413805a983e16376fbec8d44ae4e7220487906efeeaa3e02ec7c7

Observation f0b29bcb-567e-4895-8630-319c76e13cbb · outbound

This paper cites beta-vae: Learning basic visual concepts with a constrained variational framework.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder beta-vae: Learning basic visual concepts with a constrained variational framework

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:51078fd3aabd3522dc82a6776615704aa6a62fb1b480e751c0d0e483d3d5b15d

Observation e56af5a1-2663-49f7-bc14-9c012f43a31f · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:27:40.598033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:b139cd74a822f158890d9cc30869cfcf1afd7b9c733724268b00ed79ef63c08f

Observation c6ba7bee-4500-4af8-926c-d3803b477c5e · outbound

This paper cites Image-to-imagetranslationwithconditionaladversarial networks.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Image-to-imagetranslationwithconditionaladversarial networks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:df9c73398c4ed82f71dabcb5835f4998a0b846fdad226a00735bc4d4a919a0bc

Observation ffcb0b9a-e1d7-4277-bca8-0e8a45bf70f2 · outbound

This paper cites Dino-tok: Adapting dino for visual tokenizers.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Dino-tok: Adapting dino for visual tokenizers

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.283600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:f3c928152436de34e8a840f4a6de246b6d03e4c6333c112cf69ee0b8a6eb9345

Observation 15cac6ff-370f-4925-956d-f63befdbadc3 · outbound

This paper cites IEEE Trans.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder IEEE Trans

Reference 19

Resolution
metadata mismatch
doi, observed 2026-06-27T13:10:55.711184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:4ee414408e9932586ba71a48ddc4059507b8271cb50c35ad1be0634713a590cb

Observation 6c2d7fc6-2cd3-4557-9d91-6999b73609ac · outbound

This paper cites Auto-encoding variational bayes.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Auto-encoding variational bayes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:a3260f537403f9f911186ed0211ee19674688945bd79297ef7dab8d8a41e23dd

Observation 61f31447-932c-481b-ab85-22acfc9a264f · outbound

This paper cites International Journal of Computer Vision 128(7), 1956–1981 (Mar 2020).

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder International Journal of Computer Vision 128(7), 1956–1981 (Mar 2020)

Reference 21

Resolution
verified exact
doi, observed 2026-06-27T13:10:55.722165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:0c23b64819ce69d28b2324337acbccdda84be8456dc818f5f1b1962aab88e2ce

Observation 1eab0458-211a-4198-96b4-33a72c865654 · outbound

This paper cites Improved Precision and Recall Metric for Assessing Generative Models.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Improved Precision and Recall Metric for Assessing Generative Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.270431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:2abd55a08a1a4d254bb5acd7ccc971d903b14c91448998a82faf0aaa4a8e9f3a

Observation 5427997b-3428-4439-85d4-924984e7d9bc · outbound

This paper cites Autoregressive Image Generation using Residual Quantization.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Autoregressive Image Generation using Residual Quantization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.275626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:3a3ac9cba6a23ffd48aa67e821b1c3003f6df28ec08397008a96169da60c4565

Observation 35d9ed1b-5a95-462a-9335-da525c8cd3a8 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.278177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:cdd435ef7ad83d33d3459bd395956f0e410e9a25ec621a6017a2e355f4dcdc7d

Observation 280769d0-f1ae-473a-925a-d3bc569611a0 · outbound

This paper cites ImageFolder: Autoregressive Image Generation with Folded Tokens.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder ImageFolder: Autoregressive Image Generation with Folded Tokens

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.280769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:0630b126e0e5fd439e38a18c0d96879bc199aec687d5a5e8a277013bb16c9c6b

Observation 0770a158-b49d-4e95-b517-b1abafc45488 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Evaluating Object Hallucination in Large Vision-Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.267740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:20efa975fc25f9d488d4c556402c42f149008600398880abc846bf6099390a4e

Observation 06c401a8-aacf-4f6c-8e64-a06402f842f2 · outbound

This paper cites Decoupled Weight Decay Regularization.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Decoupled Weight Decay Regularization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.238479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:109018d9adab306c234fb88b19639f5a234445c13e82ea21f83bdad8bd69458d

Observation 91a599a7-5ba8-4feb-a32c-3c19314911bd · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.217854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:d890dcb239b4ac39d913476f1ca3668461c14ac4aacec2cb839c1275dfd505e2

Observation bf7e34d4-748f-450b-bd7f-e26c7d52d446 · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.204519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:48419815b9d8c3911c854f5635820f6fcd7477d0c248255d34dfc8ec3280934c

Observation 80798ca6-a46a-486a-895f-f7cd2c7c48be · outbound

This paper cites SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.229595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:0d3810594b00bda5951b6489e33a2a6690f36e66e3edd68b99c39780c8545363

Observation 68859522-0407-4f53-af23-8ed249757b50 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:afd8085b9b7295e3a2793463c92d0e85835fd4a2999b2bf7c80357c9cf4cc7aa

Observation 1b5aea5d-75da-4423-9297-ff48dbb72219 · outbound

This paper cites Docvqa: Adatasetforvqaondocumentimages.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Docvqa: Adatasetforvqaondocumentimages

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:b9de11e49423b2acfc466b0ddd8dada1b322b5fb58baae47ecd1febe20794f4e

Observation 1d135103-6b18-490a-b525-88df7750b719 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Finite Scalar Quantization: VQ-VAE Made Simple

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.254393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:2ea1d250a05585afce5ba63f648d5b18a38e82ee89a8b1096078d7c37570001b

Observation c9a3ae1c-d32b-4e07-883c-3941afa5c276 · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Dinov2: Learning robust visual features without supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:689c07544f53711427f0ac7b75334e1db73313465e782ca166bd38ef035a641d

Observation 2ef515e8-27f2-414e-a768-5e56e8441a70 · outbound

This paper cites Scalable diffusion models with transformers.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Scalable diffusion models with transformers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:f5a30ecb689194f9e2a9b58ea3dbaeb77f732e516ded76dfe662da22e24ff02c

Observation 645d9226-af78-4eab-8101-401f697c3ac0 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Learning Transferable Visual Models From Natural Language Supervision

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.201979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:a52af26c3615a26cd1a1441998d722237d21709f9de8e1d35da23c93e13c1584

Observation 85ae22e2-031f-448e-ac44-1642b3a56903 · outbound

This paper cites Generating diverse high-fidelity images with vq-vae-2, 2019.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Generating diverse high-fidelity images with vq-vae-2, 2019

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:1c0b366428aa18424377ca16972d14fb81fda9ac2b3913926b7676c5b0d8184a

Observation 27012b5f-b880-4d38-acb4-d2ab42fb7377 · outbound

This paper cites Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.221265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:24fb42a4437cf79d7b1cb6e611d75dafce32baea26d1326baa481a2ef4122c56

Observation 04d5aa3a-92b2-4fb7-ba7a-5d72ac9ca1e0 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder High-Resolution Image Synthesis with Latent Diffusion Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.243542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:a0a8edfcff9b8c3ce6ca3d7f288604a2778c290261a9d7f5bd3f2162d7db35e4

Observation f574d2f9-70b3-400d-924f-340fea066395 · outbound

This paper cites Improved techniques for training gans.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Improved techniques for training gans

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:0f6661a0c721244e1818de9d805c1cb85e8206524e1c1074eaf1c95615c92c67

Observation 4d2b0276-6fbd-476a-b91f-72f1a8142944 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder A-okvqa: A benchmark for visual question answering using world knowledge

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:b6445c7ebe5b25bc51c542a63658be9aea2a8aa9165795c51b9902d00724e70d

Observation 0f377163-2b12-469a-aab1-b8df0b0e9e2d · outbound

This paper cites Latent diffusion model without variational autoencoder.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Latent diffusion model without variational autoencoder

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:ce88471490be71e330f73f18e964bdf63c7353c180d7bced7c33abb8aceee507

Observation c34d26b8-8024-4de0-92a9-4a6ca2db5991 · outbound

This paper cites DINOv3.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder DINOv3

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T05:37:40.225112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:7c6924f9e6db642f6c1c349a8de4ce2413298614694486aabfd2c662cdd89722

Observation dc92ee85-2095-4ae6-957a-2a24db50dc8b · outbound

This paper cites Towards VQA Models That Can Read.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Towards VQA Models That Can Read

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.212732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:ca5a97144de40913bceeca3ee46f2dd53c10645973032b1a1e1a73d97b0476e4

Observation 710f038f-f07f-482a-9292-0d1a144b9ba3 · outbound

This paper cites Dualtoken: Towards unifying visual understanding and generation with dual visual vocabularies.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Dualtoken: Towards unifying visual understanding and generation with dual visual vocabularies

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:5f074f18b558fae6c364a90c86c8b20aa6dbf7de443a5324cf96f9b03fd6898e

Observation 2057130f-78dd-48f7-9fb8-d10e9f279a6c · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.220334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:d42055d07a03e931b254f7c248ef80f052a9d4fdd4dc224fd3ed567230ccd0e2

Observation 16b262ae-4a66-49ad-961a-7f3be5a36dd0 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.259673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:cbceb20786ddf245c0b9126031fb1c930fa8a08bcbde6a5bf8457ac5707f62c8

Observation c9faba58-3a7f-4f41-bc13-2a93e00bb89d · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:e0aa1cc22413bf4febfb338eb31b2a63dd2f3671fc0814bf9bda79204914b219

Observation 933766f4-ca43-44d4-bb30-019b59d10094 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:ffdaeaf48e0aebf40b0ec9c2339ba607f7c7d24901228b7cf4950379be5dfe71

Observation 9479af7c-bb79-4a33-b089-80f4764bba12 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.251926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:33aeee47920232c7f37b76113567605bb6f24fe94685d6575a6f6cd8bd08640c

Observation 206f7527-d29f-4e23-9b64-878c060dcb67 · outbound

This paper cites Neural discrete representation learning.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Neural discrete representation learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:59b1ff37aeeccbab2cee06407979050b14348febfb2dbecbce5fdba5cdfaf68b

Observation ef620b3b-092e-4edb-b4a7-b18f0869d23e · outbound

This paper cites Neural Discrete Representation Learning.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Neural Discrete Representation Learning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.249509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:0c9a898f2a889ac9c790ca0f4d92a40886ed8eed002437996ff6b2ea8b806dc6

Observation 983ba479-7068-436a-8a45-247638c713ed · outbound

This paper cites Omnitokenizer: A joint image-video tokenizer for visual generation.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Omnitokenizer: A joint image-video tokenizer for visual generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:6ceb24dd426c563a98b29e190d6ee96d6c8a0842b870df2fc0451a04bae08a3b

Observation e2b1eae2-d31c-4c63-8d71-7af4e4f2e51f · outbound

This paper cites SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.254498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:94a3cc73bd7f1d35a14a08ca4fdfac6dd082e2b91b0266bc011aa5acae11ac1e

Observation 3a590ae4-e225-4455-94a5-95dc1de28959 · outbound

This paper cites Omnigen-ar: Autoregressive any-to-image generation.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Omnigen-ar: Autoregressive any-to-image generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:4a2a4871cc3bee46fda210caa731597b641c4bd82854053e92aafc020575aaf8

Observation 5fae378a-170a-40fc-a943-38bf7090da2e · outbound

This paper cites Representation entanglement for generation: Training diffusion transformers is much easier than you think.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Representation entanglement for generation: Training diffusion transformers is much easier than you think

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.257308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:595b18f51540b0b7c2aac51a6de015fcf7b3924009f713563a6d1e3daa63b942

Observation 0ff29571-d063-4e14-9be4-5b8da79761cb · outbound

This paper cites Grok-1.5 vision preview, 2024.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Grok-1.5 vision preview, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:bb2cef9a49ff0f68b29c6a5f5349defe6b64c5dee73417a8dd34b8623a9ab1cf

Observation f37b1219-676d-41bf-930f-686a8fbc47ef · outbound

This paper cites Vision transformer with deformable attention,.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Vision transformer with deformable attention,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:3789abff68893a43c1fc9bd42b643e466fb891ef446a384c6829f798d6b21ef9

Observation 2b4e737f-8506-4735-a904-a7407279b68c · outbound

This paper cites Vision Transformer with Deformable Attention.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Vision Transformer with Deformable Attention

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.262470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:680e298994669b1fbd730f6f9051b6b98f6d2de3d70f705dd411cd284b9ffd23

Observation f893c78b-399f-4961-91cf-efbd2244dc8c · outbound

This paper cites Videogpt: Video generation using vq-vae and transformers, 2021.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Videogpt: Video generation using vq-vae and transformers, 2021

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:461f8dd7c1c4bbf9118b60d5f8bcb4d31903f18505beefd0229b42912df612cc

Observation 71145cbf-138e-4362-bb5f-bd5b55f768a1 · outbound

This paper cites FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.238986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:fe6af3f68aaf9ac968e4f6dd7a9c2f2b0657c1be71d5a18dc3733cfde8ac8da9

Observation 93018c37-be11-4763-8fa7-f072f62d6b3b · outbound

This paper cites Reconstruction vs.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Reconstruction vs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:fb833173b84466ac7e62ad481b79945c333972f98e8a0e0ce121b9d224e49131

Observation 01cc8e07-66e7-4cf2-bfab-aba86488ab76 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Vector-quantized Image Modeling with Improved VQGAN

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.241892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:e54b19299ea423a02d34a79cfeffb59375f1775882b1a4182734d28bc383d71b

Observation 5a89d216-e243-4d74-a8de-d66eca737997 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.244433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:d6920659e302ef9541688c243300ecc9ab276c650051ab4d07edfa418f900a0f

Observation 7aeacc02-a825-47a9-a700-c0469cee457a · outbound

This paper cites Language model beats diffusion–tokenizer is key to visual generation.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Language model beats diffusion–tokenizer is key to visual generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:bbad24afa025329268d26bd392f9ee8a0029a8fdb872f169093069d20049dd58

Observation 7dc2bc2f-ae86-47ed-a132-8bdeb4826b94 · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder An image is worth 32 tokens for reconstruction and generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:3648630cf628670683fc5100b7151f35c0ead745c8436bff349f83a6e977eb64

Observation f3b63dc6-6667-448c-85d4-d60709bea98e · outbound

This paper cites Representation alignment for generation: Training diffusion transformers is easier than you think.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Representation alignment for generation: Training diffusion transformers is easier than you think

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:45139cd5d05a4a182e2ff52190adc43d403313745e4e6b27969db3402d8526dd

Observation 559fc6a7-bbc5-414e-82a8-f9eaf0a68718 · outbound

This paper cites Sigmoid loss for language image pre-training, 2023.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Sigmoid loss for language image pre-training, 2023

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:d18d371864a86359b128dac631dd078c57c1e0a0f692e526d7c9010400880493

Observation 5e869353-1d06-49d3-88e4-c37ff811eea3 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder The unreasonable effectiveness of deep features as a perceptual metric

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-27T13:10:14.308216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:dfc75c1b1224b5eb720b00fd320de588dc26f1c0a7329ce8d525bfcac4141cfb

Observation 5c4e9324-6488-4379-adaf-e8e7fc994412 · outbound

This paper cites Spherical leech quantization for visual tokenization and generation, 2025.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Spherical leech quantization for visual tokenization and generation, 2025

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:37:40.227816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:74fa961816049f4da58f61d2478ea0924623dcb485157c99ed4bb1e395b00e8d

Observation 8e48652e-9415-4f0e-ac53-12c6a299112f · outbound

This paper cites arXiv preprint arXiv:2507.08441 , year=.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder arXiv preprint arXiv:2507.08441 , year=

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:37:40.246115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:f666806680dd78e7f9cbd87fc62433d2cafa4b326de926444c7b56a31fc55a30

Observation da2f3a28-03e6-4da7-963d-580fe207e836 · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Diffusion Transformers with Representation Autoencoders

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.233253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:6482a21fe75a921ee751a43a6abbfbf848e28c2f1e0dd45e3dc39bd83d738b22

Observation 69277d24-07ac-4c3c-a60c-59bdb0c99a53 · outbound

This paper cites Fast Training of Diffusion Models with Masked Transformers.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Fast Training of Diffusion Models with Masked Transformers

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:37:40.236125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:559de9bfc6899118f6be81ff40c3c7436eccf9869519eb043f3425a71079d9b8

Observation 80afaa1a-edb8-48a5-bfd4-a95221c22ac1 · outbound

This paper cites Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:37:40.265134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:5b6103dc4e3d8d1ce05417a7f60029826268e9e3d56569ba0ed47749310e4911

Observation 9325f93c-ee1f-4110-b324-ed3a0dd9fb6a · outbound

This paper cites A Proofs for Section 3 A.1 Dense RD-AE flow We derive (6)–(7) using the notation of Section 3.1.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder A Proofs for Section 3 A.1 Dense RD-AE flow We derive (6)–(7) using the notation of Section 3.1

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:37:40.247132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:e5767cc4361bafd18254babdbfa5f35ef01398ecbd88fa9bedd826dca7b1339a

Pith citing papers

No inbound Pith citation observations are available.