Pith. sign in

Paper Citation Record · LEDGER

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

As of 8 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 3 inbound Pith citation observations for arXiv:2505.19602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19602 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:16:20.927090Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:03:03.062322Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T07:58:07.594523Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved75
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 293a0609-f139-43f6-8037-2b43a274623d · outbound

This paper cites Longformer: The Long-Document Transformer.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Longformer: The Long-Document Transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.031316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.031316Z digest=sha256:21fa01a490def0b67fc7a9d92f2af67c48da282be8643eb06f8d4f5bb02c2478

Observation a6f5049a-cb1c-4012-92fc-54079f5ef1b7 · outbound

This paper cites Improving image generation with better captions.Computer Science.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Improving image generation with better captions.Computer Science

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.089620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.089620Z digest=sha256:ae3843c5e2d3cdad09d033bbdc22edb49e9702ba5eab2e4f8f2c62386767bb48

Observation dffea209-3858-47ca-a9a0-3affff8ba103 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.217814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.217814Z digest=sha256:a8cb58349689a5dfbfbbab73c3596e780f777f05096ad02de55dd902105e82b2

Observation e9ea3931-fa9a-4d77-8236-6c48319287f0 · outbound

This paper cites Maskgit: Masked generative image transformer.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Maskgit: Masked generative image transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.340496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.340496Z digest=sha256:622c2feb063b8689d1a4ef7a459214d6a59ed2ed71ae91ea0d73070f329a3d4e

Observation 646555be-1f9c-4df6-9c80-d2e0bfdc8a28 · outbound

This paper cites PixArt-Sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression PixArt-Sigma: Weak-to-strong training of diffusion transformer for 4k text-to-image generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:24.117531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:16:12.414680Z digest=sha256:4effa72e4d799fac142cd61c7b3be305b322bf3fef75982612854905c0dfd058

Observation a14d7039-3002-422e-b24d-f6c3e6c903e8 · outbound

This paper cites Generative pretraining from pixels.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Generative pretraining from pixels

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.543575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.543575Z digest=sha256:83adc5a4e3905b6c1625f0cad1e6dc24daf5c8546ccb7638e4ffcfef53bda037

Observation aa1f79b0-8e6d-406e-8dfa-1e8a18d509a4 · outbound

This paper cites Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.670870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.670870Z digest=sha256:184f94e257f16cd02ffb0b99714fc2bfb8afbdcbc0eca127e4cfd632c4306e91

Observation 0fad3e3f-9146-4a5f-beae-d73abe0a52dc · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Taming transformers for high-resolution image synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:12.835884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:12.835884Z digest=sha256:5bba788de4e176286d3d694622dbaf77966bbf904ab7fb0142e76857f68439ae

Observation 0a678a28-5f88-44c4-8487-7a8823fc6d0f · outbound

This paper cites TinyFusion: Diffusion Transformers Learned Shallow.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression TinyFusion: Diffusion Transformers Learned Shallow

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.014126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.014126Z digest=sha256:a8b0a8870a49b09fbf7e258f5188ca3a94ed931b03837fd7ede4fe6b163e82a1

Observation 2e6847e7-cbbb-4746-be02-7c97c4414363 · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.138387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.138387Z digest=sha256:343d0f666a4f2eb077380df274970ebe6e25cf4be03df00bdd0705a5f460d436

Observation 8fd9a677-964d-498e-8609-a9f486163db0 · outbound

This paper cites Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning.arXiv preprint arXiv:2410.19258, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning.arXiv preprint arXiv:2410.19258, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.258212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.258212Z digest=sha256:b81971309a26a9aa13f8e5e3252ef676fa2bde78433b66f785164a00d3df2f64

Observation 814de7f5-fdf4-4a33-9933-76925e214302 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.413395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.413395Z digest=sha256:2707e8e3ef7ddd5c2fa61022aeda65b1629d24ce61e5bea8db8e28dc05a1b117

Observation dc481280-ccd0-491c-af0e-50b4974a2be8 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:23.862112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:16:13.530092Z digest=sha256:1f020d86649fb874b7ea6cc90cb42832969d38dabcf63aad1074ea5cdfd2761f

Observation 22f15021-513d-40bb-a2ad-7b7af0fd1226 · outbound

This paper cites DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.622854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.622854Z digest=sha256:172a9bd765e12373281d4f7358264a5af5faf1abc36382f72255d4f346a013a0

Observation 0d51690e-64bf-4e93-9317-19ab2db69ce6 · outbound

This paper cites FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.693484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.693484Z digest=sha256:819fd18a685df845c6873a4a43d38492096a0f56c2738b626c1bfcc0c61365e3

Observation c8144f2e-57d3-4053-bc65-a0aff67de974 · outbound

This paper cites Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.770499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.770499Z digest=sha256:6f10162e83c1c2c7a935de806bea315b9c006a645de03313b6f709ea0cc176ac

Observation e9cccbfa-a369-4d8f-bf85-2dddd7291d5f · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.857590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.857590Z digest=sha256:679cdc0c7154262af059a1d1f4dcff8ce43fb05a46fd0e27c3ad0cae2f32acf4

Observation 90c1a4be-0e2a-4a36-8855-c9899a98340c · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:13.948063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:13.948063Z digest=sha256:cc47f8eb4e5764f5e0f579011a89a8c6893a7b559b61175590e033871bddca5d

Observation b128b49e-cda3-4642-9df2-5a338a5297d4 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Distilling the Knowledge in a Neural Network

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.033748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.033748Z digest=sha256:dff3f4b301c981735913d876a98615be7ffd32d026dbc58c91ccd4f14cc00764

Observation a26ea0ad-6364-4f16-92b9-5d354a338c16 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.124213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.124213Z digest=sha256:0d40bbedb0b948e0598d2d2a0679de0e621df1f8ce9dc94843cc07856579f454

Observation e4045744-1d75-4e87-b00d-b7d213d10524 · outbound

This paper cites GPT-4o System Card.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.194123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.194123Z digest=sha256:3321a8191896f2b2fbf56ceeef676e425911944bd43c643dba88fdd1bb960d78

Observation 38eff18d-581c-469e-8f75-d9b88ac6bfaf · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.282281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.282281Z digest=sha256:0b5fc3b9da8cacaf76111ee3e60a73697cfec37488c6f925085f97b76970ae0d

Observation c1402ed7-f490-4546-badd-4fd7f608e0b5 · outbound

This paper cites Scalable autoregressive image generation with mamba.arXiv preprint arXiv:2408.12245, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Scalable autoregressive image generation with mamba.arXiv preprint arXiv:2408.12245, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.444559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.444559Z digest=sha256:de5893a6c998c711b33eca503e6c7f1f8fab226a852740cccb41550688b91925

Observation 2b8bf013-ee0a-40cf-bd57-999abc287a7f · outbound

This paper cites Svdquant: Absorbing outliers by low-rank components for 4-bit diffusion models.arXiv preprint arXiv:2411.05007, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Svdquant: Absorbing outliers by low-rank components for 4-bit diffusion models.arXiv preprint arXiv:2411.05007, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.611793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.611793Z digest=sha256:169883a0dff73d8251cb371368b6cbde9864819618ed4fd87ca5a9b06bb22802

Observation 06423f51-cfd0-4ae9-88ef-0b4c16487a02 · outbound

This paper cites Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.731278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.731278Z digest=sha256:f5c08d36bf0e1ded57d1ca494a3a179971fd1620a624709b41fbd5521ca93099

Observation cb8defb5-873d-46e9-9eb1-3f9540d81714 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Autoregressive Image Generation without Vector Quantization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:14.884901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:14.884901Z digest=sha256:8de34aaf1857539f9e361fefcb7151ecf50cbed575f93616036a6444716b3d66

Observation 62903964-d2cc-40eb-8e34-f2ca28e166b2 · outbound

This paper cites ControlVAR: Exploring Controllable Visual Autoregressive Modeling.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression ControlVAR: Exploring Controllable Visual Autoregressive Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.022702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.022702Z digest=sha256:d0362a61b48de9a284f77a591c759c942fabe0f7da265cc648f5a65fcdd5e209

Observation 4026be12-2d5f-44b1-a227-95b768acf7cd · outbound

This paper cites Q-diffusion: Quantizing diffusion models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Q-diffusion: Quantizing diffusion models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.125418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.125418Z digest=sha256:b1eb66f4726242b614a77fb3528aabbbd7cdf2208ab8fd0c428025e521e0d8b5

Observation 29d6db59-5830-46f1-a198-ef87d091c6e3 · outbound

This paper cites Snapfusion: Text-to-image diffusion model on mobile devices within two seconds.Advances in Neural Information Processing Systems, 36, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Snapfusion: Text-to-image diffusion model on mobile devices within two seconds.Advances in Neural Information Processing Systems, 36, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:23.563170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:16:15.281994Z digest=sha256:25d94a91257d5d2684fc1028f39a64a144d6535137cf9ce5d2e42357d3088e13

Observation 8e0543e4-440f-4b44-b65f-4a42fa18ac23 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression SnapKV: LLM Knows What You are Looking for Before Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.378042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.378042Z digest=sha256:e91b472b4bc6df86910738914e7d6eb3d8cb0925c54ae6ae1bd3b412c1187bbd

Observation 6fd371e8-a244-47b5-b834-3ef570d38900 · outbound

This paper cites ControlAR: Controllable Image Generation with Autoregressive Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression ControlAR: Controllable Image Generation with Autoregressive Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.523883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.523883Z digest=sha256:98a32495105d2924f36353745c7fac871e02f1bbf7c572dcab7fa6fb3abdba4c

Observation a495cb52-b5cc-4ba3-9c05-bbf931b8765b · outbound

This paper cites Microsoft coco: Common objects in context.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Microsoft coco: Common objects in context

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.657471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.657471Z digest=sha256:4302c334f668d0b1525dc0ba6551ef768c347cc5de42df0f3ae7a8118c0e7413

Observation 46aec037-8f7c-40db-bf10-3ec3a2194891 · outbound

This paper cites MiniCache: KV Cache Compression in Depth Dimension for Large Language Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.805734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.805734Z digest=sha256:4d01ff11499cdd0f490b9b70203ad07f8db365d3cf42e55f0e6cd9ac9618fdf2

Observation 1adb9af2-121c-41eb-81d2-c622f8fed257 · outbound

This paper cites Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:15.891755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:15.891755Z digest=sha256:ca55fccdc279bf2cbe4374a3c03aa0ad5a162252d074209a0aeaafed4cd4d39d

Observation bebf9bab-5a0c-4ab4-b687-370c985bcc0a · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.Advances in Neural Information Processing Systems, 36, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.Advances in Neural Information Processing Systems, 36, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.016184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.016184Z digest=sha256:0ad6d563a36b771e9f9bd27d05e047421820c62f7ed10a6a5308eeca000875a5

Observation 6c237317-f6f2-4a3b-8185-92f5aef59083 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.195264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.195264Z digest=sha256:90e01763471b4c62d45ea1c2c443460ed189a0a5eff90c8ca3b718edc6c208ac

Observation f0675899-2855-42d8-8823-4a6794e46234 · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.358206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.358206Z digest=sha256:d2a82c016c26a466aa2b9bc33b385e8fbffd7a66a21a39d1ceb258938f16d770

Observation d5edde25-459b-4fdb-8432-9e94b657d3a9 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.536400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.536400Z digest=sha256:b10ac39831953d5d5f5b0c8709d1e47220197fec4ac5d91c3b2a8f726cbafa05

Observation 6133ddb2-fc52-429d-8bbf-4dcf6ebadf01 · outbound

This paper cites Accelerating Diffusion Models via Early Stop of the Diffusion Process.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Accelerating Diffusion Models via Early Stop of the Diffusion Process

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.665647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.665647Z digest=sha256:1c76dd37e636f3865832879110e66971efad2855df479a78606bfbdae1b43f38

Observation 4eb1f813-b788-4468-b191-0668b04482d3 · outbound

This paper cites STAR: Scale-wise Text-conditioned AutoRegressive image generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression STAR: Scale-wise Text-conditioned AutoRegressive image generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.780411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.780411Z digest=sha256:c4d57fe0dfe4ca169043f3b9b911047bd040ee6a834af742308df1094474dca6

Observation 2eec77e1-4b91-470e-85ac-3b668055466b · outbound

This paper cites DeepCache: Accelerating Diffusion Models for Free.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression DeepCache: Accelerating Diffusion Models for Free

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:16.890592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:16.890592Z digest=sha256:bcaee91b8ccd19cca949bec805e7d0405b28220a225eaf0f9cf9ffdca182b3b3

Observation c33b358f-77c5-4223-bb69-6eaff1831127 · outbound

This paper cites Transformers are Multi-State RNNs.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Transformers are Multi-State RNNs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.009571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.009571Z digest=sha256:976a460a9505a63d42aa2528e8ec6101e0f112c84f8cb321f13816d7fbac0f36

Observation 3c17afa5-d337-45c2-9e6c-6c4729a30143 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.118171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.118171Z digest=sha256:ac51afe2efd964fa69eb865190f66ca67a3bfb2b41d7cbdaa9f0e3fbfb1b55f9

Observation ba685361-ef66-4fa7-89e4-8b276a81e4f1 · outbound

This paper cites Head-aware kv cache compression for efficient visual autoregressive modeling.arXiv preprint arXiv:2504.09261, 2025.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Head-aware kv cache compression for efficient visual autoregressive modeling.arXiv preprint arXiv:2504.09261, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.201417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.201417Z digest=sha256:aeaf9e6ab6ab7d6faf0b4230369f71921d96510a79794a13ec972abd1f3841cf

Observation e9c6dae5-05e4-4040-a6e6-e81f9537114c · outbound

This paper cites Efficient Autoregressive Audio Modeling via Next-Scale Prediction.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Efficient Autoregressive Audio Modeling via Next-Scale Prediction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.266768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.266768Z digest=sha256:a324b329144d176cda8be89938182dc8c036af25dd21fb220a333c1847a4c447

Observation c21181a6-5cea-470f-a97a-05134b04e7ca · outbound

This paper cites On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.384573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.384573Z digest=sha256:8001b77cf2dca004726d5deec92cefda0abfb411517a60239119890b02ba3558

Observation 4a14f778-46c6-4d1f-a4ec-dcbf3ab8e720 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Progressive Distillation for Fast Sampling of Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.490839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.490839Z digest=sha256:8e8711b96f1aa92d59c36d65ffac3e49ceca982bbed35b4d7444df2f4ec08ef3

Observation aa9df552-5935-49c0-8aa1-7d1032654f6d · outbound

This paper cites Adversarial Diffusion Distillation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Adversarial Diffusion Distillation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.578156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.578156Z digest=sha256:55fd5643e3ad2938945d8bff66d55ac08b8ec8e2228fed76c7158c229e101e56

Observation 3d0a08cb-ef61-469a-b213-c19bf6a9ab4c · outbound

This paper cites Post-training quantization on diffusion models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Post-training quantization on diffusion models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.679607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.679607Z digest=sha256:2ce98edc671a60000835d4745d9b81c5281d029fd6245fafa7473906832d27e7

Observation a8693224-84aa-45a6-b7c3-dfc107049171 · outbound

This paper cites FRDiff : Feature Reuse for Universal Training-free Acceleration of Diffusion Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression FRDiff : Feature Reuse for Universal Training-free Acceleration of Diffusion Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:17.772212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:17.772212Z digest=sha256:45aa722324c0fc9ca6696ada07be266875ca9ebe9088542d9f8d08e607a0cc6b

Observation 353a2226-eed9-4e36-903e-a49527443e35 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.009369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.009369Z digest=sha256:5e758a91f4eacb0a323ffdc8591dec5d16bd96bc46a0e2ea8a5bedbb83ee0d6c

Observation a8d75859-c32a-486a-af2c-e1c44e89e10f · outbound

This paper cites HART: Efficient Visual Generation with Hybrid Autoregressive Transformer.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression HART: Efficient Visual Generation with Hybrid Autoregressive Transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.127539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.127539Z digest=sha256:caf13fd85ca682794463865fab33b4184c54c743869dc99d4c4bc0e9c81f7918

Observation 4080b004-8286-4e92-9045-62ccd4f9bc2c · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.201368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.201368Z digest=sha256:6e115d9535daaba0800be92ee43a12f3066c9366c1ecf4f5bf6d601f06a3546a

Observation 456bd8c8-8727-471a-935e-d761f58b42f2 · outbound

This paper cites Conditional image generation with pixelcnn decoders.Advances in neural information processing systems, 29, 2016.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Conditional image generation with pixelcnn decoders.Advances in neural information processing systems, 29, 2016

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.298675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.298675Z digest=sha256:8e14b8ddcbe30476f7e75969ccb65a471b7f4a2cb71fd0d78584da6b18558bab

Observation 1cb61cd6-bea0-49cc-9630-78d1767074d0 · outbound

This paper cites Neural discrete representation learning.Advances in neural information processing systems, 30, 2017.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Neural discrete representation learning.Advances in neural information processing systems, 30, 2017

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.412301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.412301Z digest=sha256:2c7569fdc953e5cfeb77ae6dcfce1a33b7839c73ea59f2072cd275d5993efa86

Observation 735fab92-1f44-4700-825d-a90d83fc65a6 · outbound

This paper cites D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.517957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.517957Z digest=sha256:3da7bf2c4faa9e47814882600486e13d3df1c034d17cac8dbe305eb00cdb5836

Observation 1cabc8b2-404d-4115-8940-3fe58a98ed21 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.606283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.606283Z digest=sha256:2a43241de6e75a21f625bb96f60414ae7ecf3170ae04bbd90c2afbce2ebb4d81

Observation aa87dea9-225f-47b8-8879-38d93ae314cb · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Emu3: Next-Token Prediction is All You Need

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.699512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.699512Z digest=sha256:e6a348ef6fba5068e8e8b2ed3bb82f173750362cce6675186e57e1ab301fd4d3

Observation 975c5e03-be46-46c4-8191-fea0a7ff031d · outbound

This paper cites Cache Me if You Can: Accelerating Diffusion Models through Block Caching.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Cache Me if You Can: Accelerating Diffusion Models through Block Caching

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.760250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.760250Z digest=sha256:6785ef8a30c28c30067f01ef6a12807b96e1d407518e501373413de8270cf821

Observation e9eb3f00-e6d4-4429-849e-e68f1db1311f · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.842698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.842698Z digest=sha256:5c0927bc0d25a7ba7b44be596e77259b2d15e74bde2dc0c2f7de80dad0f3ca93

Observation 545bebb5-615d-4ecf-b2f1-754fcf5878c5 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:18.961783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:18.961783Z digest=sha256:cc8bfeee5778aa177547aa68f26ddfb458ac29909ef8faa58a2fab80534c55b5

Observation bacfe7e9-9139-4137-abab-317b6f2257e6 · outbound

This paper cites Efficient streaming language models with attention sinks.arXiv, 2023.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Efficient streaming language models with attention sinks.arXiv, 2023

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.047855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.047855Z digest=sha256:821ea624668609554739173fd95f2d3b82b39d04fff556a32309683b1017fbc1

Observation 8655804e-2349-42e9-9902-0e4d05f68341 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.134844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.134844Z digest=sha256:5c0e9f3d3b291c7a5ef4afd4f675bd82f2e01e42e18faf34997a374750f447a0

Observation f56e95aa-f6be-4e5f-994b-a7f82c45c05c · outbound

This paper cites LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.250138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.250138Z digest=sha256:d760e1c72e8e3c0104b2a1344b5e7ed860d601ea09e6fc55b9931396c213323a

Observation 827a1ae8-2935-4ac9-96d7-895870c2b1be · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.360068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.360068Z digest=sha256:5e8740bc632ca5151e9ee57aa89c0becce4f48945954532148e146471c0d33c5

Observation 8d6893d3-7a1a-463f-9bef-f80db39d814d · outbound

This paper cites Hash3D: Training-free Acceleration for 3D Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Hash3D: Training-free Acceleration for 3D Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.475008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.475008Z digest=sha256:9099d85914f34a57643da247b18340bdd521ba5b4ce381148b1f751190d613e6

Observation b7a8c97c-e8de-47ce-bb5a-d8e082b59bd0 · outbound

This paper cites Diffusion probabilistic model made slim.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Diffusion probabilistic model made slim

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:23.209684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:16:19.548827Z digest=sha256:a66fb608259b41fa084c981172109224d8a5b707a99b8c087f3873046a93d14e

Observation f523d14b-49a3-43af-9877-e8c9330bc139 · outbound

This paper cites CAR: Controllable Autoregressive Modeling for Visual Generation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression CAR: Controllable Autoregressive Modeling for Visual Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.608040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.608040Z digest=sha256:b0a9a267abcb43b46e09896ae405091995841a3b951ba76d8cc4598c195896f8

Observation 0f4c3011-d88a-411d-8bb5-7ed536fb36b2 · outbound

This paper cites One-step Diffusion with Distribution Matching Distillation.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression One-step Diffusion with Distribution Matching Distillation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.701607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.701607Z digest=sha256:2ab19115448a754e9b09d556e7f58de18abc7a0d125126fbc93054c74d84f105

Observation 0e775c07-916f-4896-ae0c-6090247f3f33 · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:19.784103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:19.784103Z digest=sha256:a2c9f1d1d55f974e63fd564d928bc9254c4d9bb326480a66bee24741e4c44622

Observation 3557f87b-5aab-4826-a750-14232c37e679 · outbound

This paper cites Resshift: Efficient diffusion model for image super-resolution by residual shifting.Advances in Neural Information Processing Systems, 36, 2024.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Resshift: Efficient diffusion model for image super-resolution by residual shifting.Advances in Neural Information Processing Systems, 36, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:22.897167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:16:19.898686Z digest=sha256:630d888b7ee645ff613ffb03d0bd107d00c759f98984993c463623aa4a9d58e0

Observation affe042c-ddf2-4a81-9f3d-0a82e5985e68 · outbound

This paper cites LAPTOP-Diff: Layer Pruning and Normalized Distillation for Compressing Diffusion Models.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression LAPTOP-Diff: Layer Pruning and Normalized Distillation for Compressing Diffusion Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.004397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.004397Z digest=sha256:8c86c854259a35ca107b9448f85b15a9798572b8c36eb44cea70a2ff2b9429a6

Observation aaa1b858-7fb0-4d7a-8413-b7a35de4a838 · outbound

This paper cites G3PT: Unleash the power of Autoregressive Modeling in 3D Generation via Cross-scale Querying Transformer.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression G3PT: Unleash the power of Autoregressive Modeling in 3D Generation via Cross-scale Querying Transformer

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.116154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.116154Z digest=sha256:3523612156ec2de862d409592b5af74160c344e2cbc3df8fe05d2db4336dca1a

Observation 91b9fa70-7ee9-4c17-b71a-2ec91cd63095 · outbound

This paper cites VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.208413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.208413Z digest=sha256:6e87861eee43bf7702459856052af2f95b5077ecb807fea9cdd40f1c8410a802

Observation 9e2adf34-da4a-445f-a409-4db203c50c08 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression The unreasonable effectiveness of deep features as a perceptual metric

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.320593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.320593Z digest=sha256:e2c283e6a79795fddaaafea4e83247fd2963721fe6036d16d78ebf3715fec80a

Observation b14eb1c8-94d5-4847-9600-356c9fdc9949 · outbound

This paper cites Faster Diffusion via Temporal Attention Decomposition.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Faster Diffusion via Temporal Attention Decomposition

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.398774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.398774Z digest=sha256:2c686d87a2ac5ec90fa37bcaa8d12b36148cee5bd0791bf46d81b9c65cd12fe3

Observation 55a8508a-c4ac-4be5-804e-d2fb45026751 · outbound

This paper cites Cam: Cache merging for memory-efficient llms inference.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Cam: Cache merging for memory-efficient llms inference

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:22.625939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:16:20.509944Z digest=sha256:17b0e5bc2832a40a961a941f3c0317f654470f9920408a21ac41901ec19a929d

Observation c37d9c57-6909-4482-b493-26005cd2d039 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.627739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.627739Z digest=sha256:2278bea4eb9218525de14f75635e64d41c3caa6ef42b29f3d1ac08a34ad9adb9

Observation 30e5a746-e8e1-4266-af0e-8df43ff5936d · outbound

This paper cites Real-Time Video Generation with Pyramid Attention Broadcast.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Real-Time Video Generation with Pyramid Attention Broadcast

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.706471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.706471Z digest=sha256:aa25baf97807bb78afe47e6dd52c5001668826d042346d287fca1417d25840cd

Observation 76094485-c0db-45ea-b5ce-5a4b36647ba1 · outbound

This paper cites MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.821528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.821528Z digest=sha256:7a01c9c2ab9d3c39b970ffd5838ed6ee9684938ec72b7d68dfa5d78f33d196f8

Observation f972b876-9b38-4a37-98c5-2063619a429b · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:20.927090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:20.927090Z digest=sha256:e2a1b610d555a4aa785464ab8847cb863d00c58d2b8227a163043ccabea98e7f

Pith citing papers

Observation 1a3ca253-3931-42c1-bdfb-79ffa010f511 · inbound

Visual Implicit Autoregressive Modeling cites this paper.

Visual Implicit Autoregressive Modeling Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:41:18.325747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T15:20:03.981830Z digest=sha256:32d02e77e061eeaea65ea65ca189e019e9c4863b015321aace527abbb3aa3eb4

Observation 348b6ba1-7a44-4b8c-91ef-076c5be59bb4 · inbound

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond cites this paper.

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:58:07.596260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T07:57:51.032025Z digest=sha256:801e3160c85c038709edba135a77f2c8b5a71f7d3a8f05860e9a25ef50da52ac

Observation 14f8186d-8235-4945-ba0f-c792e5ebda9f · inbound

Token Radius Attention for Efficient Video Generation cites this paper.

Token Radius Attention for Efficient Video Generation Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T06:03:03.062322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T06:03:03.062322Z digest=sha256:1b63b49d103e8bdb1065f3e75c5b1656cd71c150feb72dcb5dc88190b9c43761