Pith. sign in

Paper Citation Record · LEDGER

ARCON: Advancing Auto-Regressive Continuation for Driving Videos

As of 15 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 1 inbound Pith citation observation for arXiv:2412.03758.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03758 v3

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:11:53.056934Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T13:39:48.225482Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T13:39:48.369773Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ca0bd72-10b2-4c9b-ad86-bd3be63cb624 · outbound

This paper cites 4m-21: An any-to-any vision model for tens of tasks and modalities.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos 4m-21: An any-to-any vision model for tens of tasks and modalities

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.297674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.297674Z digest=sha256:417807896ea939c2d62d770e27ef3c81b7c88ec15dfbf3f99d42bc92812cee85

Observation 9e947db1-1ad2-4d3d-b071-b06f79762e65 · outbound

This paper cites Sequential modeling enables scalable learn- ing for large vision models.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Sequential modeling enables scalable learn- ing for large vision models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.305995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.305995Z digest=sha256:1136947e6c20309441029fe1c2787c845c6fbffe0501071e8b7a49fc8c8e5f50

Observation 31befbeb-4b35-434b-b1d1-e235947529eb · outbound

This paper cites Video pretraining (vpt): Learning to act by watching unlabeled online videos.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Video pretraining (vpt): Learning to act by watching unlabeled online videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.312549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.312549Z digest=sha256:c58eff8316d40d9c5af1ee00cf39d44f611cc1476d0ea01221945655dd17dc26

Observation 1f506d08-5f11-4717-8f22-93b7a2495b5d · outbound

This paper cites Beit: Bert pre-training of image transformers.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Beit: Bert pre-training of image transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.320069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.320069Z digest=sha256:39e16ceab219569bf98c684415d20470b3f54c6e8ccf7498be2e9c30f0abfab4

Observation 61762b9d-4ed9-4a66-8d49-9be50bab4b6c · outbound

This paper cites Video generation models as world simulators.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Video generation models as world simulators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.328107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.328107Z digest=sha256:ab5f5364878cef5b6e0c528064384c4d063d2e2a0945fc39cc9cd7993f57fbaf

Observation 4cf59159-3bff-4ab0-8344-a1957863856a · outbound

This paper cites Language models are few-shot learners.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Language models are few-shot learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.340294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.340294Z digest=sha256:419e4b71f0378c9873e1b1de02ea1d0d08913af5886d7657b8b3c10018f064f7

Observation 9a913dc4-cbaa-4a16-a7a5-7dff43bf0af2 · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos nuscenes: A multi- modal dataset for autonomous driving

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.349287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.349287Z digest=sha256:35cfa7ffd7080d48cf96344bf5b8e2f2449e98085504a5cbdb7baa71200f3ec7

Observation 010ab84f-bd40-4c1f-92c0-b7f188af3b6d · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Quo vadis, action recognition? a new model and the kinetics dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.360824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.360824Z digest=sha256:b26b9f9803030f202c3b9ffc4eeab2de541e9faae0df5af267f6d7a75058f2dc

Observation 59435cbb-24e8-4fd4-801d-fc2112c07d8c · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos A Short Note on the Kinetics-700 Human Action Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.374092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.374092Z digest=sha256:a790db80eb3bfa26dc3f899335ffa025be3a286ede94c408ae8d7076d653006b

Observation c9d73ed8-ea3a-4b3c-8899-bd749913039b · outbound

This paper cites Maskgit: Masked generative image transformer.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Maskgit: Masked generative image transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.382383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.382383Z digest=sha256:e17348d45c3d169add59a43058f6363a6b18617325405ddb8aaabb8534f1664a

Observation 6d2dc819-01cb-4cbe-853b-d702564254de · outbound

This paper cites Generative pre- training from pixels.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Generative pre- training from pixels

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.387557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.387557Z digest=sha256:f5fe8ae63c67d1c00dddf2551e71444473c2f955b060d602c50a8f3d6b31aba0

Observation 91d892b5-28e9-4cb3-90f4-3172399cde5b · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos An image is worth 16x16 words: Transformers for image recognition at scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.405864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.405864Z digest=sha256:f2c43642f127260546a1a5d0808cff8affcea70af54877ea5776eb380b61959c

Observation 3bdfc64c-4c0d-4bc2-87b8-1fcfc78fffae · outbound

This paper cites Taming transformers for high-resolution image synthesis.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Taming transformers for high-resolution image synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.414114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.414114Z digest=sha256:379f3fb3249bb12f24f8ad9c63a333f793e1cf9da9afa18a896aa4ef3d13fca0

Observation 8c3a5c70-61d2-494e-aa1f-6ab3d30426b7 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Taming transformers for high-resolution image synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.423018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.423018Z digest=sha256:b908819b41f96cee93fb1e4afb277b569507894d9f7b417584be022843f2fcae

Observation 59a65472-ec2b-4512-a023-4b9d749e5a87 · outbound

This paper cites Restructuring Vector Quantization with the Rotation Trick.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Restructuring Vector Quantization with the Rotation Trick

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.437454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.437454Z digest=sha256:fccc7ab78b0f9ca37082ae398d94a49e20a457d93ca5b0718de4089b0a1c275c

Observation 220940eb-0a4b-41a9-bbc9-6effe629c0b6 · outbound

This paper cites Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.444677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.444677Z digest=sha256:ec746c4343b8dac7cc6b3c1eb97a6b616f82b51f4ce42bdb914693b08fcee072

Observation 71366443-ffff-443e-8ee4-2f80e2d9142b · outbound

This paper cites Simvp: Simpler yet better video prediction.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Simvp: Simpler yet better video prediction

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.937035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.450682Z digest=sha256:99c160162aa5c22728c2470cbbcaf369b94889e1d7935a1156376075fe75a5d8

Observation dfd16095-f421-49c8-b3a9-94e8fe1f87fe · outbound

This paper cites Worldgpt: Empowering llm as multimodal world model.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Worldgpt: Empowering llm as multimodal world model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.918504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.457671Z digest=sha256:459f7bcf9cbd9a23510ec646a98effa9bfafcc535e177439b5e2a854d00daaee

Observation e445c23d-9858-4d3d-988d-a8b201528cdb · outbound

This paper cites Imagebind: One embedding space to bind them all.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Imagebind: One embedding space to bind them all

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.894646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.474801Z digest=sha256:5904742a98da8d28e7ca6b2144a3ec269f8ab8d1d3e73a93c1d8c09f5d4203e9

Observation b124ba76-0c48-474f-b54b-9c607d9288c6 · outbound

This paper cites Photorealistic Video Generation with Diffusion Models.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Photorealistic Video Generation with Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.484101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.484101Z digest=sha256:31f05357eceb6d6a4c88692aedd579a51dd7cda328114cb6e98c6f975c2f20c3

Observation 0efcf9cd-b4bc-4348-9ca6-fd4bed4ff06b · outbound

This paper cites Recurrent world models facilitate policy evolution.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Recurrent world models facilitate policy evolution

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.872614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.492422Z digest=sha256:870f9ff3f1be1fbfc0b9d7b4cd73a6b1ba6e8e047251d9fcea243ed443523e00

Observation db28c11b-4964-4825-8e3d-3dac0c8b56ab · outbound

This paper cites Mastering Diverse Domains through World Models.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Mastering Diverse Domains through World Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.502368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.502368Z digest=sha256:9743656109615ca7bbb4611c9cabc4a68b81eaac7a2c94b798fb1d76e1ab0c0a

Observation 56f010cf-b3cf-4a15-8230-8b9d28715e49 · outbound

This paper cites Masked autoencoders are scalable vision learners.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Masked autoencoders are scalable vision learners

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.510644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.510644Z digest=sha256:c11cbb54d387ce8c36949ba542717f4c1f81a875460377458c2733daed025685

Observation 9fe52c10-76d8-4a16-9143-c8f081afe4cb · outbound

This paper cites StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.518291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.518291Z digest=sha256:da7001da2575b1b73ab4472a4d1ea98ca939f3cc36ac099e7210026ff607397d

Observation 2169e6cf-dfc4-4a12-a849-2f05c3835ce7 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos GAIA-1: A Generative World Model for Autonomous Driving

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.524233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.524233Z digest=sha256:4c1bbc31d667c4206d1a462128cfa6a1895222bc867020d1ab5f595aee35f82b

Observation 3a5c4acf-af3a-46f0-b3f7-b16466cf5001 · outbound

This paper cites A dynamic multi-scale voxel flow network for video prediction.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos A dynamic multi-scale voxel flow network for video prediction

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.839343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.541876Z digest=sha256:56ffa9f7b42a779b442ef5bcdeb4cb9b59db8c7f8041455ff05bbb074de82d09

Observation 9acad089-1f21-497d-9a75-b60439efa4ba · outbound

This paper cites DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.549073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.549073Z digest=sha256:e9e8bf890a3da411e01e85eecfce8943a3ce5ee8bce53416a2395a73a5c93ea8

Observation bc66bf69-6e96-4243-8ffc-252775dee326 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.558137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.558137Z digest=sha256:e6da870c55be88604b8052ba3966ed6ad5e3020b85a9b1f0b049c26e1db1ac89

Observation b6e9f28e-5bca-4e47-8bfa-838266e33b5d · outbound

This paper cites SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.567200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.567200Z digest=sha256:69dfe774f776bb304266b5a1cc963701086b81b91738e62dddaa6ef9a3a908f2

Observation 61d637d6-27b8-4ea2-9a24-08228bf1b978 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.814780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.580870Z digest=sha256:76cb5d6baadea461022694552aa72696e3da03d7de6f8e0b3f83808f710f21ba

Observation 8c0d0e86-7dfa-417d-b201-c0071c70df32 · outbound

This paper cites Drivegan: Towards a controllable high-quality neural simulation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Drivegan: Towards a controllable high-quality neural simulation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.595832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.595832Z digest=sha256:aa695d49877fccb2d588ea64f97f1c0d0b444d47ae290eb42808a650d331fecb

Observation a013a7e9-65ce-4223-b9ac-3592cbda47c0 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.605598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.605598Z digest=sha256:7616605413f76554ce1d0cfd97254bd58408eb14a84139feb0c4975fb368287b

Observation 166e6dcc-650b-4a0c-b2eb-ff66d72b9ddf · outbound

This paper cites Uniformer: Uni- fying convolution and self-attention for visual recognition.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Uniformer: Uni- fying convolution and self-attention for visual recognition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.783133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.617290Z digest=sha256:d1790c417f3b100b064964a8822f914187f735b002bd0dbef3dd3c07ac7ae339

Observation a0fc5fd5-fcb8-4acf-a030-ba2158849880 · outbound

This paper cites DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.623864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.623864Z digest=sha256:7bf7c5ffff195c14a62b729bb3b4478b8293cd780513b067415bee282a0cf133

Observation f6e91049-1e55-4efd-b715-1b5b4d001bf4 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.631572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.631572Z digest=sha256:530580c5468e89d5806f477ee0a607d288ad2af454eb65aa063aa098c3f69631

Observation b941f491-f909-4c8a-981d-78383969674f · outbound

This paper cites Video frame synthesis using deep voxel flow.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Video frame synthesis using deep voxel flow

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.758689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.639594Z digest=sha256:95769cf3a2159b88cbed0572af390d52c57cf002f8e2e74fd9ab5eaca87695d9

Observation e8f5b1d6-03ca-43bb-bde2-bd06d94cbc77 · outbound

This paper cites Vdt: General-purpose video diffusion transformers via mask modeling.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Vdt: General-purpose video diffusion transformers via mask modeling

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.738155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.654428Z digest=sha256:1d09631359873979f338321ce7b0fe93e71ddeb167c2b05cb132d37ada2d097a

Observation be619297-449c-4bae-803a-08a87feded31 · outbound

This paper cites Wovogen: World volume-aware diffusion for con- trollable multi-camera driving scene generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Wovogen: World volume-aware diffusion for con- trollable multi-camera driving scene generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.719005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.660691Z digest=sha256:1861ebaf141956358526a24ca8a683563fe4766fbccc64140f8bd642b3bdc3f6

Observation 4a7dd1c6-ebd6-44a8-a51b-29eedd3a59fd · outbound

This paper cites Masa-sr: Matching acceleration and spatial adaptation for reference-based image super-resolution.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Masa-sr: Matching acceleration and spatial adaptation for reference-based image super-resolution

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.701181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.668555Z digest=sha256:13fc527addac078b0b78a2059691cbe221ab83bbd938ece59d715f807c0fbd0f

Observation e18c1ab8-b42a-4b92-b71e-076cb7290abe · outbound

This paper cites VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.676730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.676730Z digest=sha256:8b90cb89f78eb87ba9211cfbc98be7f71585caf5a89506f1c0fa986e92f37a8b

Observation 81c6410c-c77a-4556-aead-686ffd165184 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.684577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.684577Z digest=sha256:06c993a926abeaa6434191dbacdb3fe83c1fa6ecf0b177b77ffba26462054bc0

Observation 9088f5d0-dca2-45db-8211-2d57682a1602 · outbound

This paper cites A Survey on Future Frame Synthesis: Bridging Deterministic and Generative Approaches.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos A Survey on Future Frame Synthesis: Bridging Deterministic and Generative Approaches

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.700387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.700387Z digest=sha256:ca5b02ff60eaf35ff04b1e43f33a75cbffc254026678d9422f351be4a1ec40f7

Observation 4d50a6da-060d-485c-a502-b4c4407616d5 · outbound

This paper cites 4m: Massively multimodal masked modeling.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos 4m: Massively multimodal masked modeling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.682921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.707093Z digest=sha256:79ca6f68d3f4b8bc5cb79d7df4ec308b96d9a26f2ce349d0c97290eee7877668

Observation ec682287-68c2-4a2d-871c-3ea251d598b7 · outbound

This paper cites A review on deep learning techniques for video prediction.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos A review on deep learning techniques for video prediction

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.659454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.714701Z digest=sha256:18d547736c0ddce0dba6bc2eea67489c249dd48b87b89b69ff34b0feb8b925d7

Observation 0376c388-34c1-46f2-96e1-f57eb5b4fd79 · outbound

This paper cites Training lan- guage models to follow instructions with human feedback.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Training lan- guage models to follow instructions with human feedback

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.641719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.723354Z digest=sha256:83b751e3a1bd01122baba01acda85f38386a83cd1820b6052c08540251fbd203

Observation 9a21298a-3ac2-4870-82fe-0f3903b17733 · outbound

This paper cites Scalable diffusion models with transformers.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Scalable diffusion models with transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.731052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.731052Z digest=sha256:8948f7970b525e1ac353d4e1e6fe033e78e73bea9c56bda85538c35a1018c8df

Observation 387405a6-6961-4e37-93ee-cf01d3337e85 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Movie Gen: A Cast of Media Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.736838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.736838Z digest=sha256:232099ff77af277bd340e6fe1776ba38b15a5271ebb5e3afd8f80512841a816f

Observation f2d64baf-3c1d-4ac9-a43d-998b13696700 · outbound

This paper cites Improving language understanding by generative pre-training.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Improving language understanding by generative pre-training

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.742098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.742098Z digest=sha256:b03d91f258cc3ce8b2efcf236c451a92b30044e91fe106f44921d01ab0c40f97

Observation 86b41685-ca24-4e5a-8347-92e04007008a · outbound

This paper cites Zero-shot text-to-image generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Zero-shot text-to-image generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.749873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.749873Z digest=sha256:ad16247904a24ccc2e775425c1119308d2cd0d1e8e10513bf68ec6d47333e8e8

Observation 91fe2c47-21cc-4c6b-85d2-13b1e42c3cec · outbound

This paper cites Gen- erating diverse high-fidelity images with vq-vae-2.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Gen- erating diverse high-fidelity images with vq-vae-2

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.755863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.755863Z digest=sha256:76f54922dc48ecbdc19e328ad6ef1ac3f8d78bf0976b0f80c0869d5c1465df7a

Observation f833eb71-60a7-43a4-b0b7-1fb6dd4da5ba · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos High-resolution image syn- thesis with latent diffusion models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.761573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.761573Z digest=sha256:801214abbddefab6ce4bc9e6cce552a119b5addb5e517f47361ef7519c4fbce2

Observation 45c0fdc8-1700-4991-99ee-3e81c1b30df3 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.768454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.768454Z digest=sha256:ea9c778dc5d34c6143768eadd19b925cf0af639defb7a060b0068cad3bfc83b9

Observation 432f6f19-dbe9-4d43-9392-c7f7c167daa0 · outbound

This paper cites Emu: Generative pretraining in multimodality.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Emu: Generative pretraining in multimodality

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.556717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.774341Z digest=sha256:19939f110c7e4829499c1893435dc834412a06a484917e01f973e1a8911d9ad6

Observation 72a91f2a-6f3f-45f1-85ac-e556f1ffec2d · outbound

This paper cites Generative multimodal models are in-context learners.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Generative multimodal models are in-context learners

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.779648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.779648Z digest=sha256:e2cde38f95d0f5d620ff8871d55ffd8fcc0a954a90441557ee8506c7baccca63

Observation d3f95508-5abb-4f43-9241-174b7b59b373 · outbound

This paper cites HART: Efficient Visual Generation with Hybrid Autoregressive Transformer.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos HART: Efficient Visual Generation with Hybrid Autoregressive Transformer

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.786214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.786214Z digest=sha256:6132a7a1e0150b964c3bc7ed88c3d74f5a59a6fcc8151ace5678d47ef4740b72

Observation 06f97568-33f6-4f54-918c-bbe14c87f65b · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.795799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.795799Z digest=sha256:cb6654837ac3a3238ba56f97248739ab529959e8ffafa0e4ae95d464721f0ada

Observation 4ae9bebc-c9d0-4e2c-ad21-16b3724a10fd · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Raft: Recurrent all-pairs field transforms for optical flow

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.806232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.806232Z digest=sha256:54cc97d45fa62d5feb18c5412d78963f2f56af4aed162932c3736120fcbfc3ef

Observation 9f863bbd-ef95-49b8-a71c-3976eab1d981 · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.814236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.814236Z digest=sha256:a94bd33e11bb32ec7a1fda6a1ec22ec880392ea2fe24d5292b2b32f39f96db90

Observation a2ad54c0-dac9-4aa3-9ef8-89c5adbcd3f2 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos LLaMA: Open and Efficient Foundation Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.821530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.821530Z digest=sha256:93402de35cb4cc1ffacf969c1f1b22c1e8651ec6a30e8850f866f7bfc44db29f

Observation 3134ff7a-5747-4adb-9dbb-0cce5ab31776 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.828747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.828747Z digest=sha256:19b96e46d594dcea6b0125608cd7f6deb3fb5914ba76860041712c758b3b9753

Observation 0ad0e2a8-cb1d-4646-b66b-af36bc7e301b · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Diffusion Models Are Real-Time Game Engines

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.838508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.838508Z digest=sha256:4c138263b5ddf26fe391ac05699579031c855004a5e67316e5b71a9281868237

Observation ebb31694-1b4e-4ec6-98d8-dc36c9b489a5 · outbound

This paper cites Neural discrete representation learning.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Neural discrete representation learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.502833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.844350Z digest=sha256:55a58f7ef37eb7e6318aae530acb8132e939d1ab6b06bf7ad869c9844e3b7719

Observation 5dae021c-4081-4c80-a806-e87ef27af093 · outbound

This paper cites Attention is all you need.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Attention is all you need

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.482705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.850145Z digest=sha256:90e1a5999e1fe8ab626d20dc425fe6ae33227d35b92e190ee01f2ab90637122c

Observation b2524ef3-cb1c-41f2-ae34-6cfbcf6e7f82 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Phenaki: Variable length video generation from open domain textual descriptions

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.460023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.859288Z digest=sha256:c453ace64348503e3208c439d4e27eec2a9ec5f073d1f52afe2992eb850602c1

Observation f7ef387e-09e9-4216-a172-44d5f050b649 · outbound

This paper cites Images speak in images: A generalist painter for in-context visual learning.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Images speak in images: A generalist painter for in-context visual learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.433568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.870981Z digest=sha256:251dec70b841851ddd47701947f7438c6c1af00889f82307d7f734932aa2ace0

Observation bbd0dc0f-4b8a-4ba5-a6eb-025be1336b39 · outbound

This paper cites Seggpt: Towards seg- menting everything in context.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Seggpt: Towards seg- menting everything in context

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.411342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.880052Z digest=sha256:11b1b8a504d0b2079b1c290312b87432741f8540dfd31f2f980d48c0a57c0251

Observation bbe938da-b11d-4d50-9499-f02f52d11f18 · outbound

This paper cites DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.887583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.887583Z digest=sha256:15026869a0e21144fccf282b0f8da246d3e32a9261b98d909d2153d3d3eadb59

Observation b25c7dc0-b0de-4c3f-86ac-43360c96536d · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Emu3: Next-Token Prediction is All You Need

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.894742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.894742Z digest=sha256:4893f48e429b095ca0b1c497307f9a2c80c40199ff08f79e30b006e0aab0cafe

Observation 93b6d96f-b5d6-4e76-836a-478f3bef1e7e · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.903306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.903306Z digest=sha256:4dca60c6103378dbdf24203e12860b8ddf6511d0b9203a13e21c9979a40b9c6c

Observation c1f1047c-6953-4c04-a97b-2443903b5a73 · outbound

This paper cites Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.390480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.910993Z digest=sha256:b27a5c5942bb7b3827ffa04b0f9b82aff40ad9ee8e780403a4a224c49ce9036a

Observation 0353e3cb-5f5d-43f8-800c-3a15b953a808 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.916995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.916995Z digest=sha256:829f685d85fcb562ae404b75f810e50188cde3e3c3a50aebe5bf5da60358f015

Observation f3383fad-0396-41ba-885f-78bf3f574bc7 · outbound

This paper cites Loong: Generating Minute-level Long Videos with Autoregressive Language Models.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.923496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.923496Z digest=sha256:a2fab8460cf334e4168f19dc75190dbbb791f76f130166bebe5340552f75737d

Observation d5edc8c8-6d9d-4a9a-bb1d-5a4912036a76 · outbound

This paper cites Panacea: Panoramic and controllable video generation for autonomous driving.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Panacea: Panoramic and controllable video generation for autonomous driving

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.369075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.930253Z digest=sha256:9e77289478179cfe704e3d66042551ba8641052e97b8d98becac8aea277f8530

Observation 01d54d45-0ad7-40c0-ae53-ce0738ec8867 · outbound

This paper cites Motionrnn: A flexible model for video prediction with spacetime-varying motions.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Motionrnn: A flexible model for video prediction with spacetime-varying motions

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.351544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.937202Z digest=sha256:f17714902b6cebcb4718099a6fb92899b41183d20d01d4d5c03e433740a52612

Observation 474787cb-7397-45d6-8fb9-952af983f087 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.333880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.944964Z digest=sha256:663916ed20614618182bb2beac769b00428e1b96ac23ecdfbbfbdc1fd9469e3e

Observation 8785395c-1dd5-455d-859f-91a2c49d1ece · outbound

This paper cites Simmim: A simple framework for masked image modeling.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Simmim: A simple framework for masked image modeling

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.312868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.955613Z digest=sha256:9d5ff18e264059bb6a6918ce1ed05242070a528f7f685f7e41a7203fe5f1bb08

Observation cb7a6224-dec5-4a3c-a09c-fdc8f3267fa7 · outbound

This paper cites A survey on video diffusion models.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos A survey on video diffusion models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.961351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.961351Z digest=sha256:8cb2f07c66721a2ec37e7b1050bac6cf3843c8b5b8fdb42391fa687d38952695

Observation 1f542c90-edc6-4cd2-a530-3631946b29a8 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.966047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.966047Z digest=sha256:94c7622d7f815a05b80d4b08e121f263b166f8c92257d41972e06ab7d9c82414

Observation 3fd68668-8fc2-4a1a-8249-076c5c3ac68a · outbound

This paper cites Generalized predictive model for autonomous driving.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Generalized predictive model for autonomous driving

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.283229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.971281Z digest=sha256:34dfd87a813b7f03c7c2cedbf397ceabe0dfcaa1d4cca5a737f4f9153086c0bb

Observation 3fa20282-9e5c-49e4-a080-8d698d887d38 · outbound

This paper cites ZeroSmooth: Training-free Diffuser Adaptation for High Frame Rate Video Generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos ZeroSmooth: Training-free Diffuser Adaptation for High Frame Rate Video Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.975570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.975570Z digest=sha256:d6c61419b5525f88b8a3004331938c99c22e8065f73a1d84f5604f7c745f0138

Observation 93e4470a-d5db-42e7-b682-40b085b4382f · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.981361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.981361Z digest=sha256:469e8cd4158674ab56acb223205e28b7cb9a53d084c62b1464104897f2d85492

Observation 2681ddb4-6280-4c1c-a05c-566d7c7f3e4e · outbound

This paper cites CAR: Controllable Autoregressive Modeling for Visual Generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos CAR: Controllable Autoregressive Modeling for Visual Generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.986030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.986030Z digest=sha256:a47307b46722f78b5c2f6a5dfff94fa8fc56aa40ecc96fdda45ce5bc05fd4596

Observation 99935dfd-3834-4960-97ab-c6efc4be1713 · outbound

This paper cites Bdd100k: A diverse driving dataset for heterogeneous multitask learning.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Bdd100k: A diverse driving dataset for heterogeneous multitask learning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.261141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.991086Z digest=sha256:fd7038ed68bbba356e5cd068433cffe78da1189a0bbbaec20f6d56123026d492

Observation a66a9b79-569b-4714-b037-b236bf54604b · outbound

This paper cites Magvit: Masked generative video transformer.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Magvit: Masked generative video transformer

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.241365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:52.996962Z digest=sha256:ccbbb8dafb6477e180b1db6140cc0a9fd6bf2a5e7b4d5861077832f707f0c42b

Observation e78a231d-688f-4d03-b621-7172f3bf7ca3 · outbound

This paper cites Language model beats diffusion - tokenizer is key to visual generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Language model beats diffusion - tokenizer is key to visual generation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.223917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:53.002593Z digest=sha256:30203d64b72aa6d217e778e06e0a8b29f11741160b5a986ff77a28abfcf32b85

Observation 0853c036-e229-4b1d-ba9e-698b9183813e · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos An image is worth 32 tokens for reconstruction and generation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:53.009061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:53.009061Z digest=sha256:98529cca45774ed25dc182fa1489e5a6d6767529f9b2e5fbc51804ce5298d932

Observation 0a71ca33-a143-4d94-8ea9-a019105aec49 · outbound

This paper cites DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:53.030836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:53.030836Z digest=sha256:a45573300157d2a8d5b0b2044a1f3920f0734b1db81f723f0ee8f44b4cb2100b

Observation cc635933-a5f7-40e7-afa8-d2599f9f151d · outbound

This paper cites Image and Video Tokenization with Binary Spherical Quantization.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Image and Video Tokenization with Binary Spherical Quantization

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:53.036778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:53.036778Z digest=sha256:ba0d42e08c68d46c40a12654c5b3c619ae21879269bc871fa822d7691d143399

Observation 49cce8df-535f-4631-b10f-2d6b423952ab · outbound

This paper cites Movq: Modulating quantized vectors for high- fidelity image generation.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Movq: Modulating quantized vectors for high- fidelity image generation

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.192254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:53.043221Z digest=sha256:7c66d08336f6b4d6d17cd3f2ad6a6e5e8b0b94f30f3522685ad9222e7d998e52

Observation f4beb73a-b1fd-4991-adc4-d3ec2691415f · outbound

This paper cites Crossnet: An end-to-end reference-based super reso- lution network using cross-scale warping.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Crossnet: An end-to-end reference-based super reso- lution network using cross-scale warping

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.165996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:53.051276Z digest=sha256:59b2e06eaf6678362a3491f27b523e87a3f2824e982a100584ddacccfcea7173

Observation 31913af0-2f91-4319-9aca-d64d0373afba · outbound

This paper cites Unicode: Learning a unified codebook for multimodal large language models.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos Unicode: Learning a unified codebook for multimodal large language models

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:11:54.134840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:11:53.056934Z digest=sha256:34ee01f33bd53471db1b0187fa18a60e19c54dc77c2fbf3a59dc5b88c2cf91c5

Pith citing papers

Observation 72d6ebc0-f803-40cb-830e-3e6e6ebe8195 · inbound

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction cites this paper.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction ARCON: Advancing Auto-Regressive Continuation for Driving Videos

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.372483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:2f5b9eb565d6fdfeb6d3e0fc1090876cd7a336185070db19ac308b0b404c4e3a