Pith. sign in

Paper Citation Record · LEDGER

Loong: Generating Minute-level Long Videos with Autoregressive Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2410.02757.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.02757 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:03:20.688743Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:59:42.133975Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 43ecb43a-4735-4712-b75f-d7a2087f9dc3 · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:07:39.816112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:5ad37abaeb7ef1446434fa79c29cb8a818fd90356081f44fea7c737784198aee

Observation 234ef67b-9fed-4f69-b73b-0b10d0e4e238 · inbound

Long-Context State-Space Video World Models cites this paper.

Long-Context State-Space Video World Models Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:20.688743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:20.688743Z digest=sha256:a7e28618347910f8b8aeff4e6d69f64ac7eacc37746785489ecc44541e5ad374

Observation 65cb59c6-2c2e-4798-9077-12b5e9ee947f · inbound

Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots cites this paper.

Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:17.046537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:00:17.046537Z digest=sha256:88dfc8bc52fc991a546589db046c3fc1e8dcf2317da9fbb37316fc9940ce73c1

Observation ecf31cbd-af7b-4861-b57c-eee64139f826 · inbound

Video World Models with Long-term Spatial Memory cites this paper.

Video World Models with Long-term Spatial Memory Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:41.605178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:41.605178Z digest=sha256:5804fcdaa1376d1edef0d097993f00e8adb5142d342502bf0361f739efafbc0b

Observation 4dad1b58-0de1-4a3d-a9d1-2327fcd5b564 · inbound

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion cites this paper.

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:36:53.131734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:36:53.029590Z digest=sha256:0a908843d2046bd1333553c6b61e0211e93329b716693d7bb5d67009e5db448a

Observation 7e651edd-6035-493a-89c6-4aad685075c8 · inbound

SpectralAR: Spectral Autoregressive Visual Generation cites this paper.

SpectralAR: Spectral Autoregressive Visual Generation Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:51.201816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:51.201816Z digest=sha256:6f4001d635896e322478fff6c5528d00f4f6414813b2726bc4b99a2472fb42ca

Observation f9215a7e-f11a-4acd-a52a-ec5a86ff0863 · inbound

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions cites this paper.

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:58.539477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:58.539477Z digest=sha256:bf99404ee1e66376405ac9c1322ac01db3937f5e75050bac4b4bb9833c03938b

Observation 7f265ae3-e3f9-4289-ab4c-441826bbe344 · inbound

VideoMAR: Autoregressive Video Generatio with Continuous Tokens cites this paper.

VideoMAR: Autoregressive Video Generatio with Continuous Tokens Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:28.919613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:28.919613Z digest=sha256:b5119484acb0f5f4a743ecd594ad26478e667cd9ee1b8291df7983d0b132aa14

Observation 08562ac5-5180-4707-9435-0fbf8578180c · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:29.715057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:29.715057Z digest=sha256:013175ef781427ac0a3ab149cd15a8542496119d6c4ef26e561d752c887a6336

Observation 7f4c72cf-7aea-430b-bf5e-81a1e9075b20 · inbound

TokensGen: Harnessing Condensed Tokens for Long Video Generation cites this paper.

TokensGen: Harnessing Condensed Tokens for Long Video Generation Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:30.643014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:30.643014Z digest=sha256:f9c97631af470a5317260034683424272c466e6764016f7078e214c4dd1d81e1

Observation 18149d3f-4904-48fb-97af-b879614ffba6 · inbound

Enhancing Scene Transition Awareness in Video Generation via Post-Training cites this paper.

Enhancing Scene Transition Awareness in Video Generation via Post-Training Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:03.279456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:03.279456Z digest=sha256:c787747eb012987eafb7a517ad9687cea9f3bc74cd0baa012d49ccc10a059773

Observation e17afe89-c307-4c59-a8a1-eae9de4956e7 · inbound

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics cites this paper.

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:49:52.758882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:49:52.758882Z digest=sha256:0a273b22c72e6ca5399ae56fefcb761e3c6790566c2851003b80124318ac8827

Observation 05edc1df-c0f8-4a7c-bebf-b1e5fa8eb76a · inbound

Rolling Forcing: Autoregressive Long Video Diffusion in Real Time cites this paper.

Rolling Forcing: Autoregressive Long Video Diffusion in Real Time Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 97

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:15:29.226122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T11:15:29.102090Z digest=sha256:9a7a784067218a7c06a9e2135766babc1837192c529c07d8ccc8d263aad8b469

Observation 4965dc40-66e5-46c6-a47e-84b38ec35294 · inbound

RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling cites this paper.

RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:15:54.297371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:13:42.934115Z digest=sha256:67dd5707739e472b99ba80071ce4728a32898bea855b748e7d4b27a3e1be2a2a

Observation 6b987e18-a487-43ca-a8c6-39df490ec4bb · inbound

Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation cites this paper.

Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:05.018987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T05:14:36.280155Z digest=sha256:a0492f16e03f37b9c03e28750ef60b75debaa576881bb08726f6b787ebabbd48

Observation ff0dd921-4190-4d2d-bbac-46e2f790ad5c · inbound

BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation cites this paper.

BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T19:41:58.807489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:41:58.807489Z digest=sha256:55ba073570c91e54110ce0aae89a4b8152383b6337bc2be9a66bcc016713e999

Observation ad757ea3-19b1-4a42-b533-57ba00fbb443 · inbound

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion cites this paper.

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:07:29.831479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T07:02:38.876518Z digest=sha256:4bf0fbb7abf93d3746496ca113bf0b3b4fb688ac8c9a45bb639ad6a3658d1938

Observation 4d353ac6-2594-423e-b748-f3962774c70e · inbound

GeoWorld: Geometric World Models cites this paper.

GeoWorld: Geometric World Models Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:40:03.337727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:39:15.308355Z digest=sha256:ffb4aa7dce03dc5f8e367a92e11c63d7e6c83b70acae9225c8f5170840af1045

Observation db661cb8-0b07-4279-8e00-1d6053650965 · inbound

EduVQA: Towards Concept-Aware Assessment of Educational AI-Generated Videos cites this paper.

EduVQA: Towards Concept-Aware Assessment of Educational AI-Generated Videos Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:54:09.103453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:52:09.262497Z digest=sha256:ef0f047bf228acddd2789a45c06be7a28ab4b954f057222cbb5ad89df38593ec

Observation a5275a79-fece-4cb6-871f-5115c91336d4 · inbound

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms cites this paper.

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:38:36.372424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T01:35:14.878069Z digest=sha256:494c68837bc418d9599b41a82b80da78cf28b5407352b5f8740b68faf5baa040

Observation 60e5bb63-2705-4696-9a2a-b466ed02e487 · inbound

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling cites this paper.

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:25:53.605274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:08:56.588282Z digest=sha256:e3cadada850ff52f31d3900f084e018d4e1f3d97f2673d59b96f3e6fc9267b8a

Observation abc3491d-9480-4f38-b430-594dc4f6170b · inbound

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation cites this paper.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.905365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:8f6f8a6931dc926363c4ea4f561f0f1b7ac3d7a25be5ed5df0794d8537ca4af4

Observation d4980422-aaba-423f-8a2a-c00d56f982f0 · inbound

Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity cites this paper.

Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:33:32.264156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:31:48.593354Z digest=sha256:43e19393363acdf5b765fcbbd5654aaa8f4b9df322d3a5bcdb32ccfa0a44ce93

Observation 007b082c-3c63-4872-99d7-ab14f035742e · inbound

OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation cites this paper.

OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.485825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T07:56:40.001739Z digest=sha256:8901284a9c7392b62716fcab02cbfda0e1fccfdf687e13ffb84ada89147bc460

Observation a6dc219e-b016-4691-93c1-1586790ca6a8 · inbound

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models cites this paper.

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:00.543053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:56:21.783415Z digest=sha256:510407363b059f850e1d95cad9f1f33adf7d76110c08ab5af18025a02086899f

Observation e517b4e0-dba0-4f18-a000-163c857f7277 · inbound

Streaming Video Generation with Streaming Force Control cites this paper.

Streaming Video Generation with Streaming Force Control Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.425890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:14:32.465663Z digest=sha256:d4aaec02841850e319620380ca43a0d469223d9e5bbdf7a15c2064f8243006b4

Observation 4339dd7b-7c9d-463d-b2de-e6ea4fb0fd9c · inbound

Towards Error-Free Long Video Generation cites this paper.

Towards Error-Free Long Video Generation Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:59:42.135414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T10:49:36.838492Z digest=sha256:2144c560a996ba0de58516b98a20334e5ccd01640ac731a3e03b66fdaf20c4b3

Observation 7624b5da-d2a8-474a-b6c4-83e8f3175c03 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:40.345090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:dd6aaae1d24bf27bb6ffc8dc2fc5027a4fecf13da6ee07dbb5d2aff8b835e098

Observation d0ac97ba-c813-4b6a-9385-ec4ecb60ab98 · inbound

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model cites this paper.

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T01:59:48.820112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:59:48.820112Z digest=sha256:119daef130792e9cd40fb1aabf5799b474588ee0de298f8bc3e1799ab968bb31