Pith. sign in

Paper Citation Record · LEDGER

ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2407.06135.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.06135 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:12.206435Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:29:56.712695Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b69cde91-234c-493b-834b-750a5bf621dd · inbound

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation cites this paper.

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:10:54.606202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:10:54.606202Z digest=sha256:3cde3751d18306c0f0d17fa6038c9bb57a120ee885af716f81f0cb5de0628b71

Observation 28b5434d-7e3b-4939-978d-a6dfd6c6ab61 · inbound

Interleaved-Modal Chain-of-Thought cites this paper.

Interleaved-Modal Chain-of-Thought ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:38.178321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:38.178321Z digest=sha256:3c0ac0af5178836f30fb7bccdff3bff59b469a62225b39cfb75c4d64c1250ade

Observation f4a5b56f-d725-4dd4-9c01-f3cc7ffd1e86 · inbound

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads cites this paper.

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:22.612319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:22.612319Z digest=sha256:e6aa10e9f01bb48048abf3e6c9af06ad5ede312b8a03bed18337f36af4ae291a

Observation 1de2283a-021d-4d3e-84ee-d413f1482a7a · inbound

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers cites this paper.

ChatDiT: A Training-Free Baseline for Task-Agnostic Free-Form Chatting with Diffusion Transformers ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:01:05.400479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:01:05.400479Z digest=sha256:86b686b65b3be7293f2bc45f20df7b9177330709d48c39e11fdf1bbdfc168a10

Observation 49ed5c40-5b9c-4ee4-bb16-d928b426a4d8 · inbound

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants cites this paper.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.851646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.851646Z digest=sha256:47636bd77cc14331f08313caabb21cb9021dc9e26a8f21d80a5a7dd792e454fd

Observation d6d10d20-c643-4d77-9942-832649580bad · inbound

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback cites this paper.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:12.980482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:12.980482Z digest=sha256:04fdc80c77a48f9c085bf68b8d2bcaca2ce655289df54c2c6e0219a1ceef3c27

Observation d549e1aa-5df8-4520-9333-471a6814f482 · inbound

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought cites this paper.

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:09:34.853178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T23:09:34.805552Z digest=sha256:4a50cf2b6a47e5cddd7821b64251f623b9b9d9042d0c3774537a8946d0d4e8d6

Observation 6a6594c2-6bc9-45b5-ba54-ba45a3b99432 · inbound

Parameter-Efficient Fine-Tuning for Foundation Models cites this paper.

Parameter-Efficient Fine-Tuning for Foundation Models ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 268

Resolution
unresolved
no resolver link, observed 2026-08-10T15:38:03.684324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:38:03.684324Z digest=sha256:438aaa1515be286b0171e81e8ad02fd13e9a5b6466345412ca333b7cc218827a

Observation 2c6f69ac-a1d0-44e7-8f5a-a00b949aa260 · inbound

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding cites this paper.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:31.797843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:31.797843Z digest=sha256:537578daa45928bb991dd13934354c78562cfd469a6d8171e260a411b76f02b2

Observation 3f978df0-4e67-4532-8c78-34642765db2e · inbound

LANTERN++: Enhancing Relaxed Speculative Decoding with Static Tree Drafting for Visual Auto-regressive Models cites this paper.

LANTERN++: Enhancing Relaxed Speculative Decoding with Static Tree Drafting for Visual Auto-regressive Models ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T15:46:21.995596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:46:21.995596Z digest=sha256:29e3d03916baf3c39e1605eef83c66ef652af9b3be3de5b5237daad04d780097

Observation 58d5da58-a9ab-4936-a875-026441128f46 · inbound

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation cites this paper.

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:30.373363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:21:30.373363Z digest=sha256:017f763c69cb82a63cbc2efa2eb926e72b4b533bf036d23494f6ad457eb8c218

Observation fad3175f-07c1-4792-bee1-036d26231f4b · inbound

MuDoC: An Interactive Multimodal Document-grounded Conversational AI System cites this paper.

MuDoC: An Interactive Multimodal Document-grounded Conversational AI System ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T20:20:13.163709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:20:13.163709Z digest=sha256:46c580e6a6636eebbfa91d4eccbcb4d1a7ace2c3332b9fc104a527b28632442d

Observation bc66f83b-0978-4db4-8cbd-fa239220af94 · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.206435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.206435Z digest=sha256:c9c6d6c9a8a0a4a916fe7b4e7de6dd170c4211c2b64607226b3d0752cdd877e6

Observation a5b652ff-d965-4adf-988f-0bda4ea9a5a2 · inbound

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation cites this paper.

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:36.354700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:36.354700Z digest=sha256:3c088169bcdda1be7847e902035399b1cb8a22471cc805d5c1fdbc8cce3ed894

Observation 2ac5ac4d-1524-4b64-9514-e256ac64affa · inbound

StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation cites this paper.

StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:46.064937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:46.064937Z digest=sha256:fd35db00c76746190a07d4ec74f4701bf8c5e2fbc53be946dd611cacc6b2ce52

Observation e4b8cd6b-a9a1-4815-9017-881ab97394c8 · inbound

Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists? cites this paper.

Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists? ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:36:37.945412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:36:37.945412Z digest=sha256:19cd70e00197ba512cc380c8cb987f1a4d18260b4f70bbc845a0a08b44131749

Observation e3caf459-b8b1-403e-a755-12a927bd2ff1 · inbound

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics cites this paper.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.262152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.262152Z digest=sha256:174fc438deb429a58f00be50eec4a6845852c6e14aa5481e0d7af83b4c75f9e5

Observation 5b1cc51a-8109-4340-8a18-1e8590487aa4 · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.618223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.618223Z digest=sha256:dfadb51aad1e7ebe458c99aab6d414adb0852b5d9498089a0358593bf6532d69

Observation d9b48a2b-da9c-441b-806f-bbcc140f388d · inbound

Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models cites this paper.

Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:32:34.455945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:32:34.455945Z digest=sha256:cd9790e0bb2629cd19458503551a0fd90761f34b94c49cb3670d7903145dd303

Observation b8a06caf-109b-4c80-9cdf-df37b84d8c32 · inbound

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning cites this paper.

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T23:02:52.704228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T23:02:14.913189Z digest=sha256:a325fd2ed753f0bf6c1f5904e03034d88fec56f137b11ae4fbeae5cccaf83114

Observation ac75b562-6416-467f-9d6e-b42ac1b682c3 · inbound

VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping cites this paper.

VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:35:17.562297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T21:34:14.371914Z digest=sha256:d91fd987cf92992a4d2d44829f4a7c3c2d5dfabc1ac9381473c6be4c44e7a61d

Observation bce8e8b5-a923-4bc2-981c-763879253328 · inbound

Mull-Tokens: Modality-Agnostic Latent Thinking cites this paper.

Mull-Tokens: Modality-Agnostic Latent Thinking ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:58:38.638012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T22:57:04.802871Z digest=sha256:fc36b054a5721c7cac9fd185879d868c860e2ad1f8244bb7edaf8ece208de049

Observation bc8e3f01-2f13-4856-871e-90ddc180c11f · inbound

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning cites this paper.

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:37:59.986665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T14:37:05.402850Z digest=sha256:d1ec42cb3cf3de6c294d87a997806868d66677a4a1a8fefc5d420e6ee3418747

Observation 171a5348-85c7-4a3a-ae57-0299994caaad · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:15.601866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:15.601866Z digest=sha256:f3e01b521f9e92a7e722282d3b7fdd703bc06974de0e0e8ac3f2e15df4085452

Observation 0eab7f02-e4e2-47e3-9b72-36d1e19959e2 · inbound

SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation cites this paper.

SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T22:30:39.049531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:30:39.049531Z digest=sha256:6d7a9425d9766f67abc132551e3d18e3e160a7be1192d0f61262f1f86ff39076

Observation 15e0d7f2-b3f3-455d-a874-08855ac0ac8c · inbound

CASCADE: Context-Aware Relaxation for Speculative Image Decoding cites this paper.

CASCADE: Context-Aware Relaxation for Speculative Image Decoding ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:55.103143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T02:08:27.374066Z digest=sha256:5b868c3fbd315af759d5b8b637ed44dd0eb00b46f7983b1e20a73d19cd4a7484

Observation 9f5be935-48d2-433d-8137-702321f38ce5 · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.811647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:e8594e977f6aa2db7765457e87fd07b4faba26de7c90c0875c6d1c38c616b6ea

Observation 7cb32506-9195-49ad-b539-69d0342741ec · inbound

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both cites this paper.

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:14:53.499825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T03:09:58.411261Z digest=sha256:874d98e623242d8bf8deeed6adb9bf43aaee2e34c75734148611daf41d6739ca

Observation 96ae87d1-6052-41fc-82c2-a29f92ff9eba · inbound

ETCHR: Editing To Clarify and Harness Reasoning cites this paper.

ETCHR: Editing To Clarify and Harness Reasoning ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:25:19.378172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T04:21:50.862780Z digest=sha256:40486720c30531b38dd73962a18c2752e8f4ef68b4ed4e5d6bd76d884cb72545

Observation 7a85336c-e2ca-4c2c-bd00-c87b7ded61f7 · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:52:36.081851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:5be08971a5b6473c85415ec321c5d45d27621bc67f832d2bc0902bad5baa9dcd

Observation bfe18290-46dd-4369-b218-fcd1ebcb5253 · inbound

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation cites this paper.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.571000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:3208b9a75c22e9f65a6aecb18b31568ff073b019b81436bac670cfe5456029af

Observation 079e8fc5-6c9e-4622-b814-4b922cf8c71e · inbound

Parallel Jacobi Decoding for Fast Autoregressive Image Generation cites this paper.

Parallel Jacobi Decoding for Fast Autoregressive Image Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:26:57.122683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T02:05:58.714029Z digest=sha256:73836f94e84c4b99e54980ff01635822980864c3496039153c06969f89549a9a

Observation c973eaab-4ad4-424d-89a3-fea897dc6c04 · inbound

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning cites this paper.

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.953878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T09:59:02.899488Z digest=sha256:430650aa693b7bd09869de5b60f2ca8effe455fae277029ef0ad48ee52ab1653

Observation 54943f0c-8254-425d-832c-86245aee7d58 · inbound

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement cites this paper.

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:38:19.352383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T07:43:09.538833Z digest=sha256:d3f108095f1b98a35bb78ce7a8fbea3bb9782cb8f7abd95f167d2fdf4ac7d636

Observation c0d3e570-86ff-4ddb-9fba-afdfc2198749 · inbound

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement cites this paper.

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T14:14:01.413340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:14:01.413340Z digest=sha256:eaa477a3456afd2cccb838e59155a8fcfb1a0d47fe305d3b625118a5f5dad834

Observation 33f4139f-20bd-4da1-8c6c-4c4b48745748 · inbound

NavWM: A Unified Navigation World Model for Foresight-Driven Planning cites this paper.

NavWM: A Unified Navigation World Model for Foresight-Driven Planning ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:56.714946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T00:42:30.237545Z digest=sha256:4f483cd4133a37b0f3a02fdfe75731a6dba3cd1dd187cdc268a253624618d3d6

Observation aed47806-a2fe-43b5-a94e-1cebd9dee218 · inbound

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation cites this paper.

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:04:20.971346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T06:04:07.327934Z digest=sha256:44737431cd0b0c59ef1e5387e4419e410ca5631bfdfbfde99a05db504b375a86

Observation dfeaea4e-7ee4-4034-859c-520013f7dba4 · inbound

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space cites this paper.

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T06:13:16.894125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:13:16.894125Z digest=sha256:0179d0a974680fe506c058e7d446dedf163b8adb05db180f7a390454fa3e88e5

Observation 41a213a5-11cd-4b02-b507-9ce1f963f5ab · inbound

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics cites this paper.

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T21:36:42.378149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T21:36:42.378149Z digest=sha256:7cce7d1b107618724e98944344d814e631bba4a5ffae4e440725aeb56fa773b4

Observation f0351427-741e-4ed7-a6ee-07d7e6805aad · inbound

Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation cites this paper.

Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T10:59:43.137854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:59:43.137854Z digest=sha256:3148d2e84781cfe64a8c6a2c1359172f9d9c66fdd5581717a15983174eac0f99

Observation 114e8193-2dd6-4389-9d07-a0d85a9da3be · inbound

VIG-RL: Learning to Search and Insert for Verified Image Grounding cites this paper.

VIG-RL: Learning to Search and Insert for Verified Image Grounding ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T19:16:58.978808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T19:16:58.978808Z digest=sha256:34865ecc0177bdb08bd7d0d60b4cf839aa15d40aa8946d4795e45269ae0cd26f