Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:23:12.587345Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2506.05501.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:23:12.587345Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T05:12:44.268199Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T08:46:26.764949Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 96a77d7f-fa98-41a8-9e07-edf007569884 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Qwen Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 200d7dc9-908e-44b5-9d48-9045e77bda97 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa50303b-20ca-4208-8b1d-4821dec5f712 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa006ef1-21b6-45e0-857e-fbe5e29c86db · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b7d0096-7748-4a23-9d69-4ac1d3a6712c · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ef8ac61-7805-45dd-938f-6a7e989dbefb · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8078d485-d69e-4472-96ea-3820844c2e71 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb7208e1-e5af-4525-994d-e3bacfddec81 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38f3067-beb7-4847-bf6f-4fa8939bcc51 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2ff5d2d-5d0c-45a8-b643-68d21febdb40 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3487a38-dc62-4b86-bca4-cf93de2dbdb1 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Making LLaMA SEE and Draw with SEED Tokenizer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43cabedf-93fb-4bc3-95f8-3054e75a5a97 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 729a06c6-ce8a-4ff9-b2b8-40ac076a70a3 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 335c3dbc-5520-48c4-9bd2-b300b359f201 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96b75dea-f716-4ad6-8b6c-2b525eb43572 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2855e21-d9a8-4b73-972d-110d7f90ec7f · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b80772bd-6c4d-4a8a-a5b8-b41a05f3c324 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3300954-a378-45ed-bfd4-cba46133dae8 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38bfa069-c6de-4947-9281-2cf7158039a3 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac842250-4125-431c-9591-e5c9565c852e · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2702fb33-86bf-49bd-aac5-964a85abd772 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL LLaVA-OneVision: Easy Visual Task Transfer
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5d08ef2-8fb8-46da-aa02-8e3f1342f3db · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 103a01cb-0bb3-487f-acbc-489b357ee2d8 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0172b25-e168-43c2-a262-98253bd749ec · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b20debb-3f5e-4a67-b117-b03d9d4b5172 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b68318a-909b-42d0-915e-5c207f2ee8c7 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c5eef34-f8f8-4ee5-a798-8c6b7248206b · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ffdc315-6711-481d-8abc-328255e8f321 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d87d19a-e1ae-4303-9f62-5672c95aadb0 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7332d79e-57f8-4174-85d8-8b70c25bb5a2 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b2e2319-32ed-461f-b9e9-6439c909aa14 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2774d770-059e-47a3-b1ad-0b65c07fcf89 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba6c5649-32ce-4b33-8167-eb10939e377a · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dca5abc9-c343-408a-8b5e-7b94d32c7641 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Auto-Encoding Morph-Tokens for Multimodal LLM
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 847cb11b-3ecf-415b-86f8-11836b42ae8c · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c139af4a-b154-427e-9360-3ae217bb71dd · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5ab22ed-e6c7-4935-acc6-2f424e9e5885 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80576249-4419-4000-b45b-25bfa0c04620 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0eaa143-c41d-4162-9181-63929e916336 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Exploring Bias in over 100 Text-to-Image Generative Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4577579a-4ee3-4c6b-a3c7-d3abc8a12f2a · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Emu3: Next-Token Prediction is All You Need
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f51ef8b-170d-4fd7-9587-a084cd6aca56 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06cef3f4-e53e-4366-a48a-ab0590a4e052 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc0bb740-0c9c-45ab-97af-f5ac02e04aae · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 002e6126-6d0f-4cb9-9fd6-a1e4c5f649ed · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5bfb895-cb14-434c-94c6-f3ca527b0271 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b21326d8-46de-41c1-a09d-972fc86c1262 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 806798e9-e5f9-4bf7-bf90-b96e1a959eaa · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2de3076-eb6e-40b3-a718-a37c94f2152c · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1cca2d8-79c8-4148-a7e6-e175d8903d09 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL online" 'onlinestring :=
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92668bde-dfec-4198-a361-bb225f2e4309 · outbound
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL write newline
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c012480a-14d9-4fdf-8d47-e429a5d16559 · inbound
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 334b2604-7141-46d0-96bc-011d1b7f2b03 · inbound
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.