Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:56:59.808319Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 4 inbound Pith citation observations for arXiv:2412.01824.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:56:59.808319Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:30:21.554689Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T15:19:51.779596Z
93 of 93 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5bfa5040-9f0b-4ece-bcaa-65e0cd1f737a · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models PaLM 2 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29acbec3-e92a-49f9-b0cc-fa34529d2034 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6312d3f-0ed7-427b-b6ab-345fe88efaa6 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Es- timating and exploiting the aleatoric uncertainty in surface normal estimation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9911dee2-13bb-4aae-bc9c-2aa0d6548acf · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Sequential modeling enables scalable learn- ing for large vision models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4546b04-90ea-46b6-9bb2-ef7a93e0f190 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Visual prompting via image inpaint- ing
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7977445-e071-441d-9a4b-87c27b80b4d4 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Improving image generation with better captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 255aa901-836c-4acc-a5bb-1f5093ac596b · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models In- structpix2pix: Learning to follow image editing instructions
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 554dec18-e5ac-483d-a105-35f235895d7e · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8204d8d1-e1bd-460a-9f34-aeecde2685fd · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Learning photographic global tonal adjustment with a database of input / output image pairs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eaec335-1ec1-4f0c-8bd6-1c0d0a55ee68 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Emerg- ing properties in self-supervised vision transformers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2138bdf9-57e4-4c73-a5cc-b739ccfd801a · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Masked-attention mask 9 transformer for universal image segmentation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95668963-d6b0-426b-a6e4-f0e74ed8ca31 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Shazeer, Vinodkumar Prab- hakaran, Emily Reif, Nan Du, Benton C
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b10193d-ce38-4e8b-bc29-bc3a35456da4 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Instruc- tir: High-quality image restoration following human instruc- tions
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af2b56f9-ff1f-4a6a-9cd2-5d273410e54b · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models DreamLLM: Synergistic Multimodal Comprehension and Creation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 047321ff-fb17-471c-b628-6501bf540668 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Internlm-xcomposer2: Mastering free-form text- image composition and comprehension in vision-language large model, 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb85be1a-ee18-42be-b289-08f4d43408ad · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55972e13-63d9-49c8-8a72-2fc5aebd8ef9 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Taming transformers for high-resolution image synthesis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c4beb1c-4411-43cb-82ef-13fa678862bb · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Scaling recti- fied flow transformers for high-resolution image synthesis
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ccd74bd-8563-46fa-abff-8de2b8706926 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Removing rain from single images via a deep detail network
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f241de02-5408-4dac-8cb2-69f89fa3e553 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a352133-5948-4b17-8b7c-4bd3dd3493a4 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Making LLaMA SEE and Draw with SEED Tokenizer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e58e58aa-8a1c-4340-80b9-ebb31f09af18 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b37ea80f-5399-499b-af47-4730d32667f6 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Instructdiffusion: A generalist modeling inter- face for vision tasks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1a13b393-9199-4a9e-9339-b6194817a20d · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Geneval: An object-focused framework for evaluating text- to-image alignment
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 15daa145-782f-4ccc-912b-ce5b9d7e1597 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23aa687-ef92-47a1-a152-162285a5093a · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Denoising dif- fusion probabilistic models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bedcf773-588e-41fa-9e76-dad461967606 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Training Compute-Optimal Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 874e5f9d-2eff-4ffd-9b1c-7bc2eaa1a9ea · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 199e7cce-eab5-458b-8b02-b0795708f5bb · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Smartedit: Exploring complex instruction-based image editing with multimodal large lan- guage models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f26e6c39-a4f8-450d-bcb3-123e9c2a254c · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Repurpos- 10 ing diffusion-based image generators for monocular depth estimation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c22d2bdd-3c91-45be-94e4-0ff28e766f5b · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Auto-Encoding Variational Bayes
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ecaca46-4ca0-44d1-926d-f17e32371c2f · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Retrieval-augmented generation for knowledge-intensive nlp tasks
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0aa70367-b215-4d66-907c-a89e72fe9042 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models All-in-one image restoration for unknown corruption
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 85f473ed-78fd-4ebd-89e7-ef3849560b0b · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Mask dino: Towards a unified transformer-based framework for object detection and segmentation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9e2cacb5-52b9-4fd8-8c1e-3c07c138f222 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832c8996-a5a4-4d51-954e-2590435a78e5 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models ImageFolder: Autoregressive Image Generation with Folded Tokens
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ace587-682e-44a5-a454-747b1d399901 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models MotionClone: Training-Free Motion Cloning for Controllable Video Generation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 142d523d-929d-47ef-826e-9084208b1384 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31cde6aa-6ae1-4f83-842c-bddaf0fc6804 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27cab28d-099d-4893-b962-9a205a7dd291 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cae051f-4df6-42f4-9088-051fc050bfbc · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Unified-io: A unified model for vision, language, and multi-modal tasks
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 74a0d82e-affe-4917-bb1d-b664d0bb0b0d · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72ccaa16-a94b-40fc-bba6-6c2f5feba1b0 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models STAR: Scale-wise Text-conditioned AutoRegressive image generation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a93f306-e501-44c1-b6e1-5fa294a51bc7 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Language Models are Few-Shot Learners
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e111547-fa6e-4197-9f10-66514e4e1ca3 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Deep multi-scale convolutional neural network for dynamic scene deblurring
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 863c6666-5de0-4b5f-877f-0322900e14f9 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Gpt-4v(ision) system card
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bb949fb4-704a-4199-8623-9cc2676b8c11 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Gpt-4 technical report
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 84848d12-4ae2-42a0-a72b-bee436e31a94 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b136387b-655e-491b-aef8-a7ef27d2d88a · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models GPT4Point: A Unified Framework for Point-Language Understanding and Generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e900c64a-1e02-4a64-bb5f-e2161ff8546b · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Gpt4point: A unified framework for point-language understanding and generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 825ef939-0a58-405d-a155-b340ced7a3f6 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7085699-5954-413c-8325-55e90c03c51a · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Learning transferable visual models from natural language supervi- sion
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ed437bf1-73e5-430a-8284-32d7eb7f78fd · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Zero-shot text-to-image generation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec5a1d25-08c5-48fb-983e-22c7894a4cf3 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c68f39c7-933f-4d12-9f9b-769b79b13b48 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models High-resolution image synthesis with latent diffusion models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 439ee319-2672-4f02-b6b5-4b13a4cc4248 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models RB-Modulation: Training-Free Personalization of Diffusion Models using Stochastic Optimal Control
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23fd0597-9589-4072-b0e0-1ef186d965e6 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 64de6525-48f6-41dc-9372-75aad7c53018 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Emu edit: Precise image editing via recognition and gen- eration tasks
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 41b9e1fd-bb7d-4783-9013-699820ed19c1 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Indoor segmentation and support inference from rgbd images
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e5966f0a-97fd-4a77-8b17-6ce0b66c6bf6 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Denoising Diffusion Implicit Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 679f3e18-3360-425a-9b0a-40f73e19dc24 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ffc111-bdf7-48c4-a07d-c2be88e849c3 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Emu: Generative pretraining in multimodality
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b766934-8a3b-40db-89c4-4453644f1d31 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Generative multimodal mod- els are in-context learners
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6133a99f-cd74-409a-b854-bfeed499aab6 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Alpha- clip: A clip model focusing on wherever you want, 2023
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation af1355ad-b7f9-47f8-87dd-2d990dd38077 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b6ba75-12ef-4c79-820a-e2486ba25db7 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb423553-6403-4e51-b4e3-d974705d62cf · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6bd74ac-59e0-45c6-8a01-ff96dcad33fb · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models LLaMA: Open and Efficient Foundation Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e233f119-89dc-4dd4-a4f2-78ee6f39916a · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Neural discrete representation learning
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcedd024-df6f-4a0c-aa2b-fa7f9767d2e6 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4b1cc3b-203c-4213-8471-e3764113e26d · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Images speak in images: A generalist painter for in-context visual learning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3b5255e3-d242-4e42-ac82-a20abc71799d · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models SegGPT: Segmenting Everything In Context
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9558197c-75cd-41c3-a606-b08d4cdbee71 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Emu3: Next-Token Prediction is All You Need
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75b4d6ec-fd1c-442b-95de-ce81cc6197ba · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Deep Retinex Decomposition for Low-Light Enhancement
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c5a393-e22f-4037-ac96-b40136512a9a · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70807277-4dcf-418b-af70-11f113fbdd21 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models NExT-GPT: Any-to-Any Multimodal LLM
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52c3278d-7f0a-419e-8630-568d51058738 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f71053a-6c9e-4994-bd71-8f1cd7587618 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models OmniGen: Unified Image Generation
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb8cc4e-b2e6-4657-bb56-9a2b6144c798 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f3ddc0-aaab-4290-9511-f751322438fa · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79a5fc81-a78f-4cd0-919a-e74d5ae98d2c · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Depth anything: Unleashing the power of large-scale unlabeled data
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c03bbacd-17d7-4f0b-a5a7-023eb186eb85 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b5763c5-0701-45f2-b971-f0d2af69017a · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Deep joint rain detection and removal from a single image
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 46906c86-bcfd-4deb-b086-409d8db8df58 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Inverted pyramid multi-task trans- former for dense scene understanding
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0d4f2d3c-152a-4334-83b2-7a10cbbabb39 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc6f3a2-9a93-4714-9cd9-11aeb2765524 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d50d04c-4829-421e-a1f5-fbfcd7359f6d · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Magicbrush: A manually annotated dataset for instruction- guided image editing
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4052ce3d-8bad-4811-9379-49d31fa50f94 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Adding conditional control to text-to-image diffusion models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d91d5ab1-9ccd-4d7d-9f55-7832c2d40ba3 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Internlm-xcomposer: A vision-language large model for ad- vanced text-image comprehension and composition, 2023
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 032a2303-6d34-4de3-9c41-6cd1f5d933b1 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d3d874a-6126-4c04-859a-4692e5ef9177 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b51fed7-f510-4b7d-afc0-9beec1bd5c25 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Scene parsing through ade20k dataset
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11cc75af-ed41-433b-b5bb-1077f47aa075 · outbound
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03e39ec1-c5dc-4dcb-a655-5aad110f1693 · inbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84696db2-b6ed-48cf-9186-844744fabb5a · inbound
CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bf4091d7-c1f8-4a74-b56b-357c74f87fa0 · inbound
T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dbba8e2-49a4-4769-8392-bf3f0ab74947 · inbound
UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.