Pith. sign in

Paper Citation Record · LEDGER

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models

As of 21 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 4 inbound Pith citation observations for arXiv:2412.01824.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01824 v2

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:56:59.808319Z

measured 97 of 97 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:30:21.554689Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T15:19:51.779596Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved70
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5bfa5040-9f0b-4ece-bcaa-65e0cd1f737a · outbound

This paper cites PaLM 2 Technical Report.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models PaLM 2 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.511711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.511711Z digest=sha256:27d590b8d9063d72ca21d6eef0af40f66a3b196b441b6a1d4eb74b68f5f32fcd

Observation 29acbec3-e92a-49f9-b0cc-fa34529d2034 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.516791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.516791Z digest=sha256:48a61aa4ee55acd06144e93b69c91b890b3865177fceae7c6e73362bd586ff74

Observation c6312d3f-0ed7-427b-b6ab-345fe88efaa6 · outbound

This paper cites Es- timating and exploiting the aleatoric uncertainty in surface normal estimation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Es- timating and exploiting the aleatoric uncertainty in surface normal estimation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.520501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.520501Z digest=sha256:310f3786fbe1e4b89767848e10b7b35440e914b1c2b2d5b750cbccad2cff18db

Observation 9911dee2-13bb-4aae-bc9c-2aa0d6548acf · outbound

This paper cites Sequential modeling enables scalable learn- ing for large vision models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Sequential modeling enables scalable learn- ing for large vision models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.523819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.523819Z digest=sha256:d97cd1edbcbf38c54a3297da996c0c11aac95ac7b6173516c3203d30ea7934c9

Observation b4546b04-90ea-46b6-9bb2-ef7a93e0f190 · outbound

This paper cites Visual prompting via image inpaint- ing.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Visual prompting via image inpaint- ing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.527244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.527244Z digest=sha256:7222c6c06a8092ce751545b6c3b474a6b48f25f00480c41417439bf83dc4e9a7

Observation b7977445-e071-441d-9a4b-87c27b80b4d4 · outbound

This paper cites Improving image generation with better captions.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Improving image generation with better captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.530640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.530640Z digest=sha256:5c187e30d80f4932ce7637a48f2ba2bd091f3af9c4d4919d9ad37df25226ccc1

Observation 255aa901-836c-4acc-a5bb-1f5093ac596b · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models In- structpix2pix: Learning to follow image editing instructions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.533835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.533835Z digest=sha256:4c5784ad562082fe684d60c78f695fa683fb8e9eaf303a2e1814d615fbe1dc4d

Observation 554dec18-e5ac-483d-a105-35f235895d7e · outbound

This paper cites ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.543418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.543418Z digest=sha256:eb6083a0c17b179d3871493c1a81797e818b9e9b6387ccfe2f281d7a05ded512

Observation 8204d8d1-e1bd-460a-9f34-aeecde2685fd · outbound

This paper cites Learning photographic global tonal adjustment with a database of input / output image pairs.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Learning photographic global tonal adjustment with a database of input / output image pairs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.546965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.546965Z digest=sha256:c0077639bda9d3a2f9ea23b955e0ad8e8028fd1d746fce960a3bd54cbba7ca71

Observation 9eaec335-1ec1-4f0c-8bd6-1c0d0a55ee68 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Emerg- ing properties in self-supervised vision transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.549903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.549903Z digest=sha256:5014883468425422f4f541a9f51517133fe12e9b1ce3929d324866c75785c0e8

Observation 2138bdf9-57e4-4c73-a5cc-b739ccfd801a · outbound

This paper cites Masked-attention mask 9 transformer for universal image segmentation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Masked-attention mask 9 transformer for universal image segmentation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.552932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.552932Z digest=sha256:902470b2e7ee5836ffc149ad74db9fa7fbd5e439ec798d812c006df01b52b184

Observation 95668963-d6b0-426b-a6e4-f0e74ed8ca31 · outbound

This paper cites Shazeer, Vinodkumar Prab- hakaran, Emily Reif, Nan Du, Benton C.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Shazeer, Vinodkumar Prab- hakaran, Emily Reif, Nan Du, Benton C

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.555902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.555902Z digest=sha256:014aefce55bf5bb44adad4b3fc6cc3e8771f6801bb2825a34487ed5dd71b9790

Observation 4b10193d-ce38-4e8b-bc29-bc3a35456da4 · outbound

This paper cites Instruc- tir: High-quality image restoration following human instruc- tions.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Instruc- tir: High-quality image restoration following human instruc- tions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.558459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.558459Z digest=sha256:f34b20e31a1b0dc7dd127754623f96f80481adc66fa711b9c49b56c09160493c

Observation af2b56f9-ff1f-4a6a-9cd2-5d273410e54b · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.560793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.560793Z digest=sha256:5fc347b7fe2034ad07c29d4a5eac942e88a65d8d9720db3f40deb7ff8888a1be

Observation 047321ff-fb17-471c-b628-6501bf540668 · outbound

This paper cites Internlm-xcomposer2: Mastering free-form text- image composition and comprehension in vision-language large model, 2024.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Internlm-xcomposer2: Mastering free-form text- image composition and comprehension in vision-language large model, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.564386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.564386Z digest=sha256:70b13e802bf2d4b590dc83e2fd72a5bfd0bbee081fafccfbd6e844ab0fbdf1b8

Observation fb85be1a-ee18-42be-b289-08f4d43408ad · outbound

This paper cites The Llama 3 Herd of Models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.567054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.567054Z digest=sha256:fc19769df52d540cde5d6f9ba9fe866bca647e778f551f7088f9eac72d6bb460

Observation 55972e13-63d9-49c8-8a72-2fc5aebd8ef9 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Taming transformers for high-resolution image synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.570056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.570056Z digest=sha256:e7d5c85d8f703ae905aaf8eb31c7d6232928df4be2c2b8ce110fb30a3b781337

Observation 3c4beb1c-4411-43cb-82ef-13fa678862bb · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.572607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.572607Z digest=sha256:ae63f66441f7e8e5723284537b6d5853a4bba2e465a5e519500102531d1220bf

Observation 9ccd74bd-8563-46fa-abff-8de2b8706926 · outbound

This paper cites Removing rain from single images via a deep detail network.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Removing rain from single images via a deep detail network

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.556021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.575359Z digest=sha256:f19013438591dfa11853c3e1edeba6c4c80fbebbbd32044875adeec90b073747

Observation f241de02-5408-4dac-8cb2-69f89fa3e553 · outbound

This paper cites InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.577864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.577864Z digest=sha256:ec31c6019d9e5487f5944eb982ebc3d2d37b02dc888485883193b1445fc33254

Observation 2a352133-5948-4b17-8b7c-4bd3dd3493a4 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Making LLaMA SEE and Draw with SEED Tokenizer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.580658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.580658Z digest=sha256:bd194b0e78e0c2bb74afafc2e6fd7e403382aa9bc59cdd325f79de6c05b955d4

Observation e58e58aa-8a1c-4340-80b9-ebb31f09af18 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.583467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.583467Z digest=sha256:37828204e87e85fb2cbf1d4af714c4af33e9ead2c2d01665c82750f948ade9ec

Observation b37ea80f-5399-499b-af47-4730d32667f6 · outbound

This paper cites Instructdiffusion: A generalist modeling inter- face for vision tasks.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Instructdiffusion: A generalist modeling inter- face for vision tasks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.546585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.586454Z digest=sha256:ae5759e7697487fa056c7a324cca42adf5032ad9f7f9ad72f573302972a70958

Observation 1a13b393-9199-4a9e-9339-b6194817a20d · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Geneval: An object-focused framework for evaluating text- to-image alignment

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.536879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.589163Z digest=sha256:c90bf0653395fbd540936f46a80105eca55eb07ce0930887832f588ef185b44d

Observation 15daa145-782f-4ccc-912b-ce5b9d7e1597 · outbound

This paper cites FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.592304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.592304Z digest=sha256:02ce39f33723fb173985de9ac83b0a057d8ddb855ea6d6b787c46c85b99ea1a6

Observation c23aa687-ef92-47a1-a152-162285a5093a · outbound

This paper cites Denoising dif- fusion probabilistic models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Denoising dif- fusion probabilistic models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.596018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.596018Z digest=sha256:e48cb8c8710b5bdf7c68cba041ce9ef5d7a7529de4758460bbfc17957cb437c3

Observation bedcf773-588e-41fa-9e76-dad461967606 · outbound

This paper cites Training Compute-Optimal Large Language Models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Training Compute-Optimal Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.599391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.599391Z digest=sha256:c1b95c89e2e7e41e884550ebaf7313b523b094638936439f04df4ffdc3bdad9a

Observation 874e5f9d-2eff-4ffd-9b1c-7bc2eaa1a9ea · outbound

This paper cites Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.603097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.603097Z digest=sha256:bfb5ff53f13a1b422e8f5c67f0d550445a0dbb93deb0ede6edc21641cbb20635

Observation 199e7cce-eab5-458b-8b02-b0795708f5bb · outbound

This paper cites Smartedit: Exploring complex instruction-based image editing with multimodal large lan- guage models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Smartedit: Exploring complex instruction-based image editing with multimodal large lan- guage models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.521225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.606976Z digest=sha256:8a4b673853be4c09a6c6c4d5dd1a13f65b5e51aed3639db7536d86a2d3e7d879

Observation f26e6c39-a4f8-450d-bcb3-123e9c2a254c · outbound

This paper cites Repurpos- 10 ing diffusion-based image generators for monocular depth estimation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Repurpos- 10 ing diffusion-based image generators for monocular depth estimation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.512163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.610236Z digest=sha256:27a33a2dee46af9e9945a05797ab98bc4c5b32c81d3471d61f6b2104dda29c78

Observation c22d2bdd-3c91-45be-94e4-0ff28e766f5b · outbound

This paper cites Auto-Encoding Variational Bayes.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Auto-Encoding Variational Bayes

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.613812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.613812Z digest=sha256:df51b20f34b9ebb41438a0967ee46fcb448f841af52f24d936af5353209d2bc1

Observation 3ecaca46-4ca0-44d1-926d-f17e32371c2f · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.503713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.617027Z digest=sha256:392b5dde78c261108533cc768b84ce875e4425f0aa1ba304c569daa9b65adfca

Observation 0aa70367-b215-4d66-907c-a89e72fe9042 · outbound

This paper cites All-in-one image restoration for unknown corruption.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models All-in-one image restoration for unknown corruption

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.493925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.620124Z digest=sha256:44b228e26afeedad1feec4068814384a90d7b311e3a70a3dd0f24395e2e6d0d4

Observation 85f473ed-78fd-4ebd-89e7-ef3849560b0b · outbound

This paper cites Mask dino: Towards a unified transformer-based framework for object detection and segmentation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Mask dino: Towards a unified transformer-based framework for object detection and segmentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.485722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.623258Z digest=sha256:85a6670734cb90d1195998c4a22d59fa3f83e50e61404f03ebfc29b6f1884ee2

Observation 9e2cacb5-52b9-4fd8-8c1e-3c07c138f222 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.626403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.626403Z digest=sha256:8c35fb89957ce1de57745784bca93deabdfe185423f5b8970fef21d27f67f307

Observation 832c8996-a5a4-4d51-954e-2590435a78e5 · outbound

This paper cites ImageFolder: Autoregressive Image Generation with Folded Tokens.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models ImageFolder: Autoregressive Image Generation with Folded Tokens

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.629870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.629870Z digest=sha256:97e716c09cba4ce9d70b3df32414a722bfe061510a6c49b02d9f2e0b324b39a8

Observation 68ace587-682e-44a5-a454-747b1d399901 · outbound

This paper cites MotionClone: Training-Free Motion Cloning for Controllable Video Generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models MotionClone: Training-Free Motion Cloning for Controllable Video Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.633268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.633268Z digest=sha256:141419ebeee8397a648fad557980205a1d378b2660e3b2a4e5cae2b660b9e692

Observation 142d523d-929d-47ef-826e-9084208b1384 · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.636337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.636337Z digest=sha256:1d8d7b8199b99186dce9dfd53d394600defeba30fe60cf2c276d9883f3efb174

Observation 31cde6aa-6ae1-4f83-842c-bddaf0fc6804 · outbound

This paper cites Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.639527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.639527Z digest=sha256:c3137a786710f93a190e6e06fbbcbe22a971356d25f65f11d52ef504cbb4ad0f

Observation 27cab28d-099d-4893-b962-9a205a7dd291 · outbound

This paper cites MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.642476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.642476Z digest=sha256:5e7dcae848d8aea364c204d4f186b484f92145b989d35481f159837f637bc4c4

Observation 9cae051f-4df6-42f4-9088-051fc050bfbc · outbound

This paper cites Unified-io: A unified model for vision, language, and multi-modal tasks.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Unified-io: A unified model for vision, language, and multi-modal tasks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.476994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.645445Z digest=sha256:304a94dcf5e24d4970d9e3a1521572f05dc15a8affc4e8d3a2fdeddf66d19f5e

Observation 74a0d82e-affe-4917-bb1d-b664d0bb0b0d · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.648332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.648332Z digest=sha256:34a6759baf2f6b71176f42065c668f245a1d5d6bbc3fe2e5c8079c86e19118db

Observation 72ccaa16-a94b-40fc-bba6-6c2f5feba1b0 · outbound

This paper cites STAR: Scale-wise Text-conditioned AutoRegressive image generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models STAR: Scale-wise Text-conditioned AutoRegressive image generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.651679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.651679Z digest=sha256:75cd0f54d69d794b2652b7f24089866c6f16ad4f8aa200aaeaf11fd50a66b7e5

Observation 9a93f306-e501-44c1-b6e1-5fa294a51bc7 · outbound

This paper cites Language Models are Few-Shot Learners.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Language Models are Few-Shot Learners

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.655257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.655257Z digest=sha256:8e4aad2a636b0bb89aa2cc16610bd477d1ce301e10474d71b2452cdb764cb2df

Observation 0e111547-fa6e-4197-9f10-66514e4e1ca3 · outbound

This paper cites Deep multi-scale convolutional neural network for dynamic scene deblurring.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Deep multi-scale convolutional neural network for dynamic scene deblurring

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.658407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.658407Z digest=sha256:a87a483e0f00145357dd8239d3036091f0fbf23d926aa6507fc0046f5d9ae48e

Observation 863c6666-5de0-4b5f-877f-0322900e14f9 · outbound

This paper cites Gpt-4v(ision) system card.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Gpt-4v(ision) system card

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.462746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.661720Z digest=sha256:061e54113861abfdb277fd6c2599da5c43080f8afb9b8eebb0e59230318bf22e

Observation bb949fb4-704a-4199-8623-9cc2676b8c11 · outbound

This paper cites Gpt-4 technical report.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Gpt-4 technical report

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.453288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.665029Z digest=sha256:38f92afe8c9bd8aad07d5a6fbfc565133331f399d3dc28b1c70452770b588d72

Observation 84848d12-4ae2-42a0-a72b-bee436e31a94 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.668213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.668213Z digest=sha256:5754b7bdb5eaff23906d91c30ca32f6d1faeb72e9469a03ee6173286efb38a43

Observation b136387b-655e-491b-aef8-a7ef27d2d88a · outbound

This paper cites GPT4Point: A Unified Framework for Point-Language Understanding and Generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models GPT4Point: A Unified Framework for Point-Language Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.671729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.671729Z digest=sha256:ad34f8814e958b4979648a21b2ef0549f81ad3b6d924cc2269191b912dec6b0f

Observation e900c64a-1e02-4a64-bb5f-e2161ff8546b · outbound

This paper cites Gpt4point: A unified framework for point-language understanding and generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Gpt4point: A unified framework for point-language understanding and generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.444084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.674793Z digest=sha256:3cc8cb53ad42f5066f48e41dc988d622c68d85747e8444ab1926c1082d516839

Observation 825ef939-0a58-405d-a155-b340ced7a3f6 · outbound

This paper cites Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.678101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.678101Z digest=sha256:73ae77f55c8a0a09a25c18553156ed6f05204f8b8210232f284223e6c0a4305d

Observation b7085699-5954-413c-8325-55e90c03c51a · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Learning transferable visual models from natural language supervi- sion

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.434621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.681555Z digest=sha256:f8aae645d7084e956199c67796c8aad92e2e42c8bf61693f795e02e992166b61

Observation ed437bf1-73e5-430a-8284-32d7eb7f78fd · outbound

This paper cites Zero-shot text-to-image generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Zero-shot text-to-image generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.684841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.684841Z digest=sha256:f7a9b853f5d8f0259d48e2e7e02c4a5614d289ce14224304f74a0f4d03257ae9

Observation ec5a1d25-08c5-48fb-983e-22c7894a4cf3 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.688002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.688002Z digest=sha256:f999873b2152ce4b88bf3fae2fd390798a2ba2a187d03175c419add7bcb872e4

Observation c68f39c7-933f-4d12-9f9b-769b79b13b48 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models High-resolution image synthesis with latent diffusion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.419092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.691537Z digest=sha256:f70d90952ecab34d610261ba8b07fde2a5739736029346a0adc194d47e87b0f9

Observation 439ee319-2672-4f02-b6b5-4b13a4cc4248 · outbound

This paper cites RB-Modulation: Training-Free Personalization of Diffusion Models using Stochastic Optimal Control.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models RB-Modulation: Training-Free Personalization of Diffusion Models using Stochastic Optimal Control

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.694280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.694280Z digest=sha256:1b36420c6a84b7a0a22c1d45d6d2000c8e814513da23738844bef4479af689d9

Observation 23fd0597-9589-4072-b0e0-1ef186d965e6 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.408917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.697253Z digest=sha256:08fb03dfda33590a000cee9ff19891a76e593dc26c43ae753c7c2c24d4a2e869

Observation 64de6525-48f6-41dc-9372-75aad7c53018 · outbound

This paper cites Emu edit: Precise image editing via recognition and gen- eration tasks.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Emu edit: Precise image editing via recognition and gen- eration tasks

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.398101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.700144Z digest=sha256:24b8d21ac40f4b136f79169f717a701971753337322a87cd0703f6a4c54c1bb3

Observation 41b9e1fd-bb7d-4783-9013-699820ed19c1 · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Indoor segmentation and support inference from rgbd images

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.389297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.702738Z digest=sha256:d1abe5d1f1f93a829b6bb9ce967bfbf100c0c8a8fa16ba6ffd796cd42327e0f1

Observation e5966f0a-97fd-4a77-8b17-6ce0b66c6bf6 · outbound

This paper cites Denoising Diffusion Implicit Models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Denoising Diffusion Implicit Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.705352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.705352Z digest=sha256:8523a633af2fc0305509206012ad1c5230d55df7234edf5cf8583c019dc9ff7a

Observation 679f3e18-3360-425a-9b0a-40f73e19dc24 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.708401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.708401Z digest=sha256:8470755aaefbfa053ec9fb99f22ba6f7117f96e7089c6fc0475ce8318eeebb45

Observation a4ffc111-bdf7-48c4-a07d-c2be88e849c3 · outbound

This paper cites Emu: Generative pretraining in multimodality.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Emu: Generative pretraining in multimodality

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.711162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.711162Z digest=sha256:cf1675eee316977c02012b51aaea67a62bc27d6ad2ca47ab2bd5baa4cd4d3d18

Observation 2b766934-8a3b-40db-89c4-4453644f1d31 · outbound

This paper cites Generative multimodal mod- els are in-context learners.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Generative multimodal mod- els are in-context learners

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.713917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.713917Z digest=sha256:46788cc56ed14e702dfcc0bc7f9d63f379be85976faddce4854a40705b7a49a7

Observation 6133a99f-cd74-409a-b854-bfeed499aab6 · outbound

This paper cites Alpha- clip: A clip model focusing on wherever you want, 2023.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Alpha- clip: A clip model focusing on wherever you want, 2023

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.370352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.716602Z digest=sha256:c13221abfcc4f721c3141fa461d9380c398e3b0e4a8d51eb5666d9e5dd5d520b

Observation af1355ad-b7f9-47f8-87dd-2d990dd38077 · outbound

This paper cites HART: Efficient Visual Generation with Hybrid Autoregressive Transformer.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models HART: Efficient Visual Generation with Hybrid Autoregressive Transformer

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.719288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.719288Z digest=sha256:dd6ef1f819d47a0743ab9fdccc3797354859f8cec11c943edae057b11a6999f1

Observation 46b6ba75-12ef-4c79-820a-e2486ba25db7 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.722132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.722132Z digest=sha256:2a6ea23ea73093439177ab760f56ca3e6c1c7b877b0fc88049123f477182ca19

Observation fb423553-6403-4e51-b4e3-d974705d62cf · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.725425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.725425Z digest=sha256:386e55af1609d2ca8d5106096098def5a675364802f4001c19e9b6cf47f88c28

Observation c6bd74ac-59e0-45c6-8a01-ff96dcad33fb · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models LLaMA: Open and Efficient Foundation Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.728864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.728864Z digest=sha256:e0421cf0b2ec9be0343ed944aa7c4667eb9e545146710a83b608568b503fba9b

Observation e233f119-89dc-4dd4-a4f2-78ee6f39916a · outbound

This paper cites Neural discrete representation learning.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Neural discrete representation learning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.732778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.732778Z digest=sha256:ab7b9465aa0c6241f307e12de8404dae981b880756e0c727f5ae0654f7acda93

Observation bcedd024-df6f-4a0c-aa2b-fa7f9767d2e6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.736081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.736081Z digest=sha256:1e5eb941511bc8d738c09dc0f1c24eaea52b6cfa41493cdeb54d5dca3229e31a

Observation c4b1cc3b-203c-4213-8471-e3764113e26d · outbound

This paper cites Images speak in images: A generalist painter for in-context visual learning.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Images speak in images: A generalist painter for in-context visual learning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.356301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.739715Z digest=sha256:621067cdbd49f1f6683a7dcec84e2fbbb14042fc3dc52c683b3a6d8f0de1af8e

Observation 3b5255e3-d242-4e42-ac82-a20abc71799d · outbound

This paper cites SegGPT: Segmenting Everything In Context.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models SegGPT: Segmenting Everything In Context

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.742838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.742838Z digest=sha256:1f69ba46f2d4527973f1ed71dd641f6a337ee19a2d6f7d0b2d359b1739d0dad4

Observation 9558197c-75cd-41c3-a606-b08d4cdbee71 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Emu3: Next-Token Prediction is All You Need

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.746232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.746232Z digest=sha256:1fd01da7c7a5ce3a11c5890a8ee07c17b2b4ad45b9abf96adfbedd8f3af9fef8

Observation 75b4d6ec-fd1c-442b-95de-ce81cc6197ba · outbound

This paper cites Deep Retinex Decomposition for Low-Light Enhancement.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Deep Retinex Decomposition for Low-Light Enhancement

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.750151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.750151Z digest=sha256:9e4ee330fd006071a46691f79916956d1d5adf3432e2d61879925da5f0797775

Observation 34c5a393-e22f-4037-ac96-b40136512a9a · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.753498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.753498Z digest=sha256:a37073eb54511545035ff9e9ddf08d6da384c4df41965c9900236bf5887aeb35

Observation 70807277-4dcf-418b-af70-11f113fbdd21 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.756955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.756955Z digest=sha256:0bb8daa49ae42e2419d9a5a5880935364ca5ed6b0c7f5c02bb65f8114513c4c9

Observation 52c3278d-7f0a-419e-8630-568d51058738 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.760360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.760360Z digest=sha256:611ef4ff372fd4ca79c68447cf8f1d1098eea01105663c69139a3c9bd33eb1d4

Observation 8f71053a-6c9e-4994-bd71-8f1cd7587618 · outbound

This paper cites OmniGen: Unified Image Generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models OmniGen: Unified Image Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.763167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.763167Z digest=sha256:11892122645b8ad5ca76e4a037296ebac4c55bfb3205776b75059e66cbd378fb

Observation 1fb8cc4e-b2e6-4657-bb56-9a2b6144c798 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.766390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.766390Z digest=sha256:f1d30ca5b89a4184621f9738958109af326c0af917c16eb6dcd61f571026e179

Observation f0f3ddc0-aaab-4290-9511-f751322438fa · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.769394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.769394Z digest=sha256:bd877ecd7c135ce127be6c47bc668195ef8a82dc5faf1cf4c72002766c6ce8b9

Observation 79a5fc81-a78f-4cd0-919a-e74d5ae98d2c · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Depth anything: Unleashing the power of large-scale unlabeled data

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.772416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.772416Z digest=sha256:533431887218ac002f0a0c21be28fc334e6e194d6a2c9a67a23ca8743dc541ca

Observation c03bbacd-17d7-4f0b-a5a7-023eb186eb85 · outbound

This paper cites LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.774838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.774838Z digest=sha256:6e1ba2dc883558e81a4b19791c4f3e05c0e69cf10d7e24b7987c292ae5e34a01

Observation 8b5763c5-0701-45f2-b971-f0d2af69017a · outbound

This paper cites Deep joint rain detection and removal from a single image.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Deep joint rain detection and removal from a single image

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.341899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.777790Z digest=sha256:2d1370659607b611abf383b3c1970d054b1c314d2b61a1927ed193e122f0502a

Observation 46906c86-bcfd-4deb-b086-409d8db8df58 · outbound

This paper cites Inverted pyramid multi-task trans- former for dense scene understanding.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Inverted pyramid multi-task trans- former for dense scene understanding

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.332943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.780553Z digest=sha256:68dffcf2f7e77a09928cade18ce372dbca16fa00e986e4cdd7fde6f8e268d528

Observation 0d4f2d3c-152a-4334-83b2-7a10cbbabb39 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.783126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.783126Z digest=sha256:77b00932bc16cf81371c20571ca10e6e7b468244e3776ff01622df721aca22cb

Observation dbc6f3a2-9a93-4714-9cd9-11aeb2765524 · outbound

This paper cites Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.785833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.785833Z digest=sha256:7c527e141de8fdb1718cf52aec37434ef96f31e474e8def20c1ee8d35c81a07e

Observation 1d50d04c-4829-421e-a1f5-fbfcd7359f6d · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- guided image editing.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Magicbrush: A manually annotated dataset for instruction- guided image editing

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.323390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.789002Z digest=sha256:55f631359b567114a42ad4df8301a49692268967d27213bad07f493e80e5ccd6

Observation 4052ce3d-8bad-4811-9379-49d31fa50f94 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Adding conditional control to text-to-image diffusion models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.792082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.792082Z digest=sha256:5be6cd0fa10f4be7b7d0146ff90baf916d7001e1fcf5ce1bfe22360e124d1bd1

Observation d91d5ab1-9ccd-4d7d-9f55-7832c2d40ba3 · outbound

This paper cites Internlm-xcomposer: A vision-language large model for ad- vanced text-image comprehension and composition, 2023.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Internlm-xcomposer: A vision-language large model for ad- vanced text-image comprehension and composition, 2023

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:57:00.307195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:56:59.795238Z digest=sha256:3667a8e84dc95bcca1348a464a55c1e251a432eaa3e0d02c413d1e338f1232f7

Observation 032a2303-6d34-4de3-9c41-6cd1f5d933b1 · outbound

This paper cites VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.798606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.798606Z digest=sha256:e40a454d798465adc3cda488a573be188b5edc01b1e3b3f7b4bd385e8b39ef7c

Observation 4d3d874a-6126-4c04-859a-4692e5ef9177 · outbound

This paper cites UltraEdit: Instruction-based Fine-Grained Image Editing at Scale.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models UltraEdit: Instruction-based Fine-Grained Image Editing at Scale

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.801968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.801968Z digest=sha256:d2fa125136e1c202b1ed16a63e3468c92ae2fe737de70d30131e36e762591b3b

Observation 8b51fed7-f510-4b7d-afc0-9beec1bd5c25 · outbound

This paper cites Scene parsing through ade20k dataset.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Scene parsing through ade20k dataset

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.805205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.805205Z digest=sha256:a96ea175e176139a4688daf0829ca3f43c4da7301bdb2df1b7834f0c2156bcb2

Observation 11cc75af-ed41-433b-b5bb-1077f47aa075 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T00:56:59.808319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:56:59.808319Z digest=sha256:d0f30a731157cc5c3eb69499060a51725df441f6bc424bbd6e5a9033d3c59cca

Pith citing papers

Observation 03e39ec1-c5dc-4dcb-a655-5aad110f1693 · inbound

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing cites this paper.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.554689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.554689Z digest=sha256:11c165634f3fea3062f35ceceeaa2eef7ed22e3bf3a8fa027599ce93aff5c974

Observation 84696db2-b6ed-48cf-9186-844744fabb5a · inbound

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning cites this paper.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:19:51.797076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.820124Z digest=sha256:ca168d57ff3c1896e3b6237e0585da0c0d35672d5da702094e59335009eabea6

Observation bf4091d7-c1f8-4a74-b56b-357c74f87fa0 · inbound

T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs cites this paper.

T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T21:15:46.208071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:15:46.208071Z digest=sha256:c9fcbb41964d3da688b15c6eecbb1127c1e8255b5d7d8faa7e99fb5ce1512bae

Observation 1dbba8e2-49a4-4769-8392-bf3f0ab74947 · inbound

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling cites this paper.

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:12.450878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:25:12.450878Z digest=sha256:3418a40d51dcc1ec36a9e7ac5ee842ab468f319bf9409fd404358159108acd1a