Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T22:50:24.053411Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 79 inbound Pith citation observations for arXiv:2303.04671.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T22:50:24.053411Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T12:18:39.399123Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
58 of 58 outbound references displayed
External citation measurements
188
pith, observed 2026-08-05T02:28:24.338817Z
Observation c77ebbf0-4fe9-423a-bccb-50552a95133b · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6fea55b8-019c-4578-b072-bd339dac97c9 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Vqa: Visual question answering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c65f590e-691d-406d-a694-e0775b53ccea · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fa1e7c72-3bf2-44cf-afa8-309a3a5c3638 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models InstructPix2Pix: Learning to Follow Image Editing Instructions
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4929d6d5-c397-488d-8ecf-d0cab6b38bca · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Lan- guage models are few-shot learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5b18d774-46e8-4e83-85b3-d6ece233d377 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Realtime multi-person 2d pose estimation using part affinity fields
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1b80f66e-0e81-4333-bc5d-0a3737e728ee · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models LangChain, 10 2022
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b1e9f1e8-ef46-464b-8cbe-9a1d6d7a3e89 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Visualgpt: Data-efficient adaptation of pretrained language models for image captioning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1b632c2c-9858-42a8-8523-5439223a85f0 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Uniter: Universal image-text representation learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f54fcd22-9f4e-4868-bf4e-7429323decb2 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Per- pixel classification is not all you need for semantic segmen- tation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3214a427-3809-45e9-a35e-9e6c26bfa80c · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Commonsense reasoning and commonsense knowledge in artificial intelligence.Com- munications of the ACM, 58(9):92–103
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5f24adfb-0b84-4b76-bb85-8fe937d36df2 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models MAGMA -- Multimodal Augmentation of Generative Models through Adapter-based Finetuning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d54d0cf7-6a98-4114-9471-058a461bb26f · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dd4a9cd9-1ad5-4b73-97cd-794981d1c823 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Large-scale adversarial training for vision- and-language representation learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e400680a-bccc-48f6-9f1c-b639df85f16e · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Semantic compositional networks for visual captioning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0c540a3a-740a-4bcd-8f05-9d5a82a54aa7 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Towards light-weight and real-time line segment detection
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c61453ba-7dfe-4b02-b3d3-afc3c5198760 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Parameter-efficient transfer learning for nlp
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a1ee5456-2ed4-4b13-badc-843e8e542b14 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Image-to-image translation with conditional adver- sarial networks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 97e19545-13fb-4656-96a8-24ff494b9133 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Bert: Pre-training of deep bidirectional trans- formers for language understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cbaf85ba-e755-416f-8cec-81a2447c4bf3 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Large language models are zero-shot reasoners
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 220664ce-1f90-4751-8746-5f3373f5f5b2 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Manigan: Text-guided image manipulation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f317e750-6e96-41ed-a8dc-7fbd94286ff5 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8178b679-9458-46d7-9bc3-b567ccc10edb · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90de9b61-0200-4d79-b066-406a1964ff37 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models UniFormer: Unifying Convolution and Self-attention for Visual Recognition
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 14cf1848-40d9-41c5-9897-88ae2056fcee · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Oscar: Object-semantics aligned pre-training for vision-language tasks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 295940ec-0f2d-464f-9e15-298014b8c0ba · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Su- pervision exists everywhere: A data efficient contrastive language-image pre-training paradigm
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a17744c3-8461-44ea-812b-4bae9ff52aba · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Microsoft coco: Common objects in context
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1fd2fc12-4eeb-4d12-98a4-57c4fd1a9026 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3900d762-1960-4828-80b9-f4fdc1be9539 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Training language models to follow instructions with human feedback
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f50a76e-e794-4ae1-85a0-159aa9316784 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Learning transferable visual models from natural language supervi- sion
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f1495288-8c3b-46f3-877d-78af38119dbf · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Language models are unsu- pervised multitask learners
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f8ca77d-ecb1-4cf1-a285-7e314105311e · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d9f87b83-ee88-4a33-857d-0e002093c422 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Vi- sion transformers for dense prediction
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 84cba0ae-38e4-40cb-b7d9-d8dbe3afdc3d · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 817465d6-ccb1-4b26-acb3-754a558b4565 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models High-resolution image synthesis with latent diffusion models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b2cd8133-0cdc-49c4-85c5-8c7ce45ce085 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e47eb0e5-7664-42a2-9606-03bbdc1ee1cb · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Learning to summarize with human feed- back
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7be9af3d-5212-42c4-a288-24a351ce6ab4 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Multimodal few-shot learning with frozen language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5a804739-3f9b-48ab-bb9d-5f4fd1af431b · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Attention is all you need
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 829ab7f8-c1f7-4016-9503-3500f33b456a · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Show and tell: A neural image caption gen- erator
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac2a14dd-4320-4516-a414-ec492b26edb4 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 54132017-1509-423c-b4d5-bc91d9b32119 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Chain-of-thought prompting elicits reasoning in large lan- guage models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e9fe70d-8324-431b-a019-2624fd180394 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5aa7a6c0-ce5d-43ee-afc7-c2f52ae3754a · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Holistically-nested edge de- tection
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bd0811a2-e2db-40f1-9d23-484580cf2c13 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Canny edge detection based on open cv
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b4069334-9df1-4d95-9bdc-3e70fc8c8a11 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models An empirical study of gpt-3 for few-shot knowledge-based vqa
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fda2cde2-ed85-4635-8610-fe15f8bdb458 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Star: Bootstrapping reasoning with reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e2b683e9-989c-401f-a02c-3c759d189e32 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models From recognition to cognition: Visual commonsense rea- soning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bde057b6-c6c3-4594-bc85-0cbdf86da67e · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Merlot reserve: Neu- ral script knowledge through vision and language and sound
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 415f539b-35d0-4f5a-a7ab-b0f3fef1ad1a · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f9af3c9-6088-4524-80d3-d0f9ed8e4714 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Scaling vision transformers
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6fade32f-a965-4bc5-90ae-6d7b7f3e0496 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Lit: Zero-shot transfer with locked-image text tuning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bd083eea-4380-47bf-8fc6-bf7329b40a3b · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Adding Conditional Control to Text-to-Image Diffusion Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b495ffd3-6e09-4754-82ff-957c6e732d25 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models VinVL: Revisiting Visual Representations in Vision-Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1465b4a4-8254-47cf-b102-b03d6dc4151a · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Text as neural operator: Image manipulation by text instruction
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f9cbf5cd-57c8-4020-93be-59669aa174cc · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Automatic Chain of Thought Prompting in Large Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7c9d7d79-955e-4dd1-b16d-b5ff256b9116 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Multimodal Chain-of-Thought Reasoning in Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6bb86c84-499c-42eb-b25a-3625162a46d7 · outbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5c680498-cd2e-449b-a239-52f49807e8f9 · inbound
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 31e91514-1fa1-4803-9b41-9ad3b1123be6 · inbound
A Survey of Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d468bce-53a2-40b5-a63e-720b491db788 · inbound
Visual Instruction Tuning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d35ff2ac-d4c7-4a1f-9534-8e9292a49e7d · inbound
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e7c36291-f5cd-41e0-a187-d197b380e0b7 · inbound
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 71cca06b-0916-4d52-a063-8629de8e0e6c · inbound
VideoChat: Chat-Centric Video Understanding Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3b82b836-23d2-42e9-b7b9-e5b2083f42e2 · inbound
ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb64f6e6-214d-4d3a-89dc-b5cec4a84416 · inbound
A Survey on Multimodal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 191
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 17df00cd-8389-4405-9147-d8a400f4c766 · inbound
A Comprehensive Overview of Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 289
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b6ca3b5d-d479-4254-8eb8-4388c5c98b86 · inbound
The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 137
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 637c4a68-9021-401d-afb6-e9bf10c27473 · inbound
Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a537c30e-4c9c-4876-80ad-a12fab1c960c · inbound
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 02ef4b95-ed54-4784-a2bd-5d92a1e68b0c · inbound
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e44ea933-e0f3-481b-90aa-2e7a407cb552 · inbound
GAIA: a benchmark for General AI Assistants Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 588a3070-2ebf-4c2e-8552-060c1c9f93ae · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 156
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5ff7fa79-856a-47f9-a33e-6eb4bc6ac663 · inbound
Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb41746b-bf36-4cc3-98f8-808cc13902aa · inbound
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c5746d66-b271-4125-8f6b-326bb46f19d4 · inbound
Understanding the planning of LLM agents: A survey Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6708187-c820-4266-b1ba-c76f8e460ba5 · inbound
TempCompass: Do Video LLMs Really Understand Videos? Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5372a0fe-d33d-45eb-85ea-329c7bd4ca35 · inbound
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9ea42463-204d-41b5-a3fe-b656a6c561c3 · inbound
Deep Multimodal Learning with Missing Modality: A Survey Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1dcfaa93-cc33-451e-85f3-1b4a45868699 · inbound
A Survey of Scaling in Large Language Model Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 222
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f66e0d49-ae06-40ab-afc5-544496faad49 · inbound
Grounded Reinforcement Learning for Visual Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 240c4979-dee3-40d1-bab1-6639fdf6266f · inbound
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fa55aacc-a30f-4dad-bec8-8cb5b7d95ad5 · inbound
Thought Graph Traversal for Test-time Scaling in Chest X-ray VLLMs Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9783126e-1221-4d01-b0e4-57aaf7e617de · inbound
FaSTA$^*$: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b2630173-c5e9-4b66-b4b2-b531121a1309 · inbound
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a13c92-cf07-433b-82ee-4ac419c8b115 · inbound
Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7853432a-773d-4f7e-8780-8574435f3070 · inbound
SketchConcept: Sketching-based Concept Recomposition for Product Design using Generative AI Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d3c5c47-b591-4072-bc94-2368ae0cee4d · inbound
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 256
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d81b29ff-cb3d-46a0-b410-15db89530d33 · inbound
PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4da4bc1a-4721-49b9-bb78-3fc4ff8826fc · inbound
TextOnly: A Unified Function Portal for Text-Related Functions on Smartphones Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa17747-8fb6-4b42-af11-ed13b8fff6fc · inbound
MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbcb9983-f344-4ba0-bfcc-b59c2c067f3e · inbound
SCOPE: Speech-guided COllaborative PErception Framework for Surgical Scene Segmentation Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcee366a-7990-4787-8542-9bed575f0004 · inbound
ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a36f0762-3aa6-4720-8c67-71e1d37608bc · inbound
MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f955a90-b0ed-4e0d-8c2d-55e0ff89a1b0 · inbound
SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0233a267-69d2-4fe4-8eb5-4edf7cdeec25 · inbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d326644-677d-4c24-a555-7035e0c475fe · inbound
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bec383b2-7a51-4fe5-ad44-c57c78f4272c · inbound
CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 93620732-c587-4df6-97d2-b31213b1739d · inbound
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a0a8960-44a2-4592-89a0-737cca64748d · inbound
Less Detail, Better Answers: Degradation-Driven Prompting for VQA Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b9549798-c9e6-40e5-958b-faa815936aa4 · inbound
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d55c2a41-761a-4b5d-9f1e-be321f79cde9 · inbound
Towards Long-horizon Agentic Multimodal Search Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6bac073c-46c9-48d0-8713-d9648a418789 · inbound
ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b48e7828-e1eb-46eb-927a-74877dbba0f5 · inbound
RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9af972ec-a509-40bc-9efc-c37fd6d4d2f7 · inbound
Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5333e4c7-cf75-4294-91b8-c41f85b06de2 · inbound
AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 138
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fb9bcde2-a995-45dc-a4e9-24d807f8df49 · inbound
Probing Visual Planning in Image Editing Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b88000a1-03dc-45b6-a838-4e416a9da89c · inbound
MIRAGE: A Micro-Interaction Relational Architecture for Grounded Exploration in Multi-Figure Artworks Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 71a0f833-70ce-4b10-9deb-24d3cd697249 · inbound
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 070e3e39-0ccd-460d-8301-9131cb5e2681 · inbound
UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2f09e039-bc02-48ec-b76b-0fdc2df0d3ff · inbound
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 84bdb04d-756b-44da-8c7b-0343ba5d244a · inbound
Cross-Modal Backdoors in Multimodal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6bb15c0c-50a4-4060-92f4-0177f37ff98e · inbound
Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9cbe3419-19ab-4ab9-a1b2-2551a758edba · inbound
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7086640a-515a-44d9-a468-7e0cb20f1cb1 · inbound
Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bb0be874-7acb-4933-9f37-f223acfd1d38 · inbound
ClueAegis: Heuristic-to-Reasoning Cognitive-skill Learning for Unified Evidence-based Synthetic Image Detection Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d07799a-5982-4daa-9e4a-375099e6c8a5 · inbound
InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a325792f-d5d8-472b-93eb-4df0f6719350 · inbound
Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1ad6ab5f-f1fe-4ee9-8edc-5450c5242b17 · inbound
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness? Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 23b5c24f-801a-498d-aae8-fd343aa2b41b · inbound
VESTA: Visual Exploration with Statistical Tool Agents Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0a06fe68-43f6-494b-ae43-ffa45e125e5b · inbound
OctoT2I: A Self-Evolving Agentic Text-to-Image Router Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9bd144f5-639f-4bdd-9a82-d455d9930287 · inbound
MUSE: A Unified Agentic Harness for MLLMs Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b31f5033-83f0-4800-9973-fe50287098ab · inbound
ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2a8ad26d-e57d-4328-bf0b-078d57464678 · inbound
Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1a940d70-a0b0-4545-88e3-d5a6aa23bf05 · inbound
HDRAgent: An Agentic Framework for Multi-Exposure HDR Imaging Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dbf10217-2ab0-46b5-aeda-1c94becce6db · inbound
MedCTA: A Benchmark for Clinical Tool Agents Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d747a8cf-3e61-4166-8f96-04c3b539c7d7 · inbound
TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data? Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a296de52-98f5-48cc-8c75-fbd9a6c11ee5 · inbound
TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data? Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 23338d62-8378-4b75-a614-716c34dab6f4 · inbound
S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2dad7be-97fc-4bc7-9dd1-c5e14cb39a8e · inbound
S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 746377dd-ed62-4931-93ce-2d2f922096ca · inbound
A Comprehensive Study of Implementation Bugs in Multi-modal Agents Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0058b84e-3db6-4097-8ba6-ec4ad83f4553 · inbound
CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0881aa9a-4249-46de-aed8-3c1b32aec5eb · inbound
SPyCE: Skill-Policy Co-evolution for Multimodal Agents Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea0ae2a3-a447-437f-835a-d59bbc8a5a1f · inbound
Knowledge-Centric Agents for Workflow Generation in ComfyUI Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd4a26c1-2b41-499d-88cb-640293d72da2 · inbound
Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfc8a1db-ba66-494e-a5c3-c8de96e31970 · inbound
Reason Before You Retrieve: Agentic Planning for Multi-modal RAG Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6569cc84-a17c-4ab0-a829-05809cf166ef · inbound
When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.