Pith. sign in

Paper Citation Record · LEDGER

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

As of 6 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 79 inbound Pith citation observations for arXiv:2303.04671.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.04671 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T22:50:24.053411Z

measured 137 of 137 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 79 of 79 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:18:39.399123Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact13
  • verified fuzzy42
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

188
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c77ebbf0-4fe9-423a-bccb-50552a95133b · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.302267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:aa5176a1ec84917bcbcdb90be0718e0115af1bde07870212ccbecae73d08aec7

Observation 6fea55b8-019c-4578-b072-bd339dac97c9 · outbound

This paper cites Vqa: Visual question answering.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Vqa: Visual question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.318107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:cb7dda081a813d29fa4bc33023ee4887aba47f56f2412bd7f13521bbe4f6f4ca

Observation c65f590e-691d-406d-a694-e0775b53ccea · outbound

This paper cites VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.137129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:ca06dc7d6f5dcb113eb60db5216b88aaafbfc70ec80a352f16014f90bf1b4606

Observation fa1e7c72-3bf2-44cf-afa8-309a3a5c3638 · outbound

This paper cites InstructPix2Pix: Learning to Follow Image Editing Instructions.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.120824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:c313bd861b2a472b3dbad4c4dae15541ad1f949df47f31c397d4107590aedbae

Observation 4929d6d5-c397-488d-8ecf-d0cab6b38bca · outbound

This paper cites Lan- guage models are few-shot learners.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Lan- guage models are few-shot learners

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.321634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:30edf0b95573fd4e3c671bbf117d7bbc0a9f1116070f1b0cac238daf661a18c1

Observation 5b18d774-46e8-4e83-85b3-d6ece233d377 · outbound

This paper cites Realtime multi-person 2d pose estimation using part affinity fields.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Realtime multi-person 2d pose estimation using part affinity fields

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.325478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:13e80f609282421d16d863a7c0e9b91d01f335d5c90c883283f0e578311ea385

Observation 1b80f66e-0e81-4333-bc5d-0a3737e728ee · outbound

This paper cites LangChain, 10 2022.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models LangChain, 10 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.329029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:7dfb6349d8be085c0ec46a14fbeb7bbfdf65d983ecb03352bd33ca17e796a247

Observation b1e9f1e8-ef46-464b-8cbe-9a1d6d7a3e89 · outbound

This paper cites Visualgpt: Data-efficient adaptation of pretrained language models for image captioning.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Visualgpt: Data-efficient adaptation of pretrained language models for image captioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.333169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:8b79922b9797117efcdf2eb5c9102157223fbdbd6496ea09d5fb0389fb709424

Observation 1b632c2c-9858-42a8-8523-5439223a85f0 · outbound

This paper cites Uniter: Universal image-text representation learning.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Uniter: Universal image-text representation learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.336794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:ef18b339c9a7068ebe781b540de14b15d9787a7637cb6569b46216dac19fc70e

Observation f54fcd22-9f4e-4868-bf4e-7429323decb2 · outbound

This paper cites Per- pixel classification is not all you need for semantic segmen- tation.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Per- pixel classification is not all you need for semantic segmen- tation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.340753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:36736648ffb2398d720b3204fe9d8ae5ed25743d89701bac85c0b13ee6fc69f2

Observation 3214a427-3809-45e9-a35e-9e6c26bfa80c · outbound

This paper cites Commonsense reasoning and commonsense knowledge in artificial intelligence.Com- munications of the ACM, 58(9):92–103.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Commonsense reasoning and commonsense knowledge in artificial intelligence.Com- munications of the ACM, 58(9):92–103

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.344620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:8de2d2654e6ef6542235ccc4b29a946fdab1f6c19e79cfea55832d55c47b6111

Observation 5f24adfb-0b84-4b76-bb85-8fe937d36df2 · outbound

This paper cites MAGMA -- Multimodal Augmentation of Generative Models through Adapter-based Finetuning.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models MAGMA -- Multimodal Augmentation of Generative Models through Adapter-based Finetuning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:50:24.154851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:b8fce71eac88120c4834462c6b5aa68868b4c4eaaad5163592b809c855abe6e2

Observation d54d0cf7-6a98-4114-9471-058a461bb26f · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.171526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:258e47af26588d1b98cdf70963adff57c2acd62825fc2ac29d53e79ddcf2618b

Observation dd4a9cd9-1ad5-4b73-97cd-794981d1c823 · outbound

This paper cites Large-scale adversarial training for vision- and-language representation learning.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Large-scale adversarial training for vision- and-language representation learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.348615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:8f653461d5377c57bbeaa028dc16d01c08bbcbdd2482555ddac4a8354360f55d

Observation e400680a-bccc-48f6-9f1c-b639df85f16e · outbound

This paper cites Semantic compositional networks for visual captioning.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Semantic compositional networks for visual captioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.352654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:87d3bf0f4bc71f34372aaecea82e13b74b044693880d10ecbadeae17b0b26b24

Observation 0c540a3a-740a-4bcd-8f05-9d5a82a54aa7 · outbound

This paper cites Towards light-weight and real-time line segment detection.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Towards light-weight and real-time line segment detection

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.356233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:c73b1564d90a7457af4028d4281ab1d5feb7354c361e3c578865bfd847688973

Observation c61453ba-7dfe-4b02-b3d3-afc3c5198760 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Parameter-efficient transfer learning for nlp

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.359469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:a9a750606041685b4452d78519a8cdc3fd4cebd900ededa5cfa00c490a1ef018

Observation a1ee5456-2ed4-4b13-badc-843e8e542b14 · outbound

This paper cites Image-to-image translation with conditional adver- sarial networks.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Image-to-image translation with conditional adver- sarial networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.362923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:17d477e8584576c9341c65d4acc5d277cda14ffbf4db5c899e824c187f4bf599

Observation 97e19545-13fb-4656-96a8-24ff494b9133 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.366640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:f73250d4d9271c9f0be7f7a1b433cfc2e526d5b31087e2ee57aeffd3af8faa85

Observation cbaf85ba-e755-416f-8cec-81a2447c4bf3 · outbound

This paper cites Large language models are zero-shot reasoners.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Large language models are zero-shot reasoners

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.194209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:143e58e703c2868f2751af7fdb0350670a2df1b5d281c2c81d66150b9d7a241d

Observation 220664ce-1f90-4751-8746-5f3373f5f5b2 · outbound

This paper cites Manigan: Text-guided image manipulation.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Manigan: Text-guided image manipulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.198026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:4bd26468408c59105fcd5bd3f50a05d7371981e3548705a6dfb89a6f5a301e19

Observation f317e750-6e96-41ed-a8dc-7fbd94286ff5 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:50:24.097649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:0a8b41d12357cfbfbb50ee796d2afdf03a99366fa57392240ea898ba085bea76

Observation 8178b679-9458-46d7-9bc3-b567ccc10edb · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.202190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:5d2fac7ef0d4e2762a55322798c2871de748bf67385228c289f47e66ffe1ec8b

Observation 90de9b61-0200-4d79-b066-406a1964ff37 · outbound

This paper cites UniFormer: Unifying Convolution and Self-attention for Visual Recognition.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models UniFormer: Unifying Convolution and Self-attention for Visual Recognition

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.130337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:67a671f5914b454431ea98f2f27f21e6aeec6b6c750aea47880e8167a16a94b3

Observation 14cf1848-40d9-41c5-9897-88ae2056fcee · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.205549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:f3bad1ca9558554832a85ea06166e531b9e1b1ed52d9ee71776a2fdcdc32bfac

Observation 295940ec-0f2d-464f-9e15-298014b8c0ba · outbound

This paper cites Su- pervision exists everywhere: A data efficient contrastive language-image pre-training paradigm.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Su- pervision exists everywhere: A data efficient contrastive language-image pre-training paradigm

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.209143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:97290fc87f25b16b4422d59355d7dd5e86c7e0d76275e5893d7ad8b96664d198

Observation a17744c3-8461-44ea-812b-4bae9ff52aba · outbound

This paper cites Microsoft coco: Common objects in context.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Microsoft coco: Common objects in context

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.214184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:c76d58d9e88d83a80bf8b56cefb656638a749bd2ba7596ffc7a58e03d6c9b7a6

Observation 1fd2fc12-4eeb-4d12-98a4-57c4fd1a9026 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.221357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:f8a1ed8c86f3dab9cb5976c217814193e3580e255d1b1987c6d67f1e1b81dd4e

Observation 3900d762-1960-4828-80b9-f4fdc1be9539 · outbound

This paper cites Training language models to follow instructions with human feedback.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Training language models to follow instructions with human feedback

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:50:24.114173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:2d48c89a7c0eb40a33ab6e3c2b3d6845a776567534fb65aeedb7d8bdac7c7b28

Observation 1f50a76e-e794-4ae1-85a0-159aa9316784 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Learning transferable visual models from natural language supervi- sion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.227091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:0fdf43ee3e4db61a36410dd2fac38d8ddcf506d3e7b10c1446d2310c905673a5

Observation f1495288-8c3b-46f3-877d-78af38119dbf · outbound

This paper cites Language models are unsu- pervised multitask learners.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Language models are unsu- pervised multitask learners

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.230761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:9aad53dd3cc04853d9fb8a46cc09249f1fbb6271d31c74af7e71d787d7a4a58a

Observation 9f8ca77d-ecb1-4cf1-a285-7e314105311e · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.235144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:6004f24f624263b252e45853d4be49efb8796ca4e9d295b4e2c2d70298f7d84e

Observation d9f87b83-ee88-4a33-857d-0e002093c422 · outbound

This paper cites Vi- sion transformers for dense prediction.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Vi- sion transformers for dense prediction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.243412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:746af2e6f85ba1dea51ac89923ff93a6947b18065e109516fce32ee4f6747d9f

Observation 84cba0ae-38e4-40cb-b7d9-d8dbe3afdc3d · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.247197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:546cb3400c295136b25c7148e242d3f67b3cf678f1aa3998b0f49be09748889a

Observation 817465d6-ccb1-4b26-acb3-754a558b4565 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models High-resolution image synthesis with latent diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.251417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:70aa382130f0dd56c62cc1cfde7a0db1131a9c0467e99793e282db7c91e8ee7e

Observation b2cd8133-0cdc-49c4-85c5-8c7ce45ce085 · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:50:24.107337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:c2711d5dd02e2049d8964c93e6375cc5057330dadb1140990486a4b17d8446f4

Observation e47eb0e5-7664-42a2-9606-03bbdc1ee1cb · outbound

This paper cites Learning to summarize with human feed- back.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Learning to summarize with human feed- back

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.256488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:1ebc57e90bc55dd37188a7cdc6571a12b47435c714091ce798cb42d2f56baf75

Observation 7be9af3d-5212-42c4-a288-24a351ce6ab4 · outbound

This paper cites Multimodal few-shot learning with frozen language models.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Multimodal few-shot learning with frozen language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.260118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:f4b649855df858ce715267fce5d24ac07419176995f5708097baf7fe02bd9fa0

Observation 5a804739-3f9b-48ab-bb9d-5f4fd1af431b · outbound

This paper cites Attention is all you need.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Attention is all you need

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.266273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:aeadea823c7718380f5a98cabd39d1dce4ba637a8a2e0c61aa9816c36013241e

Observation 829ab7f8-c1f7-4016-9503-3500f33b456a · outbound

This paper cites Show and tell: A neural image caption gen- erator.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Show and tell: A neural image caption gen- erator

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.269985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:d64e5d174cff8df2789cc823a359e007e04dfffc3820a2461992d17800501231

Observation ac2a14dd-4320-4516-a414-ec492b26edb4 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:50:24.148645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:95f046084847ff14f94629c47ff43b27c7f02c004ba53f14c41297d02cc5d77a

Observation 54132017-1509-423c-b4d5-bc91d9b32119 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.274754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:308e830fe350910dfbedc80766b8a408b4cf3518f3e8ff989dd75e4aff9d40d9

Observation 6e9fe70d-8324-431b-a019-2624fd180394 · outbound

This paper cites an unresolved cited work.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-13T22:50:24.277687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:53adfc5a7334a0e91843883888ccf07724090421b5c02e1f30e006ffb5ac2c45

Observation 5aa7a6c0-ce5d-43ee-afc7-c2f52ae3754a · outbound

This paper cites Holistically-nested edge de- tection.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Holistically-nested edge de- tection

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.281135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:af862817cc577b9ad9e18f53d5d65f52a318b6ac68db17560f53624c6633e2a4

Observation bd0811a2-e2db-40f1-9d23-484580cf2c13 · outbound

This paper cites Canny edge detection based on open cv.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Canny edge detection based on open cv

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.284175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:1776d7d8945f4e21b82b9dd5cf8fbfcc0a5810e4382dbbcdb2838addd2317076

Observation b4069334-9df1-4d95-9bdc-3e70fc8c8a11 · outbound

This paper cites An empirical study of gpt-3 for few-shot knowledge-based vqa.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models An empirical study of gpt-3 for few-shot knowledge-based vqa

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.289407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:b4b314b399e8b10ef14995615518c61a190976f953f4ee6349ab47796306e5ac

Observation fda2cde2-ed85-4635-8610-fe15f8bdb458 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Star: Bootstrapping reasoning with reasoning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.292535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:ce72f5be38050fe417283ee61cbe00451115ffeadb64acb1928d687428be101a

Observation e2b683e9-989c-401f-a02c-3c759d189e32 · outbound

This paper cites From recognition to cognition: Visual commonsense rea- soning.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models From recognition to cognition: Visual commonsense rea- soning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.295501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:1aa499a89c143c9f7a4ef6bc194ecdfa2d215487a71ff81e19af373201851346

Observation bde057b6-c6c3-4594-bc85-0cbdf86da67e · outbound

This paper cites Merlot reserve: Neu- ral script knowledge through vision and language and sound.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Merlot reserve: Neu- ral script knowledge through vision and language and sound

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.298644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:0701622362b9b95b08a77e7fbbcd44e4756d786a799c13dff411195637e341a0

Observation 415f539b-35d0-4f5a-a7ab-b0f3fef1ad1a · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:01.041344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:71035abbe0eb6b4210636b05f24af21f230c445f4b82ca02f4926193f9c43541

Observation 3f9af3c9-6088-4524-80d3-d0f9ed8e4714 · outbound

This paper cites Scaling vision transformers.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Scaling vision transformers

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.315058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:ee78e4a70f81124bd9b70d630b7256596d0391d4b1dcfad6d17d9e271b5c4e4f

Observation 6fade32f-a965-4bc5-90ae-6d7b7f3e0496 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Lit: Zero-shot transfer with locked-image text tuning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.307416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:680d66d16606f92b450627b4bacae0f89d1a8413746833703dee6133a170edb9

Observation bd083eea-4380-47bf-8fc6-bf7329b40a3b · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Adding Conditional Control to Text-to-Image Diffusion Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:43:11.169114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:2651420a8e044e1e6813f8f3a78bce3a6ec99bd8bee28b652824945093d53c4c

Observation b495ffd3-6e09-4754-82ff-957c6e732d25 · outbound

This paper cites VinVL: Revisiting Visual Representations in Vision-Language Models.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models VinVL: Revisiting Visual Representations in Vision-Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.165819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:55c79106ef206faf36163c878c0b400956be96f8e26a3ff2c88e0474594dc0bc

Observation 1465b4a4-8254-47cf-b102-b03d6dc4151a · outbound

This paper cites Text as neural operator: Image manipulation by text instruction.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Text as neural operator: Image manipulation by text instruction

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T22:50:24.311758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:38a36500c04c68e8f4c7c7342e05e524407a3c2217df8112a5d4b1fd1fcfd8fe

Observation f9cbf5cd-57c8-4020-93be-59669aa174cc · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Automatic Chain of Thought Prompting in Large Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:39:17.100641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:488a9a9e9e5975bb71bfc9407b8ac831a756b57bdde322c2aacff4de3e11ce07

Observation 7c9d7d79-955e-4dd1-b16d-b5ff256b9116 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:50:24.184228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:b1bc98618d50d1f6b97a327ca3f5cad29a978c841b6ae2ac4d41402353e2f9d0

Observation 6bb86c84-499c-42eb-b25a-3625162a46d7 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T22:50:24.189906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:b659fbeb2083d809fae284eea483c8cb37ad6361ca9d6696732a4db670dc7b0b

Pith citing papers

Observation 5c680498-cd2e-449b-a239-52f49807e8f9 · inbound

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action cites this paper.

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:17:58.826194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T01:17:58.678036Z digest=sha256:7c2a58d14aee1f0e5069b3e307a8c8c35d7c1231b574bf84599b1e85a47663f9

Observation 31e91514-1fa1-4803-9b41-9ad3b1123be6 · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:a0e34ebfedce6659cfd53605d7eb78a7f800e58351a3bb49865f2f74ed54dcdd

Observation 8d468bce-53a2-40b5-a63e-720b491db788 · inbound

Visual Instruction Tuning cites this paper.

Visual Instruction Tuning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T08:22:03.403362Z digest=sha256:8454df7e75ff9717f9ea6a240c1c125c3b27a1e5ec7aa99bdec678055a0c5e42

Observation d35ff2ac-d4c7-4a1f-9534-8e9292a49e7d · inbound

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models cites this paper.

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:37:01.617345Z digest=sha256:7fa22bcefb158acf1440e2cc4795c50ce9f4ecd4b29d24d439843b7f91f034ad

Observation e7c36291-f5cd-41e0-a187-d197b380e0b7 · inbound

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model cites this paper.

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:41:04.849759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T08:41:04.743886Z digest=sha256:15a802fcf9a1416db45f9c843795f6d8548e83e4b12b041ce03fabfb88ad76cc

Observation 71cca06b-0916-4d52-a063-8629de8e0e6c · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.663677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:ef4852eaf749225a9c6f9f531ef6047d41731c1a8e5c0c34c6451101e0e523ba

Observation 3b82b836-23d2-42e9-b7b9-e5b2083f42e2 · inbound

ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language Models cites this paper.

ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:15:55.711045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T18:15:55.525596Z digest=sha256:03eb12692bfe1c659d14d57a1a799fc1256a3d8d9ff0343e10b42261af3c1bf3

Observation cb64f6e6-214d-4d3a-89dc-b5cec4a84416 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 191

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:56:42.180990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:df8f487169524589f30e34dbe0d683244cd8805b3ce48eb3e760da87217aefc0

Observation 17df00cd-8389-4405-9147-d8a400f4c766 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 289

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:28:39.272777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:47e19cb453b43fcfe1a2ecebbf5176a5b072fe0d2259e5d5626cc4ef94d2a546

Observation b6ca3b5d-d479-4254-8eb8-4388c5c98b86 · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:26:06.435149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:ebd652a8fd021ad1f16fd9a246a9dd368878e09415d5ef80a2f275dd9c5ec9bb

Observation 637c4a68-9021-401d-afb6-e9bf10c27473 · inbound

Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V cites this paper.

Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T14:01:49.854238Z digest=sha256:c4ed9a4515f6241bec863e6dae9a3c1c35660165f5fd02104f0d5029e0abda91

Observation a537c30e-4c9c-4876-80ad-a12fab1c960c · inbound

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models cites this paper.

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:03:27.012613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T03:03:26.723464Z digest=sha256:d74e8c132cdeb88df7d54d78f3457568440edf4fcd175b2a329135e27ca81528

Observation 02ef4b95-ed54-4784-a2bd-5d92a1e68b0c · inbound

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection cites this paper.

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:08:01.311610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T18:08:01.166072Z digest=sha256:27b1e77e119fda38fc6983bb277ed6423f62891a3f651d9bdd1262d0ed581e1c

Observation e44ea933-e0f3-481b-90aa-2e7a407cb552 · inbound

GAIA: a benchmark for General AI Assistants cites this paper.

GAIA: a benchmark for General AI Assistants Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T15:46:03.247029Z digest=sha256:2c67034c1b16582258ae4a7916dd1cf833cf9aad2d66ade05096e44f06ded499

Observation 588a3070-2ebf-4c2e-8552-060c1c9f93ae · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 156

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:82b999030228eac5ff0bda9c232f240df01db355132972c01a9adf90c6b452c5

Observation 5ff7fa79-856a-47f9-a33e-6eb4bc6ac663 · inbound

Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks cites this paper.

Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T06:20:15.656356Z digest=sha256:fd6f8d30fead27154d2ce3992a854c86c3b805dd06b4066e9df0fb35993fa74a

Observation cb41746b-bf36-4cc3-98f8-808cc13902aa · inbound

Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception cites this paper.

Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:19:27.991850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T00:19:27.965902Z digest=sha256:e1ea171080c83386cdc5e63252dbc897b63add43a04f6e1387763213ec8b8483

Observation c5746d66-b271-4125-8f6b-326bb46f19d4 · inbound

Understanding the planning of LLM agents: A survey cites this paper.

Understanding the planning of LLM agents: A survey Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T18:12:57.568144Z digest=sha256:6f2ee46307e1c140714c210132af9876fcd84a41758de050163ae9c0ba605c7d

Observation d6708187-c820-4266-b1ba-c76f8e460ba5 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 124

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:46:16.841199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:4dfec14f3202c22efd00c90b8539a4260609b44d9cc2d1f89d6a66d390b7ed25

Observation 5372a0fe-d33d-45eb-85ea-329c7bd4ca35 · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-17T07:44:47.575767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:05d808165b04ada8075391b73629fea664f69f4513c02c32ef5476cfb175329b

Observation 9ea42463-204d-41b5-a3fe-b656a6c561c3 · inbound

Deep Multimodal Learning with Missing Modality: A Survey cites this paper.

Deep Multimodal Learning with Missing Modality: A Survey Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:26:03.946063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T22:26:03.706725Z digest=sha256:cb73f7d1d095ac09935fd05e6509194aeac45fa61bfb91bbe22dd153cc8e2330

Observation 1dcfaa93-cc33-451e-85f3-1b4a45868699 · inbound

A Survey of Scaling in Large Language Model Reasoning cites this paper.

A Survey of Scaling in Large Language Model Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 222

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:22:09.012108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T21:20:07.238992Z digest=sha256:c5184f09dae9f6f4f12c9547f6de7d6869f5d192b0519dafb23b646d38539ec3

Observation f66e0d49-ae06-40ab-afc5-544496faad49 · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:05:52.185645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:de4659e7c90c6e52c5760e5b469a9be2ac0071754772063fd9a610a4e5e2d992

Observation 240c4979-dee3-40d1-bab1-6639fdf6266f · inbound

Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving cites this paper.

Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T00:20:50.495942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T00:16:35.823270Z digest=sha256:b538514374b4615d97291f298f3b73dfae0bc6579bfe188b5f67ece292a80325

Observation fa55aacc-a30f-4dad-bec8-8cb5b7d95ad5 · inbound

Thought Graph Traversal for Test-time Scaling in Chest X-ray VLLMs cites this paper.

Thought Graph Traversal for Test-time Scaling in Chest X-ray VLLMs Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:17:13.990834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T09:15:46.135084Z digest=sha256:83075a151a8c5cf8e8188f4bcc1751056b4ab2e464d1967e1899e22c40415515

Observation 9783126e-1221-4d01-b0e4-57aaf7e617de · inbound

FaSTA$^*$: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing cites this paper.

FaSTA$^*$: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:17:10.725881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T08:17:08.223736Z digest=sha256:34026b792cd4bd4026a75a28c5a95b8bd34995577b0688a887d8fb5510acf873

Observation b2630173-c5e9-4b66-b4b2-b531121a1309 · inbound

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning cites this paper.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.399123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.399123Z digest=sha256:c0d05672f8ac2137f6f246014eb18a8e95e85530cd171641879c958a2324c1a8

Observation 96a13c92-cf07-433b-82ee-4ac419c8b115 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:33.590502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:33.590502Z digest=sha256:86062c1ecaf74f2ac0a4dad903fd65d938151cc274003fcdfe435cf663c7aac9

Observation 7853432a-773d-4f7e-8780-8574435f3070 · inbound

SketchConcept: Sketching-based Concept Recomposition for Product Design using Generative AI cites this paper.

SketchConcept: Sketching-based Concept Recomposition for Product Design using Generative AI Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T22:21:39.627240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:21:39.627240Z digest=sha256:2e58f4924c1484e42080d442c4d0a07e2c3b13c3771605afacc1a9264f6f46da

Observation 8d3c5c47-b591-4072-bc94-2368ae0cee4d · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 256

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:08.632796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:08.632796Z digest=sha256:6f99e9eb156fa92f7b352e327cb1929d203268b894ec874650ed1d7220e83b62

Observation d81b29ff-cb3d-46a0-b410-15db89530d33 · inbound

PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction cites this paper.

PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:23.360014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:23.360014Z digest=sha256:9c50978d11c5a8f3abe52d5b6568640ee6209c42fde53299b5f51561e0801aae

Observation 4da4bc1a-4721-49b9-bb78-3fc4ff8826fc · inbound

TextOnly: A Unified Function Portal for Text-Related Functions on Smartphones cites this paper.

TextOnly: A Unified Function Portal for Text-Related Functions on Smartphones Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T17:12:47.124406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:12:47.124406Z digest=sha256:4af44965e702ac1a9b76e32dfe952cc2921666746df2852711142cde17934249

Observation afa17747-8fb6-4b42-af11-ed13b8fff6fc · inbound

MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI cites this paper.

MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T17:56:57.832286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:56:57.832286Z digest=sha256:859aa7108d4f506a2886e332777d470e27e3d4d0f44d0c785570482e85d3f883

Observation fbcb9983-f344-4ba0-bfcc-b59c2c067f3e · inbound

SCOPE: Speech-guided COllaborative PErception Framework for Surgical Scene Segmentation cites this paper.

SCOPE: Speech-guided COllaborative PErception Framework for Surgical Scene Segmentation Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T17:37:43.500195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:37:43.500195Z digest=sha256:c7c0bfe65cab4394f3ed4d917141f8d0021440732cc5025fffb66050fa1c8c72

Observation fcee366a-7990-4787-8542-9bed575f0004 · inbound

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering cites this paper.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:14.983970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:14.983970Z digest=sha256:f1be614953062e3254ce2bcf3c4cd6fc2277a3224dbdc2cc9f49570e75fb10de

Observation a36f0762-3aa6-4720-8c67-71e1d37608bc · inbound

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models cites this paper.

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:06.728321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:06.728321Z digest=sha256:b441085945ad2f781b292f2f3305db8e793f6eb8609122a8bda3a878fd7b4a14

Observation 8f955a90-b0ed-4e0d-8c2d-55e0ff89a1b0 · inbound

SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses cites this paper.

SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:20:16.374629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T19:18:36.314029Z digest=sha256:8a4d300709a9303e22e16c0b059f073eef54158e73aec7f9673b20b387b17983

Observation 0233a267-69d2-4fe4-8eb5-4edf7cdeec25 · inbound

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models cites this paper.

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:26:26.964302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T18:25:21.621268Z digest=sha256:90aba130649058d15900f6d65e47a69564fcc31e60c5e24a483ba6adf33f6cad

Observation 8d326644-677d-4c24-a555-7035e0c475fe · inbound

TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration cites this paper.

TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T16:50:50.552104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:50:50.552104Z digest=sha256:8c52eddab12d267bb90fd4bbb447580d171046aeadbaea9df7d5ce4802df57f3

Observation bec383b2-7a51-4fe5-ad44-c57c78f4272c · inbound

CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator cites this paper.

CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:58:59.346867Z digest=sha256:180ad3591d410b018f7ad3ebfec00a2cca6f1da817168c7a3f9dca893e9a600c

Observation 93620732-c587-4df6-97d2-b31213b1739d · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T09:40:35.631188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:40:35.631188Z digest=sha256:a546d5addba0593136df827d7678a03556400a87c6f7fdf1f5c274140659dcd7

Observation 9a0a8960-44a2-4592-89a0-737cca64748d · inbound

Less Detail, Better Answers: Degradation-Driven Prompting for VQA cites this paper.

Less Detail, Better Answers: Degradation-Driven Prompting for VQA Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:17:01.867903Z digest=sha256:25414c1a42fcfdc1e52e9d686666d8a40d18b19d9d0a24276033ee6b4b04d226

Observation b9549798-c9e6-40e5-958b-faa815936aa4 · inbound

Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding cites this paper.

Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:24:17.872120Z digest=sha256:4e65d13629261fdc194290b7f5290ad06d5163fa466d39b2f2c0c3cdc80cc828

Observation d55c2a41-761a-4b5d-9f1e-be321f79cde9 · inbound

Towards Long-horizon Agentic Multimodal Search cites this paper.

Towards Long-horizon Agentic Multimodal Search Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:40:32.137708Z digest=sha256:6d4f06e6720094f481309cbcadbca332f5c3f860fab4137144fc00e7f78f1475

Observation 6bac073c-46c9-48d0-8713-d9648a418789 · inbound

ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution cites this paper.

ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T14:05:34.931798Z digest=sha256:f2bcdf3d31d5d1e2522b5ddb81ceeaa29c19817a81d852c6f26b4ef3b0c81379

Observation b48e7828-e1eb-46eb-927a-74877dbba0f5 · inbound

RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models cites this paper.

RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:24:18.061622Z digest=sha256:a283f85393992e1daa5fae76073488970fe15e6ec67440b733a8a103fc8c6e4d

Observation 9af972ec-a509-40bc-9efc-c37fd6d4d2f7 · inbound

Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation cites this paper.

Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T06:40:02.167172Z digest=sha256:0d61817c6980a66f22a3a4136ca9b62991f051112a07538b01da2ddbffb7ee22

Observation 5333e4c7-cf75-4294-91b8-c41f85b06de2 · inbound

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation cites this paper.

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 138

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T05:10:44.608959Z digest=sha256:a3427ad70b56abcea5b3a82589e22a4294920bfd3836e212a1780f110f565d7f

Observation fb9bcde2-a995-45dc-a4e9-24d807f8df49 · inbound

Probing Visual Planning in Image Editing Models cites this paper.

Probing Visual Planning in Image Editing Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-09T21:40:23.460129Z digest=sha256:9e8b37f87938b17cf5cdb99810f2d497e4ea692062eba0daa29fb26fe15bedca

Observation b88000a1-03dc-45b6-a838-4e416a9da89c · inbound

MIRAGE: A Micro-Interaction Relational Architecture for Grounded Exploration in Multi-Figure Artworks cites this paper.

MIRAGE: A Micro-Interaction Relational Architecture for Grounded Exploration in Multi-Figure Artworks Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:32:57.198044Z digest=sha256:c7fba34d96c370d6a9ade0887f31cc947108071b525e430da6f1212ff440be7e

Observation 71a0f833-70ce-4b10-9deb-24d3cd697249 · inbound

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration cites this paper.

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:11:11.410051Z digest=sha256:f463613edd492496def1b9f543bd823ca0028d624b768fd2028fab888756a31c

Observation 070e3e39-0ccd-460d-8301-9131cb5e2681 · inbound

UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning cites this paper.

UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-07T17:35:28.050906Z digest=sha256:57d3212954a76ddede4e2856f5206052ce9022c95144d7cb7eb6e02db399e34a

Observation 2f09e039-bc02-48ec-b76b-0fdc2df0d3ff · inbound

Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning cites this paper.

Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T17:07:53.418355Z digest=sha256:4586103515e96568c97c67dbce3736f9862166350730871610b80b48d35cc042

Observation 84bdb04d-756b-44da-8c7b-0343ba5d244a · inbound

Cross-Modal Backdoors in Multimodal Large Language Models cites this paper.

Cross-Modal Backdoors in Multimodal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:50:24.368060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:51:11.432424Z digest=sha256:07a3c640e942985a3208f8ecb2ec80d169613fbcdab21e1166f6adb004ac0ff4

Observation 6bb15c0c-50a4-4060-92f4-0177f37ff98e · inbound

Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models cites this paper.

Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:23:44.409801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T21:21:11.388050Z digest=sha256:e21585f5a76e84e453dd0867bb33f664fec67317a2e1fd97efc2def3f3334c85

Observation 9cbe3419-19ab-4ab9-a1b2-2551a758edba · inbound

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs cites this paper.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.707113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:1c331070118964ec6d1c14f92f736cb69e240fa69019bd63179f25193cb16e39

Observation 7086640a-515a-44d9-a468-7e0cb20f1cb1 · inbound

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles cites this paper.

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:51:15.303382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T07:51:13.362986Z digest=sha256:fbe046ea18f350044d2253dc14ce80a71573863081a3f03c0b0d2b81a9082f2e

Observation bb0be874-7acb-4933-9f37-f223acfd1d38 · inbound

ClueAegis: Heuristic-to-Reasoning Cognitive-skill Learning for Unified Evidence-based Synthetic Image Detection cites this paper.

ClueAegis: Heuristic-to-Reasoning Cognitive-skill Learning for Unified Evidence-based Synthetic Image Detection Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T12:34:38.624208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T12:31:43.052626Z digest=sha256:b2b6100578d6de9d50aacd817c2f6a0967239ef0854f8ffab5a025fbddae1250

Observation 6d07799a-5982-4daa-9e4a-375099e6c8a5 · inbound

InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward cites this paper.

InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:46.830016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T17:52:48.888733Z digest=sha256:d9a860daf7c042f41a25a39b15478af54a2ddc40bbb84ccb871467b32917abfe

Observation a325792f-d5d8-472b-93eb-4df0f6719350 · inbound

Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents cites this paper.

Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:13:49.077016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T18:05:59.091231Z digest=sha256:714ca1676cdbed2a661a96e6c563e77f145218b1c7b88d900e09102361945032

Observation 1ad6ab5f-f1fe-4ee9-8edc-5450c5242b17 · inbound

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness? cites this paper.

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness? Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:03:29.601869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T13:57:21.387438Z digest=sha256:7782e260be95b0a34980ea0e8a81b2a7ad7de69e855a384f92db792ba629e181

Observation 23b5c24f-801a-498d-aae8-fd343aa2b41b · inbound

VESTA: Visual Exploration with Statistical Tool Agents cites this paper.

VESTA: Visual Exploration with Statistical Tool Agents Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:56:10.485302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T21:58:11.339217Z digest=sha256:827a27f6ef005f9b93b4461045d4dc63ab0dcf03a414480587561f7b7dce0b9b

Observation 0a06fe68-43f6-494b-ae43-ffa45e125e5b · inbound

OctoT2I: A Self-Evolving Agentic Text-to-Image Router cites this paper.

OctoT2I: A Self-Evolving Agentic Text-to-Image Router Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:26:22.446985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T14:22:57.822123Z digest=sha256:ad61dfadcea0cf02f8c7f196c7d14c77055e3bd4d853caa22ab016878bce7e7d

Observation 9bd144f5-639f-4bdd-9a82-d455d9930287 · inbound

MUSE: A Unified Agentic Harness for MLLMs cites this paper.

MUSE: A Unified Agentic Harness for MLLMs Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:26.849039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T11:14:31.395762Z digest=sha256:6cc1ad349f701443226bdf48b49feaa2428baf7392118ae2139d6d78b06bdbc8

Observation b31f5033-83f0-4800-9973-fe50287098ab · inbound

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents cites this paper.

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:46:28.929685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T10:38:41.735204Z digest=sha256:b7bd664232c2e0bd5d4133c4d6a9f2b0ac7f9b2ee5949caf15e0463d33a1c7e6

Observation 2a8ad26d-e57d-4328-bf0b-078d57464678 · inbound

Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning cites this paper.

Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T17:37:14.745432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:56:19.901930Z digest=sha256:2fe006f010b9856acb287fb66277dc9d32768e96b2909773fce600889c1b5b16

Observation 1a940d70-a0b0-4545-88e3-d5a6aa23bf05 · inbound

HDRAgent: An Agentic Framework for Multi-Exposure HDR Imaging cites this paper.

HDRAgent: An Agentic Framework for Multi-Exposure HDR Imaging Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:17:28.657216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T17:24:18.764625Z digest=sha256:a4c381df05fed78c1c63f8c893f47286bafbff8e8490214cf258ccb027bdbf67

Observation dbf10217-2ab0-46b5-aeda-1c94becce6db · inbound

MedCTA: A Benchmark for Clinical Tool Agents cites this paper.

MedCTA: A Benchmark for Clinical Tool Agents Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-03T09:07:47.988823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T10:32:01.387453Z digest=sha256:44699220395820d5873c10ab96ec948f4e9c7fedc84d74c7b20ee2f23856f2d9

Observation d747a8cf-3e61-4166-8f96-04c3b539c7d7 · inbound

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data? cites this paper.

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data? Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-03T13:58:22.234578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T07:17:43.188693Z digest=sha256:29b9884edd637bed6c936672cfdf9d3fba4749b529da9f7baa5c716f92b6033d

Observation a296de52-98f5-48cc-8c75-fbd9a6c11ee5 · inbound

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data? cites this paper.

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data? Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:37:25.466494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T22:29:58.492918Z digest=sha256:11b5d705ec966fb228c784576d5f12c573e5211eb3458a5d34f57b0c06879a65

Observation 23338d62-8378-4b75-a614-716c34dab6f4 · inbound

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence cites this paper.

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:29:31.594823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T17:56:07.864580Z digest=sha256:4e56e1ec984a1a9d4e50bc506e5f00825765bae69204abf17eaf25f4e925e00b

Observation f2dad7be-97fc-4bc7-9dd1-c5e14cb39a8e · inbound

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence cites this paper.

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:54:39.022751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T10:22:08.535453Z digest=sha256:8423b3da2208d4b19feaec7c225301826908a19aa932ba1b25d8f690b9f7123c

Observation 746377dd-ed62-4931-93ce-2d2f922096ca · inbound

A Comprehensive Study of Implementation Bugs in Multi-modal Agents cites this paper.

A Comprehensive Study of Implementation Bugs in Multi-modal Agents Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T10:41:33.954036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T10:41:33.954036Z digest=sha256:214d24b2754dec52fe00cb30843941cec4ed906405f8a17312ab0cb7a1a347ed

Observation 0058b84e-3db6-4097-8ba6-ec4ad83f4553 · inbound

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration cites this paper.

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T15:24:53.016199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:24:53.016199Z digest=sha256:1fc8765e7e0697092aedb1b3f951c858b6f99c752184946cfbdc4e81980f0195

Observation 0881aa9a-4249-46de-aed8-3c1b32aec5eb · inbound

SPyCE: Skill-Policy Co-evolution for Multimodal Agents cites this paper.

SPyCE: Skill-Policy Co-evolution for Multimodal Agents Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T03:36:00.317683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:36:00.317683Z digest=sha256:fb51e4abc70f68cc59a11e4f21d7bd8c77133171aff6766e147c9bf9a529c15b

Observation ea0ae2a3-a447-437f-835a-d59bbc8a5a1f · inbound

Knowledge-Centric Agents for Workflow Generation in ComfyUI cites this paper.

Knowledge-Centric Agents for Workflow Generation in ComfyUI Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T22:12:55.547636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:12:55.547636Z digest=sha256:a1d693edefd601d73d347dbacd16b4f91766bf4f813ad946fc4f8c734fc6ac26

Observation bd4a26c1-2b41-499d-88cb-640293d72da2 · inbound

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning cites this paper.

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T06:08:22.141248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:08:22.141248Z digest=sha256:03b5a58380147192de30fdd57eb23aaf50ad2e913d62b86255dc0df09bf82b06

Observation dfc8a1db-ba66-494e-a5c3-c8de96e31970 · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:51.524195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:51.524195Z digest=sha256:f2003e518030fb842dfa278fae12316af479d693f6eb9d14e514aca2e2ff061d

Observation 6569cc84-a17c-4ab0-a829-05809cf166ef · inbound

When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence cites this paper.

When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T07:56:28.755281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T07:56:28.755281Z digest=sha256:4d7e1f6d6a6ef94c00069caa2659facc1855c0ab3b1c43cc8b6bb31e7e25db68