Pith. sign in

Paper Citation Record · LEDGER

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models

As of 8 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 1 inbound Pith citation observation for arXiv:2505.20021.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20021 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:05:42.456961Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:55:55.359488Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T03:56:21.523747Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved49
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 795b52be-22ff-4369-977e-1c8890156936 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:34.222869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:34.222869Z digest=sha256:2643002c1cfbffb8873b15d6e430dbcfd2c085a40773162ed84cf47d8c61dd0c

Observation 20e25e54-91e2-4695-8a64-5e24b4b2a1c0 · outbound

This paper cites Allen-Zhu and Y.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Allen-Zhu and Y

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:34.332475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:34.332475Z digest=sha256:3e2908dcaa46cd81be791f1a2ee5bf24bc090222424a05957b0d47dc88823ae8

Observation 0e44503a-5e38-46b2-8b14-a4a40bfdd685 · outbound

This paper cites Antol, A.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Antol, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:34.421047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:34.421047Z digest=sha256:e2b0f9e0669bb2ec9a45b4384c97473aff4d144bfab810a53d9472a03b111c95

Observation f6967756-261e-48d4-aa1a-a52a0175bce6 · outbound

This paper cites A Theory for Emergence of Complex Skills in Language Models.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models A Theory for Emergence of Complex Skills in Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:34.543572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:34.543572Z digest=sha256:c0921d6e4a621e73fa084f8388469df19ec603bccb277e2a7844d6506ded0513

Observation b068dc6d-adea-42e7-bc0c-efda2aa181ac · outbound

This paper cites Program Synthesis with Large Language Models.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Program Synthesis with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:34.683309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:34.683309Z digest=sha256:8f740da2fe4e418d4f900c0038968f0f4e1a5cbbee4a7cd93c25156e05a88ff9

Observation 7ce5cc7b-d201-482d-8b99-9a2247b4a6b9 · outbound

This paper cites Qwen2.5-VL Technical Report.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:34.765882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:34.765882Z digest=sha256:4d0140b1a9640b7f401e5cc5b6f2430129093691f8c4c593431d1b06b0517de2

Observation 7c3df1fe-2758-4fd3-82fb-bf47cc7da3e1 · outbound

This paper cites a is b" fail to learn.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models a is b" fail to learn

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:34.894746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:34.894746Z digest=sha256:a29d5b7e414c8d9d924da19435511c94a115fe32f4250c6a82eb61bf188efd6e

Observation e4d8c62c-0d45-4f12-9cc8-0161ea0c5337 · outbound

This paper cites An Introduction to Vision-Language Modeling.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models An Introduction to Vision-Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:35.000146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:35.000146Z digest=sha256:15132a1cef779bc1c686599a8498de7355c3468a66c77629f890c240b386113d

Observation 358c7624-7bda-4f12-b4ff-8efa597a0b1c · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:35.094711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:35.094711Z digest=sha256:e74cf3b01e22f368724b4373a6442a755039edddd94042f5551545fdcbe1bdbe

Observation d9a4eaaa-ecae-4eec-9c81-bdc2cc603662 · outbound

This paper cites Cao and J.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Cao and J

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:35.228731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:35.228731Z digest=sha256:940855a1d649ef25d2f17b572aafb091beb068663a028b1daef908322b7e3309

Observation 2e0cf66b-54e6-4532-ac01-a8a0e44b1e57 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:56.048996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:35.328449Z digest=sha256:2beccbfe26a756b2a2c319c3a3cef7731cf4834a9a0d03efed6f42668bce97b8

Observation 59fa14f8-f2cd-4313-bfce-dde95930057e · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:55.924448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:35.427288Z digest=sha256:44e53dac1bbb17b1580127a2dfbc225445047d80069cbb9a0fbe857e11bd93a3

Observation 9414f5fd-52de-4cdc-919e-3f73b080c0bc · outbound

This paper cites Deepmind.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Deepmind

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:55.777345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:35.514643Z digest=sha256:b78d0ac321ff5ee0ff3dc6cca56f3f92721a0907e1948180406e3a9c010477b8

Observation 4aff6f04-589c-4557-ad9c-2e3f57ba0d9f · outbound

This paper cites Deepmind.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Deepmind

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:55.637881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:35.613946Z digest=sha256:3f94781a1a8a58d7f5a46732c6c8a67a940130c34e098db0d247f64779fe948b

Observation a2ab2c41-9595-48bf-a4e9-1f93b9b59d85 · outbound

This paper cites Dhariwal and A.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Dhariwal and A

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:55.516045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:35.736149Z digest=sha256:23c1d09209fe1b8e46ed81132c5cc4fc805908a011f2ee1a44243fed48eb8da8

Observation f4514c5e-b950-43a9-9500-e2f9b5c10df6 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:55.386058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:35.838571Z digest=sha256:78e0204fc991df4acd05c04c9ff4b27ede5aca0a270b66b312f287a2b4d1b866

Observation c686a3c9-81fd-472a-ab70-28b50d05b991 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:55.249246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:35.971460Z digest=sha256:5f45a189ffbb0fac454f281bc6ee08232a6ee92a5dc1edbf333d2bef2b7119ec

Observation 44fd03b3-e36e-4eee-ac02-9422185d368d · outbound

This paper cites Golovneva, Z.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Golovneva, Z

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:55.150396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:36.084307Z digest=sha256:682ec38658f7856d85b237cd4c61251af57c3472380f72d60dd4d4edd42c7d4a

Observation 771fd9a2-4ae4-4175-a099-99d78da84ebd · outbound

This paper cites Goyal, T.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Goyal, T

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:55.049173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:36.180142Z digest=sha256:b3d2af00347ba1e091b854e7898776f78346caa3179377ce56b40814a249fe2f

Observation 090f2275-d5ed-4061-938b-6d95e84a3065 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:36.273148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:36.273148Z digest=sha256:9f7be60bffd6676550ef524e12908ab40efe7fd58424c2df81b5dce0d81a1540

Observation 83f3ec76-d058-42c8-89f3-3a5d7f9e5569 · outbound

This paper cites Gurari, Q.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Gurari, Q

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:54.907590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:36.371175Z digest=sha256:66950906cf5750387347f5396361634186a92ee4ecddf7fa3519e887f45177fc

Observation 36a57854-9c54-4459-ab07-f40317530ed0 · outbound

This paper cites Hanna, O.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Hanna, O

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:54.739548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:36.441104Z digest=sha256:3610993c949045cbd563271925944d4679d8da3441dd2210c0113bc5fd8f9063

Observation cf492c02-6a9a-4966-bdee-f851caa8fa55 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:54.594042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:36.494586Z digest=sha256:e35a5619a8457997d47618016a534dd58230d57178085a1da3d7e9c3cfdb55aa

Observation 124dd037-4006-459e-b9e1-439b2e4db9c4 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:54.451057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:36.570584Z digest=sha256:2365f4c834c8c7ca645d366e16e242515188f405549b27ca64fca1e09e1eb4a7

Observation cfc3f413-2d1c-4967-bc9a-32acf0820737 · outbound

This paper cites SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:36.645985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:36.645985Z digest=sha256:3ea76681b0a3657e73e614dc55f58dc043a6a2cdec57197da824b8f4a59ac586

Observation 696adc68-3b63-4f9d-9c0b-51f8143a11d2 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:54.340760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:36.698623Z digest=sha256:c60cb5274ce7ecbf96c6a167c867c2afe9917122028aae20ee7bb452ff47a205

Observation 587a6e81-e059-4432-897b-6d3f0cf90c41 · outbound

This paper cites Flux.1-dev controlnet.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Flux.1-dev controlnet

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:54.191439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:36.757449Z digest=sha256:d387b5fa0cb897dc0d88cdad81980b4b4bd728b6d1d42ac2864e58a5a5c56ff8

Observation 813e62f5-8e22-4500-82ab-62a7fe7de83b · outbound

This paper cites Kafle, B.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Kafle, B

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:54.083349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:36.827681Z digest=sha256:78dec020ca2b73a806cc6b550a68bf33efab4a426bcf6c6d06d68f31c21da977

Observation d3140256-ee1d-42e4-acca-471c1fca0a04 · outbound

This paper cites Kazemi, H.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Kazemi, H

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:53.949724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:36.871244Z digest=sha256:473aa1cb97acb752822bdadc195678b4ce319fc983cca8d37d3bf22693bdbb00

Observation e7743a40-1715-4e6f-8623-38fd92585201 · outbound

This paper cites A Diagram Is Worth A Dozen Images.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models A Diagram Is Worth A Dozen Images

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:36.949501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:36.949501Z digest=sha256:24bfc99655def09c3125180e780b17f62c83b2ca0286fcfdf77ccfdcdd848668

Observation 6a2d6922-ece0-4ec8-99d7-e2b87899456a · outbound

This paper cites Kojima, S.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Kojima, S

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:53.814841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:37.022596Z digest=sha256:7a70ccadc7a66dd49737f686b414f32f3fede433d4b75e7c8affa385f75bffc2

Observation e347c95b-fe01-422a-b23a-48049d7a9bc5 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:37.136608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:37.136608Z digest=sha256:839ab4cc93de62378671ae0dd01a77fc5e4a48e0805b1ca53fd138ea12d2066e

Observation e0eed4e1-8c0c-4a80-b808-7dc16247024b · outbound

This paper cites Lake and M.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Lake and M

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:53.598176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:37.202288Z digest=sha256:d57bd4e49ea97532c4e43b5975d02c55b3e32128f2c26576ea14b3836206ed35

Observation 0334fa58-7991-461e-bbdc-c934078dfe2c · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:53.451969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:37.253468Z digest=sha256:aa3aa7383737b2723e076e02b31b8d2a17b442b1bc6728f1897d3e33e3cc5add

Observation e7838624-6517-46a8-845c-793c97f1435d · outbound

This paper cites Lewis, N.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Lewis, N

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:53.198326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:37.329177Z digest=sha256:ef0e7f8920a728cc21bfe420b586eb19674014ae735c433cefff78eac008778b

Observation 50c00425-cdb2-4659-92ac-8812ce71a223 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:52.938778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:37.431121Z digest=sha256:761b2a09332687fd2c1ee21b0123083397e2a7d4559098ecb1876d872fc3ebb2

Observation 04b7d7ec-d4f9-402a-ae47-7cdbe860d0e6 · outbound

This paper cites Revisiting the Role of Language Priors in Vision-Language Models.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Revisiting the Role of Language Priors in Vision-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:37.515954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:37.515954Z digest=sha256:436f4b7b3437396a8d31513686f11033aed22a70117dcd0d8693696de43ba01c

Observation 9bf7053a-96e9-4a40-8c19-89b15847866d · outbound

This paper cites Lin and K.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Lin and K

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:52.686973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:37.558439Z digest=sha256:f3878b6e17822ecb8e76fcc9015d19ec0056afb77ff67b3cf7e1b2730948e285

Observation 527e31b1-f555-4f39-92fc-33db36e646ac · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:52.440764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:37.664224Z digest=sha256:5de99bb3afda04b266ab985da58708e3dccb01b32404acd473f1b58513c57c97

Observation 27469fce-a8e7-4d8b-9c1b-fa5029723cd0 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:52.177661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:37.764682Z digest=sha256:02e447269c8c1b5dae94ac6c3e1bb3544f991f62dc08d068f4421ddab1500271

Observation 26e7c2ca-2100-4312-839e-931c10bef6c3 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:51.862532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:37.845651Z digest=sha256:86864344bf8628095c78ef8e295862fca19c5775b1c761bd635d3e4031da0432

Observation e36b919b-c129-44b1-8a11-a9037a854b80 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:51.505626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:37.971786Z digest=sha256:bf3467369aeb7198c6f6c16c0394d4923b25f541dd140e473c263e5d2ab99db9

Observation 06ee6e18-7cd1-4d54-a346-44c4512dfb3d · outbound

This paper cites Masry, X.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Masry, X

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:51.184852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:38.154480Z digest=sha256:c4cba2a09a2edeeb5bec9587fe0d362585c4ff5c28f87488d43cf6c140e55b09

Observation 196f6d38-4993-41df-8872-b98be6224394 · outbound

This paper cites Methani, P.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Methani, P

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:50.848389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:38.248190Z digest=sha256:876f888242da2c91f8393438b203cd5f1115ef08d99e444da6e02c0fd0f37650

Observation b6d7e9b8-c14f-4437-a0a8-bb74f7ec8cf3 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:50.544657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:38.333927Z digest=sha256:2e33dea6d1c8e099331c2f66a13e4f7e22c4ba7db9cd23689f25291207695e7b

Observation fe102a5a-e16e-4c0b-8303-5dfa42fc70bf · outbound

This paper cites Making Transformers Solve Compositional Tasks.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Making Transformers Solve Compositional Tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:38.474789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:38.474789Z digest=sha256:0e128fedfaf710e324f2baa107d92d2a90de3c1206a2360f8d7b38863e5992ab

Observation 9f149caf-19a9-4704-a454-dd81cab183ce · outbound

This paper cites Okawa, E.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Okawa, E

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:50.249267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:38.580911Z digest=sha256:3077d46d6e5590a4fdd4db37bafcb59c93e7be1b1a99ea5fccab62992ddb220b

Observation 6d71d73a-c66a-4afc-96ee-a41d6b3cd622 · outbound

This paper cites GPT-4 Technical Report.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:38.734949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:38.734949Z digest=sha256:7495239058a72471ca736ef52e612cc2f003c7313a227117643d373a4e7b7a24

Observation 09de245d-0142-4c67-93d8-0ee05fee4b95 · outbound

This paper cites GPT-4o system card, August 2024.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models GPT-4o system card, August 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:49.887153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:38.891448Z digest=sha256:f3a6350941e2faa9b04d178af811f3dd973eb4f9fc736aa295a22fe69f2eb672

Observation 653b0dce-0cff-4c5e-be9c-8ce200a203c8 · outbound

This paper cites o3 system card, Apr.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models o3 system card, Apr

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:49.504220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:39.033465Z digest=sha256:012301907ce0fc2a4ac54a5fc0b6b2a8334c88a98566d40bea056215a9fb3fb8

Observation 52c008ae-f50c-434e-be6b-c6847f3e5d7b · outbound

This paper cites o1 system card, Dec.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models o1 system card, Dec

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:49.235917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:39.156434Z digest=sha256:f008ab04fdcccbec58009b73ae50d50bc4804f8e5670094c5d22ee83262e9c2d

Observation 47820dbd-d0d4-4a2e-bee9-1ebde996c18e · outbound

This paper cites Ovadia, M.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Ovadia, M

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:49.039219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:39.309489Z digest=sha256:5a530ec9495724b30052d06a4416b5b159fb8a16073fe963cc2e9d53d9b342fe

Observation 55371cfd-89f6-47e5-b393-db04ed947137 · outbound

This paper cites Paiss, A.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Paiss, A

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:48.874953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:39.397867Z digest=sha256:cf1781e5fc1c9edfdf5d8f2412fb8ef90f3513675ee244cdda88a68af7d464e3

Observation e73d9d31-a226-4304-8698-fd69a1085cf7 · outbound

This paper cites Press, M.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Press, M

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:48.708003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:39.481124Z digest=sha256:6c888abef946b1d98f9b19fcc7059248c491b551ef58e4238cae73829a204ec9

Observation 6a9de239-60a1-4205-86b4-3b8bb3168eb3 · outbound

This paper cites Rahmanzadehgervi, L.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Rahmanzadehgervi, L

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:48.463697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:39.566351Z digest=sha256:9092757c791c1c8f32d41a5f79e00506b76e8c392fae973085b17cb5a080be9c

Observation f2bde55f-507c-4d53-b4df-587b7e174c78 · outbound

This paper cites Ramesh, E.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Ramesh, E

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:48.264217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:39.681541Z digest=sha256:281239bb942b824e36a8a25e899091e803782ead21eea07aa22efe57b60aa352

Observation 98ee4969-b97b-4c7e-b469-574c6f3b5b59 · outbound

This paper cites Roberts, K.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Roberts, K

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:39.794510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:39.794510Z digest=sha256:60f82ce8f22ac92bc6615161554646272dd331cf418b9321384261e04081e981

Observation 6c97a82e-e4ca-45d9-99c6-6132f8a6d19b · outbound

This paper cites Roberts, K.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Roberts, K

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:48.082975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:39.855561Z digest=sha256:6fa5751540f9d1c1f4c13730247762dfad78e1fd14ddff42b28383a7a1d3a887

Observation 5ad61aff-3e58-4879-b7ea-b05517db02fd · outbound

This paper cites Rombach, A.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Rombach, A

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:47.873471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:39.964791Z digest=sha256:96232c359ae014e6ca7365edec0c81b6d5f72e3bfb53122b80b4933f8b37f6cd

Observation d6712501-79ea-46c8-b1db-fed43894d700 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Code Llama: Open Foundation Models for Code

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:40.074623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:40.074623Z digest=sha256:4b0e4bf9e147dcb98f0d62f85712755cb5c55418aba57fc2fe764243b4bf2396

Observation 3c8a478e-caac-44b0-b84b-ba822c2ec787 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:47.654058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:40.156806Z digest=sha256:8e91d6b9506498053376609e496d043611af4208fcee70670905b4b581b4f9e3

Observation 45e93225-70a3-41df-bdcc-c154954c7115 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:47.429234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:40.262147Z digest=sha256:680ea39a3bfd167ec97def0196847f4ca97ba263b4bff5c6eb850677071e89ec

Observation 939cde22-2f98-49e4-9f8d-624e3c256e9e · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:47.256602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:40.369673Z digest=sha256:16cc004e191e68ec94dfc3b8e54104b110f9e030d13fd5275fc88cf3154f6a42

Observation f8570060-37ff-45aa-bf45-ead4b4bfe9a8 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:47.069010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:40.429951Z digest=sha256:27edd3f4c0c605763b5da81b89c054e094e4a6e72ca895c9c10249e82f68b4dc

Observation cd200761-bbe3-40d8-b4bd-58e560800f0e · outbound

This paper cites Thrush, R.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Thrush, R

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:46.822160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:40.557953Z digest=sha256:d103fbaf32d731eea915af58fc4ba2b8c24b38ab825637260971389d29b90306

Observation d0884905-4640-4c4c-b461-91aac94e8089 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:46.574360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:40.639278Z digest=sha256:0f908d2d89e811913325dbcecc9754380c0fe882a0d4852a524dba6d38dec890

Observation 9b144d5b-065e-4896-b268-0e5268f97006 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:46.340750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:40.740361Z digest=sha256:41be854a7a1c80a4d711f31f57c2fc2d5d9d4bdec2d5e306676ffc05631c31b4

Observation 30abef7e-1847-42ca-93bf-06453b6ee176 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:46.106985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:40.803871Z digest=sha256:d8768ccba70b57b58d21b2ba0e34f227d8c67c50af22e9489b5e13601870b815

Observation 42533b49-7b12-485d-9694-ba1ab1c80c11 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:45.873727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:40.913742Z digest=sha256:9ad7832eb7bc62fa42ccbda7bb5bf233cce49fac7014cd09d9520fd5b2119b7e

Observation 6aee163f-e352-449c-aef4-d1e77d65644f · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:45.585731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:40.997673Z digest=sha256:7f6f48dd6efa88a0ebb8f1089f33b1e0e1b3354103637314853aa878da94fd9d

Observation 9f70d4d5-f90b-4b00-9902-7d318c83adb7 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:45.376294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:41.095837Z digest=sha256:3b53132b08f0c00cb20209bbf8522ade973c6be593b6c58fee31c830dbdb2fc2

Observation 269d781c-7b4f-450e-963a-617ca9416b4e · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:45.146074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:41.161769Z digest=sha256:4fde64c70b9dc605cbecf6fc0c72cc63e44d6b0808b44d9aa7f4f6f232969199

Observation 59871319-7a82-486e-ae41-bd5314374d71 · outbound

This paper cites Yuksekgonul, F.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Yuksekgonul, F

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:44.955938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:41.258830Z digest=sha256:6a9d299e5ed7db2f5eafc4fb23ba342b5556d08709624ea97ecb8bb3a1b26588

Observation 6a39e83f-74bd-4625-9c10-0e87e15c6203 · outbound

This paper cites Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:41.379301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:41.379301Z digest=sha256:10774080cf4597d5a6727b1adb4418ccca771a389a21b7c93495dd3f1b34f67e

Observation 04d26674-4d12-4231-8e7c-91f6db26eaa1 · outbound

This paper cites Zhang, A.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Zhang, A

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:44.707065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:41.494083Z digest=sha256:86887adb1798384b50205c3e245161a1763961c9c11073633016cfca6175c47a

Observation 216b50ac-f24e-4cdb-9b21-24e589a7828d · outbound

This paper cites Zhang, D.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Zhang, D

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:44.526771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:41.555609Z digest=sha256:b6094be8330cb4627f6fd045a17edabde906b7a1de1a3c982ffcd9b011e122a9

Observation d05d4abb-9423-4097-9b59-01f634055b42 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:44.363072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:41.663048Z digest=sha256:3f6238f017ef3444e59311f4362359b4c3cc5bf04a6e8e567f7323e3082c7760

Observation 7b581a1c-981d-4577-bf25-65524f03fa8f · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:41.725304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:41.725304Z digest=sha256:fd74bc547225dbd1aeedaff23f4fb7d0d5b752ba15884e934343d3a2a0ce2222

Observation 8ab0d4a5-0106-4b42-9ce7-830c8a5c0364 · outbound

This paper cites Zheng, X.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Zheng, X

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:44.174978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:41.813669Z digest=sha256:e93f651bbb3085261e5499c6d479098dfa2a8bb2199d1c8a18d907d609a8d884

Observation cc231e01-55d9-4eae-a0d3-2305d9c4f1f0 · outbound

This paper cites positioned in between,.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models positioned in between,

Reference 80

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:05:44.004784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:41.935542Z digest=sha256:b219189add0e09c3e24073eed5cbc7d9cf0ba541c4dc93d0d38ae05e99b66a27

Observation f1fe5e20-9612-4d33-8f16-4e2af5428854 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:43.787447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:42.026236Z digest=sha256:f0972291ad4432a8beff8b9c2bdef1fced59fa74cfd89966ad85a76b6c0af287

Observation 2ce947f5-b8aa-4a49-ae98-0ae21ceb9775 · outbound

This paper cites an unresolved cited work.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:05:43.574047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:42.125407Z digest=sha256:321556d0097237dad9258e0ba894953e33b943672c0a428b58d7272d2120a169

Observation 16def75c-ac7f-48fc-a6e9-a83d06060753 · outbound

This paper cites Diagram on {BACKGROUND}.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models Diagram on {BACKGROUND}

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:43.385816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:42.207499Z digest=sha256:60677bcc87d609bb42d897bf061f7a4b2f0e8d72255b75e3ba1bea6b98ca4a0e

Observation 13172535-6613-41cb-ba5f-f5c5e78e9c8d · outbound

This paper cites For instance, we forced the distance between any point or text to not be too close and indistinguishable.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models For instance, we forced the distance between any point or text to not be too close and indistinguishable

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:05:43.142955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:42.322183Z digest=sha256:d31b104e279a4e746c0771536b3705c84cf1c0377939833bcdb87cdcef130a25

Observation f32108f6-6526-4bd2-9ba6-d54fff3708da · outbound

This paper cites true” or “false.

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models true” or “false

Reference 85

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:05:42.692797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:05:42.456961Z digest=sha256:61741160a03a9257f748631183272e2c6df34024e9c5c721baa90fa486cb4e9e

Pith citing papers

Observation de076bad-79c7-4850-81b7-5e0ffc96fdbe · inbound

RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology cites this paper.

RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:56:21.525579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:55:55.359488Z digest=sha256:31b0d5351e5bd1d35cdb7407000245a10d0e12dc04004c5aa19b6bd412539553