Pith. sign in

Paper Citation Record · LEDGER

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

As of 7 August 2026, this Paper Citation Record lists 100 of 129 outbound references and 45 inbound Pith citation observations for arXiv:2402.11684.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.11684 v2

Coverage vector

measured 100 of 129 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T22:20:21.427717Z

measured 145 of 145 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:28:02.771576Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T17:07:25.737454Z

Reference resolution

100 of 129 outbound references displayed

  • verified exact4
  • verified fuzzy89
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d2d683a-73e5-49f9-aecc-1b27beacc083 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.924640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:885387d53dac2e80ce85fb3e1fb215e9b6c88c5264a4433053fd849817f29f0e

Observation c2537725-e74f-4c91-8305-e37a9d132f3a · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.969058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:6e1057125ee70ebf151ac0d449fe992328d53c5a329c601947256b6098ba95bd

Observation 491845b6-90d9-44ed-b7b6-947560ba8a38 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.972356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:72893749e96d335bd2dd0484367d9487e29edbb4cb73c1a8ecf0a3bf3ef9fe76

Observation e7294995-6413-44a0-ac14-6c1f788230f3 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.975254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:761dcbccfe092342292b3ebaa2053710c2a9e4e249b437611dc89c102471980e

Observation e6b9af04-0852-4d05-8783-832f9e0234f0 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.978398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:9b18ca943f776dc4153b38340d4a72591d5bcc353dc28fe2b2ba8cf6b88c2701

Observation c7096c37-3649-4320-af67-c54aba6502b0 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.980983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:4d85bd08b3bd48b1a1162a543bc5449db9c4398992cd77d31e071bd24218171d

Observation 2b775ada-2493-4478-9f65-54978b561c9a · outbound

This paper cites Introducing our Multimodal Models , url =.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Introducing our Multimodal Models , url =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.984180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:3cdb8a29a0cd4936a07625adeb22e4545a62589351b3e7a2fa9068da52d9e1aa

Observation 8a57a259-5d04-4df7-8fe9-a9a1f15f1fd3 · outbound

This paper cites 2023 , howpublished =.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , howpublished =

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.987109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:b0af77fc179332823a3f2abea515baf8a468d7d2e2b0d2638c71b1f4de335e9d

Observation 4041ad40-fa85-4071-92f4-61c4401e36b7 · outbound

This paper cites an unresolved cited work.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:20:21.989872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:f0c21b5b00751b1c2094088fea511e9744927862bb9a4bc47d8615237c338af2

Observation 81503d65-50e5-4020-b37f-944af054eb09 · outbound

This paper cites Vision-Flan:Scaling Visual Instruction Tuning , url =.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Vision-Flan:Scaling Visual Instruction Tuning , url =

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.993173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:8e91e29340ed30d7415ce2b2ac83d494ccd40822715115a8cb6cb1d23bbf5a14

Observation 44388c28-940f-4b21-a68e-ae5388db61f2 · outbound

This paper cites LIMA: Less Is More for Alignment.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models LIMA: Less Is More for Alignment

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T22:20:21.561755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:f20f901708f490dc5b324fe5f00a0bd409ab7958307d40002030e268a212ec08

Observation b2b65ead-b4b2-4f62-92f4-f60724531658 · outbound

This paper cites 2022 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2022 , eprint=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.996128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:782453d72b3a69bacf2fd587f22c46c6d22d86c2e182d6899822768c07ee3332

Observation 653a7255-9f02-4f2c-a643-a9efa08ca86d · outbound

This paper cites 2022 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2022 , eprint=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.999078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:aa376a976ed0bbb3506ce9f678281babe9324ced41311e01774c72372e641ea1

Observation ac03eb89-2e31-478d-b1de-444eba8ecece · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.579450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:f0ee2e2865a425360405edb5bb6e18933ee5ac1e5c3b6e5336fd2097e3226ffd

Observation 2bb80d9b-327a-430c-b4f4-cc186303cd69 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.583650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:74b0b6e121b00a94c63cd9e05a755a267e82465ee8510250a87fa5587111e123

Observation 5529c882-0365-40af-b32a-025c0c678ceb · outbound

This paper cites 2015 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2015 , eprint=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.587726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:4e6b7f0c78acdf49fd608a12cd7b134c0751b2f0df9502d0b466625132cf0ca5

Observation 2761e794-4eee-4be3-8331-a62df4cd818a · outbound

This paper cites 2016 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2016 , eprint=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.592062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:7a7ed34267293b4f4fa82b09b379bc8459d3594936c9c8aa0d573cbe1313b071

Observation 9a364230-b94b-4b0c-824b-194c81bb0575 · outbound

This paper cites 2019 international conference on document analysis and recognition (ICDAR) , pages=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2019 international conference on document analysis and recognition (ICDAR) , pages=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.595577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:ca679655111478b5f96da9536ebbd27fc8e743bd4cca1bf5bbf88bb3a021dfb3

Observation f7de95bf-1e31-4817-9aa5-654de6c00192 · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.599458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:650337e03a875e30673476fbaa52d9b054f8d7705bb79f90782fd48aa3b91599

Observation f314c186-052f-4e68-b298-b7de2e46e4c9 · outbound

This paper cites 2019 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2019 , eprint=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.602734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:035940da356d07ad1906a3ebd2ad0cff7d1d8e323b0e75236acef1dd2cb559a0

Observation b52e2ce0-e127-4caf-a304-db723afd055f · outbound

This paper cites 2024 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2024 , eprint=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.606652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:8131ec5aac8e461cf23c6229f6101b081478e7168b95bf3527fced7954696b30

Observation e66aa8fc-b5a2-4309-9f2b-451c59ba5630 · outbound

This paper cites 2021 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2021 , eprint=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.610156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:ab3609c0646de5c947dc9c9140d598249f5b09fca6681c5d98aea72b1a36aa53

Observation f2d30a15-bf42-4802-8f6a-bf172d9396f8 · outbound

This paper cites Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.620429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:9fae369cb94cfe1d85d6a11d96bf40fde904dde72196e33ce55d6d252cccec67

Observation 342c2df9-f52d-4545-b04a-326a3135fb98 · outbound

This paper cites 2021 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2021 , eprint=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.624409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:057a0ea7d61c6d129140a5f87c96cb1ccbee78fd8260e17b7200bd2bcdd81055

Observation b15b2755-52d0-4faf-9afd-3eeea75888bf · outbound

This paper cites Advances in neural information processing systems , volume=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Advances in neural information processing systems , volume=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.628255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:7e26b87f157705620c2f177175a9dde3c60fbba874f94819cdf9bca291e2eb99

Observation 31dbce92-e27c-4df3-95d2-a976e0054382 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.632109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:f1c4129452b187964be4cf4f747c5130bf94a247a4ad196420a5c1b8ec6e8819

Observation ec4803c7-c36e-40f4-8e98-0077178dccf5 · outbound

This paper cites 2024 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2024 , eprint=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.635823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:2a2a50207007af1568b43e00d73f520db17c44038d320d44845b5ed7f5ebaad9

Observation 6c41fa27-360f-4d73-a9e6-ac8b34d64a29 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.639697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:577a84660ab7c0b5a5faabd81a267a5eee362883ba5150787ffee9c9ad19d14e

Observation d678331f-1226-4d5f-ab12-1497f4f11b06 · outbound

This paper cites 2024 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2024 , eprint=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.643190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:aa3050cba78d391d898963ec81a4d46954a61810c45fcd5731d85c094497a7d2

Observation 9c763088-936e-4117-846d-c9005c092b69 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.647196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:f78c808d0114912f201ccb4442cd6ec5ff88be23c31006291c3d569b9638c1d0

Observation 64bb3308-17a0-4d28-a27d-974286800dde · outbound

This paper cites an unresolved cited work.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:20:21.650231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:7b89c1b08f1481efaa68550529232f5711e903756a1e4cf58a50243572f20c77

Observation 03275804-3167-4c0a-b26e-f08b8cd5cde8 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.653987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:15fb1cf299a35d3c6dc3af642ece11bb4b713f140377e9fdd5f2f4b3935bb30c

Observation 640d4cdc-74d0-4c4a-ad7f-02a9c05bccd6 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.657674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:625dec8dfddc2bdfe56f0b74ed60308d3cbb9c2288a1c08e7853afe9b12a9a84

Observation 77284240-dda2-46c6-9465-5f2bcf952931 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.661012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:a6fcbd672fbeee1c166c3f29a36deaab86b3f9399d9212a2f7b1900c9b96096e

Observation 20e6a517-f5e9-496e-a7ea-2e883cc060cc · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.664482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:9c369de641dd52f823b1daa24e364af2127e3258c7977e4f79f08c6ecabf47a4

Observation ea12edbe-a67f-4603-917b-44961e0e5a4c · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.668136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:d330355eb1b634e14ea2e8e177cf1be2f9229e812c5c8c7ae386d027fa89d2e6

Observation fe31deec-7e16-4f6e-ab73-c43d7230412d · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.671283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:96cefa62bedb6ab355428b7466f6018553d0cb94c058375672fe1def7d314b64

Observation 08637f11-df65-4b36-96e4-e25aee153dca · outbound

This paper cites 2021 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2021 , eprint=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.674153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:95c562ad3820c4de7a03e9f93f23fbabf979b1f1b4eda6fdfd7ce4d8a0fdd2ad

Observation cccad769-eb23-4b6a-ae05-73e822704d42 · outbound

This paper cites 2021 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2021 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.677836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:72770ca83f9fbaa0e1243a634691a15aa1460c8b3fd4a8ff1852c737dff47c81

Observation 2867262c-0806-431e-a96e-d69978d76030 · outbound

This paper cites 2021 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2021 , eprint=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.681685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:d9cf30962c1251f3056bed224d4a69f891064ad24282c94776a21c64a8c85580

Observation 81396910-be46-42da-a8bd-6562720adabe · outbound

This paper cites and Stoica, Ion and Xing, Eric P.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models and Stoica, Ion and Xing, Eric P

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.685194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:0db9ae9bf18d49ba09f19f30a6b407aef9a8bb4071006b1a890709391c0bb434

Observation 2c88cf11-40d6-4f60-b450-75541e1df7d2 · outbound

This paper cites 2022 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2022 , eprint=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.688017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:bb6dea2df77a925c45ec557820a466bdd42f96cb5da80feaa0b5cd4730004375

Observation d8512e99-c20c-4292-8635-f9e2904e77ad · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.691395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:d464bcd7b12ee595ca01c11225eda508ef5fe46e44002ca740415733dd0b52ff

Observation 2a90798a-c33c-488b-95ab-a06655f2fa64 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.695671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:69569face427c1c1b9be839641e9ccc6f415b7bb274640203a6adb2e3f8c0ad7

Observation 42b7c44c-d132-4065-8939-01ec6b74359d · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.699400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:526f2a923b48c20b9995c675afe152be5e4bf863881e77d9a5d13e2c9ac0c39c

Observation 95697ec4-af41-44c1-a16f-2cf994274dea · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.705715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:a09fb9de5e4dcda47ca06752764e53ae97791d251072431c5ddc27d0e971ad3b

Observation 130973af-461b-44ad-9635-7c731621ff65 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.709343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:eacdada9d2a8102cb5ce03dd74cfcbea383d930db88e07374b7ce5102fb50091

Observation e985cbda-1e97-4a7b-8112-bf29668133b1 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.712716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:bd579066bc2d2710fc5cd04268922e028713c3851e4e4a25caf920e229eb322d

Observation d5dde8f3-90df-43a0-bc4d-767f24ea125e · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.716507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:147e4d8291b60d15b9b6eea998d7e82e1f6f3fcf0d006a7d717c46899e482b47

Observation dcd1b7e0-0f8d-4a22-80f1-8d1e31b630ec · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.719656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:a66a76602399f74f6aea4ef04de624d3f660a33b5ba24177a2c6a6d13c5dfc1b

Observation c97d7154-6b70-4f32-bcb3-ab6216e28512 · outbound

This paper cites 2024 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2024 , eprint=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.723153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:1b28e275a3c6a32b0c045c685973244d2277ad578ba9a388cd150fb38862691a

Observation f4384be8-2529-4326-856f-7c6c2200c731 · outbound

This paper cites 2009 , publisher=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2009 , publisher=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.726986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:15fbcc2434c5edc27b77ad83ebc7595554c2aa4314050c5bbdf197c783c49b1e

Observation 868d901a-f915-4ae0-9030-ffb1fa32bd3b · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.731508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:6b43958a00eb27431c3f77af50590cda118297a61e57d949196ad9fc38605f54

Observation 1a780d89-ddf8-4c5f-a88e-b4c7b7e40652 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Transactions of the Association for Computational Linguistics , volume=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.734954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:247d02f5005f225f29755f7a241f36ec4cafbd3a75b66c3b284621a0debbc060

Observation 53b17f81-28c1-4ed7-9d11-dcb53d741470 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.739284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:eaed6b40d7322bb214253e8b047e7897edb577f46b77a80d83c4eec95eda4968

Observation 08752f88-a987-485f-8c0d-d5ad64c66248 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.742709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:17f929b1f1fd60662334e1a07bfbb82029c17f656b82a8683e57eb98db59b31e

Observation 85893a63-bbdf-4e22-be26-8255eba7e608 · outbound

This paper cites 2022 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2022 , eprint=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.745665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:d00f9eb3f80648eb631b17de5b9de183d8826862f83d6ef3ae13e21b7dc01e45

Observation 5b28d85a-173f-48f6-94bd-92dfa9fadbb2 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.748785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:1b72ca9fce653420a91d06a44a63ce87db2de758ffd29d6c2d14be12d0a22ad7

Observation def1f119-3f35-45eb-9521-c7fe23ebc60f · outbound

This paper cites 2022 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2022 , eprint=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.752010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:6c1b5f95845d2444fdd32fd2d4bf90ed57198877953cb89662820865199db433

Observation 50e4f4d1-5156-47d7-a703-82349913dca7 · outbound

This paper cites 2021 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2021 , eprint=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.755292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:2cee3a617e8c9919b0a1952ec45bd91810cc8ac5d66ecbe8caf97aea5bf52510

Observation 889f30d6-66ff-46e3-a8cd-ac864de3c0a6 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.758464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:40a7b56d2ef87d83bd0261c28fd5df146d71fa5d3c14ba5748c13444310b97bf

Observation 1d1b284d-f935-45ae-b303-4fc392fec446 · outbound

This paper cites 2022 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2022 , eprint=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.761731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:fda69e8ce8ae414bb17defbc11b8e11a622a557c76052cb2937c6f468578658a

Observation b23d89ee-492c-4e7a-a013-8dba0bbd83cb · outbound

This paper cites Ilharco, M.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Ilharco, M

Reference 71

Resolution
metadata mismatch
doi, observed 2026-05-23T22:20:21.494761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:1d52d5174bc808b7979cb516ecaf4c199c6ebb2fd30a416cb1e04ad9c802ead5

Observation 3c0f7ea7-d416-4146-8cf0-75fb7c87e0c3 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T22:20:21.555822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:682651c198e52f9d5556f50ab0c59778fa910a3777458ea7b9f61296656f98e9

Observation 0e9966ce-37d1-44f0-93b7-2048ccca50f7 · outbound

This paper cites Journal of machine Learning research , volume=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Journal of machine Learning research , volume=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.764916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:eed7143974c96809231f2d9ac288173e7ef904c9be1d1cb59862ac8ff96bd30c

Observation be2d446c-6f5f-4912-ad96-f1a291d8d32b · outbound

This paper cites MALLET: A Machine Learning for Language Toolkit.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models MALLET: A Machine Learning for Language Toolkit

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.768773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:7bc05c4255b4b2f177c8235f8c85ff0aea9aa882272007ae135d001dee7fbc34

Observation 10cf12b4-4978-4ef0-9fcc-fdc6bd908b85 · outbound

This paper cites 2023 , eprint=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models 2023 , eprint=

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.771895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:95edc717fcfb405505144bc99d6e8d0297817afb63a34a6578553c1727012620

Observation 32d13e77-f61a-4c37-a4a6-e532db40460c · outbound

This paper cites author=.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models author=

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.774842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:2e1ffeae0f4decb48232f3e97b9a138bcb2c44e879a8b416bd3f8ef1743a660a

Observation 7b8f70a5-b063-4cdb-aef6-48d3396a4cce · outbound

This paper cites Chandra and Dexter C.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Chandra and Dexter C

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:21.505274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:90673aba378363a4bef0d0774d13525469eba78f5290b45759f9bdf6acf2f505

Observation dcac1eed-0ebb-491a-b960-166c4efa90cf · outbound

This paper cites Scalable training of.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Scalable training of

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.777780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:c12c2cb0657be1302a83bf923db742fba8307b7c6093eb2a64744f0d3e8cd33c

Observation 1ede3114-024a-4da3-927d-086b40a14301 · outbound

This paper cites an unresolved cited work.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:20:21.780927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:96ea45b5ec3aa5393e795f14c0a6503a420991e8e58e305b280d2d79d9710285

Observation 3b1fd3fb-9745-42ea-aa2f-0e51e5f14d2e · outbound

This paper cites Tetreault , title =.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Tetreault , title =

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.784256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:c9c450ac1a744097236cd86a76f99a0b5b8c1065887cb522f757a06fed007ed1

Observation 1cb50b23-0e7d-42c6-9dc1-56bd9a7a53d5 · outbound

This paper cites A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.787537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:104c004c9d45a812a70530fa6d2c1432eb811150bfc9440ed564e1f8a639b8e7

Observation e5e25d00-0079-4f9b-ae9e-13f995e26361 · outbound

This paper cites and Tukey, John W.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models and Tukey, John W

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.790362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:dbb5f0b78f47547678167d86aa6aae4a12860d75f66bed6b27f4fa524468df07

Observation 66ed7113-b106-4834-9797-bbc5901cf189 · outbound

This paper cites Aho and Jeffrey D.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Aho and Jeffrey D

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.793548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:fbfbbb165a8eadadb9d5f5496f18a624984f5f7584ec0c7203c75069df898921

Observation ee221ca3-c181-47f2-8079-9466e7e6e69e · outbound

This paper cites an unresolved cited work.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:20:21.797329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:332fe852b317b3918567ce9687a6f411385c474b817131bd0440d543a37fd037

Observation 3817f296-48d9-494c-953d-d93ed57dcd4c · outbound

This paper cites Hewett, Jamie Huynh, Mojan Javaheripi, Xin Jin, Piero Kauffmann, Nikos Karampatziakis, Dongwoo Kim, Mahoud Khademi, Lev Kurilenko, James R.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Hewett, Jamie Huynh, Mojan Javaheripi, Xin Jin, Piero Kauffmann, Nikos Karampatziakis, Dongwoo Kim, Mahoud Khademi, Lev Kurilenko, James R

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.801272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:02cb57882524d0e1613813e7c036de44cea278bda2fc2e480162cad94970bcbe

Observation 1ccf5454-28e5-461a-a525-038fe49142fb · outbound

This paper cites Nocaps: Novel object captioning at scale.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Nocaps: Novel object captioning at scale

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.804481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:0406c37ac6f6a330dda38aa76b601aab5f53e589a1d7198fe033ad11bd4b1164

Observation 6d274f27-fba9-4712-b910-459621df9cb5 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023 a.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023 a

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.807697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:8793872a8498669657b1f929d1e534b51a46e4e1f6966450157ead8ad5ca3480

Observation 20ec7aa3-2814-4e87-9e78-c87026f775e7 · outbound

This paper cites Touchstone: Evaluating vision-language models by language models, 2023 b.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Touchstone: Evaluating vision-language models by language models, 2023 b

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.811178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:795a7239daf15c3b37729c35b5fc3e558d504bc60e491e56e2a5bffc4b29af0f

Observation 9e0f0d3b-1789-4b9e-8e98-1e974de303a3 · outbound

This paper cites Latent dirichlet allocation.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Latent dirichlet allocation

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.814593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:9cae87b6590858328b9d4d02b2d77232cbb49137c453fbe231fa5a8c18887081

Observation 5eb180f7-2442-4c7a-9999-de2859a1ce72 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.818572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:41ba34b7bf9cbe3588d91736e04efb790113ac7aea385e8421e402c683724112

Observation 1356a2e4-adec-4f61-a46a-02785d6caf6b · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:20:21.543720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:fad27d07e0d17e85b91f1e2fbf438bf31a647742ed2db7f1322e949d247cf859

Observation 12d96111-2720-40b2-9b76-e3cbc86953f5 · outbound

This paper cites Phoenix: Democratizing chatgpt across languages, 2023 b.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Phoenix: Democratizing chatgpt across languages, 2023 b

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.822026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:b3ba2f92fdc58ca19c312e201900de411e0829e0f6366533b2d7d9262808ed17

Observation f88c942c-3803-4db3-ab38-982550ec80c8 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Gonzalez, Ion Stoica, and Eric P

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.824882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:bf4078692dbe3fd75262e4b9a73e5fe98b14fe70b3993327798b42e267b8e771

Observation ced2eb04-bf6e-4629-bf64-b210f31dadbd · outbound

This paper cites Mobilevlm : A fast, strong and open vision language assistant for mobile devices.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Mobilevlm : A fast, strong and open vision language assistant for mobile devices

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.828154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:e57aa2098c22fd8fc1cd2dbe7e2c7c7f1ff8fa908fdf4cac03bd52b302d50ea5

Observation 62c3fc62-8603-4219-8272-b505e84d3f81 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.831564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:ed4880c985efde712e1221475fcd7233bec69d28d9a2c16d8b3d751822c4e3c7

Observation 6702dafa-4bf8-4cc9-b834-a48acdd1b6c4 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.834696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:464b1387f7f8720fec8768e3380d47e80fab29afa8af8daf3cb8981f5b62c314

Observation 0515903f-faf0-488a-aa9e-39721797804d · outbound

This paper cites Mllm-bench, evaluating multi-modal llms using gpt-4v.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Mllm-bench, evaluating multi-modal llms using gpt-4v

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.837932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:1d21e1f158ce136dd402e73261c508593243290bb8ae87896682a8667e9de7f7

Observation 66c33652-3752-4a31-bdee-ff5ab1ee64e4 · outbound

This paper cites Hallusionbench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Hallusionbench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.841508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:192f251513755cf5546709beeb74bd3fcbd03595fcea5ea09beb78d4ec367e9b

Observation bacb8e41-b547-4573-b237-c2c20e2998e5 · outbound

This paper cites Hudson and Christopher D.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Hudson and Christopher D

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.862301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:7b530633d3a7392bdba53712461d0572f542a6cda03c7c5ad78adbf465d66892

Observation 53b20947-e357-44ba-a96b-8a9779ef5b46 · outbound

This paper cites Introducing idefics: An open reproduction of state-of-the-art visual language model.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Introducing idefics: An open reproduction of state-of-the-art visual language model

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.866025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:840b1306666e2bb6d92b980154657be12ee676ea0d7db7053bf1ebce205cac32

Observation 901b63b4-68e7-4a75-9ce4-8859f4c7549d · outbound

This paper cites TinyLLaVA Factory: A Modularized Codebase for Small-scale Large Multimodal Models.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models TinyLLaVA Factory: A Modularized Codebase for Small-scale Large Multimodal Models

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:21.537698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:0d79c5c42d2770f6319bb71c1d23d8675461f76ddc088eeca35d20a6f02620b3

Observation a981e278-73ff-494c-81c0-d73501d91d8e · outbound

This paper cites Shamma, Michael S.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Shamma, Michael S

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.868871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:1cad1f1298aa19c8e1051eb6580520774734e3a16842337b6d900073f6469117

Observation cd9df9c1-8a33-4183-950a-690d86f4bbe4 · outbound

This paper cites Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023 a.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023 a

Reference 103

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.872055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:e32e813b210dce142f78336c56cfb7c866747ff7a822a1a5f6441844ed0583fe

Observation 26b3f7ae-e8f5-4f3e-a94c-4e7c50e5880a · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 104

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.876392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:8434cae057cccb5f84e03c78550b31d64c21e1d3027901b94e1d16bba51330fc

Observation 319d4a34-5fb3-42cf-beb0-dfa1ba506e61 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023 b.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023 b

Reference 105

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.879581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:0fda8c6120a2850a86f689c9ac0b8e0b92f7d0693d1b329de514e5aebe2ffd64

Observation 2c650b40-4f4a-4d41-bb7a-cc3b06e5423c · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Silkie: Preference Distillation for Large Visual Language Models

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:21.522610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:39c622af912fdc9614b4bc63b8c8bc69694233960dc861fe46d93e90273ff490

Observation a78c0d9a-9103-471e-83d7-a84e24bb6b72 · outbound

This paper cites Lawrence Zitnick, and Piotr Dollár.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Lawrence Zitnick, and Piotr Dollár

Reference 107

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.882748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:80a360a0b33b4fa413dd2076c6e975736d821231c96261c1d220b124a50cdd38

Observation 71ab9a6d-b6f1-4d9d-80df-8bff52901a14 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023 a.

ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models Improved baselines with visual instruction tuning, 2023 a

Reference 108

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:20:21.885325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T22:20:21.427717Z digest=sha256:050be9d642f440b1b94a5905786a3c05023a53a9ee53aab4ae350e27d5bd0e8d

Pith citing papers

Observation 7f484300-e510-404c-b8f2-cf66eb5029db · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:f0deff77cd82a45d464c2a2cb1c014f068817b8279fd81b750acaad804e4add8

Observation 64a706b9-3249-4f6c-aa99-61d74a1960ea · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:f6351ff86efaf00033aa4cb405bcd2d754e012fca94b64df2d06179d5c60a736

Observation 0843e0c7-7283-4bf8-b8e7-4ace62ffb8b5 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:137a04b85bbea61297206cb556ae7a688bd06b342238f6ab5adbdef0d0112bf5

Observation e750d7f9-9365-4719-9f7d-3889b382ab3b · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:53ecb0edce0f168e7ec596a2e0beae14eee13b6d57624cfc7b6f7e8b472385cb

Observation 57684e2e-a9c1-4d60-ab39-4700ec59b21a · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:17b743e29ad665dd777c6dc1b16297a3eea8f77dff8753622aeeda2e3818a448

Observation 56e4d9ec-0f37-4b91-a8f7-79e175ca21f7 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:a0984d3ddd0a10865e06b3104522ce060984b5678e344da621cee4d18c20dfdb

Observation 19822b22-57b7-44b1-9d38-ca6665b1f39d · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:dc1de758d848da7b4e01727e79cafe7362ef5d0e3fb4aa8e23c0a89adb490a3b

Observation 188228ea-7320-4e61-9843-0f905066800c · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:1f65a2601adbc4361b53e14e93c3478c5e87ccf28997822aca3043d4b4494e94

Observation 1b9acf5c-c89a-4316-912f-b40cbaa05b3a · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:fbf45ad5d73e40755254fc6a42aedc4d18d20b4cc67c419547b0f40ea6c4b081

Observation 7c203caa-1a8f-46c5-b985-4fe3afd5b520 · inbound

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents cites this paper.

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T15:37:25.781240Z digest=sha256:43f80aa76bb579a832602b13be3b5e7be13b7830101c76680e161aa1343717f8

Observation 12465fea-7643-4fdb-a10f-7759182e5100 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:a35915cb7b7f215bbe315d061d45bcaf5f1c4a1c2c50206233ed23691742e840

Observation 60562b2e-a27c-4d56-b8c2-6b73887b1e0f · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:59fdf539a3cad7ed46245028a4768e1ed50584bf8db8d83f302e304f9102e75e

Observation 8a85821c-0631-48c4-9033-01459026a816 · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:0e08be8e9bcd0f607e62f8198d4de3a1ba3a509441fbcaff11e827cdc185fd74

Observation f0646d9d-8115-40e0-9c7f-0ae1c7abd54f · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 98

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:cba27c63c965c227462341a57eadd86aa6ff79d50ffeced9b7d8d0778da06306

Observation 7311a77c-640e-4ac4-9681-d6e43ef63cfe · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:e548522043a574bec4864e7152ba9a0ea4ecd0a1fe0ec632ec9aebfc1af78efe

Observation be7b2f50-db56-46cf-a2b5-3f3ee2c0b42b · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:b60138943f3180b70f9072a68db596fd4bef5f51aa7427f250e630e17182b5ee

Observation 29948f7d-4168-40ed-9555-073fad16b4fd · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:0836758af4b2bc550444923894fbb2a16275cd8e76cda0e85960d6d2d3ec7aee

Observation b8c5d58f-23a4-4a3e-9389-833e49223ddb · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:a5d18c6158401bf9f0bc23b549c7b6656a5a4f382306c4fcb2bca7a297c7ac96

Observation 09d2fd83-eeef-4386-bd3f-8c938035f13b · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:9efa24e7a847f3ed38c603caea688c03b6f831cc1d1539fd2d16eff0115a52d9

Observation eac4d61d-d032-4427-a682-7161337aadda · inbound

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation cites this paper.

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:02.771576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:28:02.771576Z digest=sha256:8c05a58818ad35a80ee1cb875d6a3dd6cb6d5fcf823a4a08cd83b82fceeadd02

Observation f09fa8bd-6a5a-4514-a61f-37094c21544b · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:38.026930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:38.026930Z digest=sha256:8958050aa41abac51e4cdbae61fc53ba81a55cc6f970076d9998878cdbc2ad2e

Observation c74c8ca2-29e0-459b-8ba7-1b55b8ab2e6c · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:26.584334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:26.584334Z digest=sha256:8e38b94cb162263a3cf47a6917c02de1ce0dd92b2ae8e6d912d22d58537ec7a3

Observation 244696a9-e355-4869-9b6a-18fa1cf73075 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.268162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.268162Z digest=sha256:05bab3a984b03198d007f0a818c15cc1c53ecb7b7e8573378d7c16b111031206

Observation ae659049-9bdc-4c7c-a9a9-ae018ab13db5 · inbound

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos cites this paper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:42.059251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:42.059251Z digest=sha256:af9d0d88dfbc5142f9431b9a11191dc2bab81c7d4f03233566b1da4b2a3c3c9b

Observation bc5c1853-1c5a-4793-94e1-21bd4be5d566 · inbound

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces cites this paper.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.035221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.035221Z digest=sha256:edd0c32a307f3474e104e44db2127f73826cea20291e42837b05c6b3bfdcce50

Observation 70656234-ccfe-4790-b2aa-9f3d34e43bcd · inbound

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs cites this paper.

See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:12:23.530239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:12:23.530239Z digest=sha256:35adfb7ddad8f7989bc957128b9ae5559d267bbe8949e1d821ffcf91c030fa96

Observation 9c68df4a-2ce6-4e09-9029-597b53eeb8ae · inbound

Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models cites this paper.

Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:38.690498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:31:38.690498Z digest=sha256:086b7f06182589e5c4148ac624a21b3b0cd37e5a3a0b6fb617bd57db4b10fc65

Observation 0f65a0b2-c7f3-4021-9f79-277a0b1b0540 · inbound

UItron: Foundational GUI Agent with Advanced Perception and Planning cites this paper.

UItron: Foundational GUI Agent with Advanced Perception and Planning ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:34.344201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:34.344201Z digest=sha256:6e2860525d3145cef1cc4783330ff3bcdc2fc98d6f2e7830af9c17791026e9a8

Observation 11029495-1e23-44bf-945f-aa1297ea12f3 · inbound

InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation cites this paper.

InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T18:57:10.308113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:57:10.308113Z digest=sha256:b743e22950a326b40d8ef6a5b10aa7b744c471d5f5c430dbfbfe9045dbdf96cf

Observation 3142315b-10aa-4957-acab-53d14926c21f · inbound

Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework cites this paper.

Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T11:28:05.491861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:28:05.491861Z digest=sha256:42f16e22b6d2288bb55cc52dc1b743e6c0cab98fcaea2d55965ab203e265f507

Observation d1b0a231-d878-4ea6-b605-325350dca75e · inbound

LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models cites this paper.

LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T05:10:56.991368Z digest=sha256:294c10ce11609b07c538e91193c81499cdf51bb5fd1817ab4b4a838252f7cb97

Observation ca2f9252-a0b9-4afb-acd6-f99770275511 · inbound

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models cites this paper.

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T02:14:59.487486Z digest=sha256:85e6e8244b8317195479d334849dfe16304d1944165c65adf1198ddb6938091f

Observation f32c4e1f-9cd0-415c-9197-f90f2246059a · inbound

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models cites this paper.

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T21:47:32.112057Z digest=sha256:77513e498b08c0bf29a40b5f769c7b36597243573ba5c9165531a5454444dc96

Observation 0345c5d3-e620-41ad-85ae-c1afd9c177e1 · inbound

A Nash Equilibrium Framework For Training-Free Multimodal Step Verification cites this paper.

A Nash Equilibrium Framework For Training-Free Multimodal Step Verification ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:10:36.290559Z digest=sha256:b234183a60b2eada97592f0224e3e9f18c5e8fab83040e50557de30858081144

Observation 6b6c3cf9-6943-450e-be68-f3ad21c67230 · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:23:28.499217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:703d970370a505e917da4affb9ab8ce296704a62930642b9da7ac90bddb3a554

Observation bf538334-bcce-42a9-92c9-b16a4252d183 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:17:29.051154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:ffe3dbe85f08067a22850d65ca5c4905c106f3dfcae76d1ffd5848b63c9354de

Observation 0b6e4a42-02b3-4b5c-98df-dfa2c8b94edc · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:45:47.633263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:61290e5e8bf3559d96cf6b5a7009a9fd67ceb7e80fdc41b3cdfce0f540070ff4

Observation 6d30b0ad-a591-435d-953b-7ea19ac1832f · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:24.218321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:31ac210d415cca0b28e910ac62c01e31501aa1ef894d31f2d475471e7c5b6365

Observation 9bf5078f-b318-4c00-bc21-b63c3a4982a1 · inbound

Infinity-Parser2 Technical Report cites this paper.

Infinity-Parser2 Technical Report ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-10T17:07:25.738616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T17:02:28.089092Z digest=sha256:0f1a1390adca1340c9d132a1879ae3a5c049a1c09ebc469a5e70ae08c406e2c8

Observation d9ce6021-0736-4b19-8f31-9f19cda1160a · inbound

Infinity-Parser2 Technical Report cites this paper.

Infinity-Parser2 Technical Report ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T08:03:52.279708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:03:52.279708Z digest=sha256:6f49aae6edf349b4b68273ad1624efe6f75c303570339998daef7aba9e622a44

Observation f52ac92f-c83b-4782-94bf-78228e7463b5 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 277

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:52.571104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:52.571104Z digest=sha256:b329e7066bb0b4f6f3c23f720c37bc5dbe3a8e2750ad770acc57d3849c982f52

Observation d0da804b-bc58-427f-b0a9-b877b11f34f1 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 292

Resolution
unresolved
no resolver link, observed 2026-08-01T04:30:12.392735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:30:12.392735Z digest=sha256:441cbaee89f904cc4edb8841489b0b0e0a5544eed89a2b432804983db57043ed

Observation fc030710-7f3d-4c6a-9c50-c2a7651aeed0 · inbound

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design cites this paper.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.411882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.411882Z digest=sha256:0b535fbc34b1c8fdd86b26bbd82d0ac33385cb5ab623b53e819d242fd1169ede

Observation c14cd281-c78b-4d79-8b87-cbae69fe064e · inbound

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding cites this paper.

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 104

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:30.121328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:30.121328Z digest=sha256:7926756c3949da56dc09462efffd03ad5aab9d3aff2d4c7f169054670515f2e3

Observation ee0a4277-ee5e-4f3e-ba16-6697ec990925 · inbound

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation cites this paper.

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T02:14:09.761937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:14:09.761937Z digest=sha256:448cd9ab8c0c216ba28c58782fa6eb79cc1a1323e6f5a7668a0a173bcf55e4e6