Pith. sign in

Paper Citation Record · LEDGER

Building and better understanding vision-language models: insights and future directions

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2408.12637.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.12637 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:49:10.837634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:39:39.660194Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0f304a0f-a3ca-4e71-9e9a-db73504ec9c7 · inbound

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark cites this paper.

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark Building and better understanding vision-language models: insights and future directions

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:51:48.342889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-14T00:51:48.163349Z digest=sha256:125becaad8bc817ba50316b436fa7efc3c1ee28cfa578b042755fc1b77b87f54

Observation 33e2b272-aef7-43ed-a3d4-b2fd70c979bb · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Building and better understanding vision-language models: insights and future directions

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.630563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:3cefeb64c06a14c46044d9e7481a8511aebdb13f97956b3e30a78c3cb05a7344

Observation 92ab786f-14a3-4d3f-91b0-43fa501d2784 · inbound

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models cites this paper.

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models Building and better understanding vision-language models: insights and future directions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:06.804474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:06.804474Z digest=sha256:076b67b4fe0935f2a71e7be58650f7bc5d32199b2876a614a36b347234aae114

Observation e481b500-b1fb-47f6-a568-685b8dc5b950 · inbound

VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models cites this paper.

VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models Building and better understanding vision-language models: insights and future directions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:36:35.917708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:36:35.917708Z digest=sha256:6de6e4413650b84caafe37534f436313265970096fba841a709b882014281d6d

Observation 2ca4f020-57c9-4508-95f9-499a99c5fe63 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Building and better understanding vision-language models: insights and future directions

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.129247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:f373a88d25112b437174a982db060c0fdf3343ddc4c8388325f1b7645355881b

Observation 7fe45a35-7a91-445e-a1e1-5e349c1c25dd · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Building and better understanding vision-language models: insights and future directions

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.029371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:4bae2025f059e74291e87787c4a285c76c4cf4b6f346bb6da6c246f98de7dd51

Observation 020c5bb4-6836-4b3f-9270-e1dfc5e9856d · inbound

Apollo: An Exploration of Video Understanding in Large Multimodal Models cites this paper.

Apollo: An Exploration of Video Understanding in Large Multimodal Models Building and better understanding vision-language models: insights and future directions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:11:10.490240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:11:10.490240Z digest=sha256:b99e2e3e2f88b67abc125ec426ed1081237e7bf384d3ff768290f867a62f1e3c

Observation 4bad5c29-2f67-490c-be1d-f7bb08f05aa9 · inbound

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation cites this paper.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Building and better understanding vision-language models: insights and future directions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.131404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.131404Z digest=sha256:99ed3d5ba5c0cc1d6f0b7e00d15bd3a613b275ed69e6485000b1deea8dcaab2f

Observation 4662635b-f70f-440c-bc44-1957f06b8306 · inbound

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining cites this paper.

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining Building and better understanding vision-language models: insights and future directions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:33.498202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:41:33.498202Z digest=sha256:4b01b2347edd8006bebeb4419abdb49d8f026204c0bd1978d44dbd3f7593fa87

Observation 627db8c1-9b90-473b-b6c5-3bff1107c01a · inbound

Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation cites this paper.

Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation Building and better understanding vision-language models: insights and future directions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:14.213819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:14.213819Z digest=sha256:59a4677a43c394a7edd05b6710079c6d1037e6058302628a16af2825d5731af8

Observation c7969d33-d702-405b-9d60-6ab0036c8e96 · inbound

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark cites this paper.

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark Building and better understanding vision-language models: insights and future directions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:25:47.061351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:25:47.061351Z digest=sha256:d307abd13bc3bd221840479178773d0a288da74e2b7ee17fe98196ac2a3a23e9

Observation 5e4f1832-6a6e-4fa8-8245-cedc8457168b · inbound

MSTS: A Multimodal Safety Test Suite for Vision-Language Models cites this paper.

MSTS: A Multimodal Safety Test Suite for Vision-Language Models Building and better understanding vision-language models: insights and future directions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T19:27:42.242753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:27:42.242753Z digest=sha256:1d15944d7832efc37b58bce561e437751e45e1c476906c2d005424590a77128a

Observation 39a1499e-823e-4baa-96a7-1963287bdc7f · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Building and better understanding vision-language models: insights and future directions

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.402880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.402880Z digest=sha256:79a9dd5ef7433984e7d4376c656052ceb6d3a3a66da6647b7c1069ebc2ea0bb5

Observation 603ad2fc-c9ac-432a-963e-be588269380b · inbound

The Impact of Persona-based Political Perspectives on Hateful Content Detection cites this paper.

The Impact of Persona-based Political Perspectives on Hateful Content Detection Building and better understanding vision-language models: insights and future directions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T19:15:41.568116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:15:41.568116Z digest=sha256:a1f002daf6d119571bf234bda7ba3888dd1689959024cb5aadb92cd8523b5414

Observation f78ef67f-3e12-4473-8176-d60d129373ad · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Building and better understanding vision-language models: insights and future directions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.706584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.706584Z digest=sha256:7015776045c3dee4643c6880f23a60e3697ebc35b3e77ba4ee83e476fa57d10c

Observation 10b5a689-0ad0-401e-a7ac-e47ca724df58 · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models Building and better understanding vision-language models: insights and future directions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:23:51.725976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:06c6df7aebd4c6bbf752a776434e5d3bb0b5f1a86ab0e97cf1c44fb8215b1db6

Observation cb1f77ae-97fa-40b4-93a4-fe4f374436a2 · inbound

Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency cites this paper.

Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency Building and better understanding vision-language models: insights and future directions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:49:10.837634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:49:10.837634Z digest=sha256:ef79841c8aed8c7dd16617fbe19761baa617b06dd26d8a12e557cf2bbd3afb12

Observation cfe53a72-f40e-43b5-b681-e17e28e4a155 · inbound

R^3-VQA: "Read the Room" by Video Social Reasoning cites this paper.

R^3-VQA: "Read the Room" by Video Social Reasoning Building and better understanding vision-language models: insights and future directions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:41:34.161930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:41:34.161930Z digest=sha256:5702dd8fbff3a6f257c38a5e5de5609b55ad35bff38897527d28473f63894745

Observation 4dd34dc0-0697-4789-a243-673797ff84f3 · inbound

FG-CLIP: Fine-Grained Visual and Textual Alignment cites this paper.

FG-CLIP: Fine-Grained Visual and Textual Alignment Building and better understanding vision-language models: insights and future directions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:58.708306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:58.708306Z digest=sha256:2b6749dfae3710bc8d7da20a2a47b671ea86bc269b9c2f395b6b9683b5844c45

Observation 2021fc53-9986-495d-b081-c67fdc160c3b · inbound

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings cites this paper.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Building and better understanding vision-language models: insights and future directions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.673554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.673554Z digest=sha256:eee9b5b464418c8655c7789e706797f0a48585b2ffda00b3a3a5cbb2c38505ca

Observation de154433-0b74-42e1-91f1-83e5f4c6ba35 · inbound

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings cites this paper.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings Building and better understanding vision-language models: insights and future directions

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.743202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.743202Z digest=sha256:0532e3768aa549992a2c09036bfd7d0719a31c760b0a046a9b70b4a4861c339b

Observation 8251c337-49a0-4d1b-ba3b-3da9da9008f3 · inbound

Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data cites this paper.

Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data Building and better understanding vision-language models: insights and future directions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:26:12.795081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:26:12.795081Z digest=sha256:7b20502f60e6e91880ab6fef17abe2d91d8ef838eb48e0774de4d6ffebe744b4

Observation 280158b9-84d2-4448-93ab-a08974e26520 · inbound

Vision-Language Models Can't See the Obvious cites this paper.

Vision-Language Models Can't See the Obvious Building and better understanding vision-language models: insights and future directions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:39.601722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:39.601722Z digest=sha256:90ac6e907a1d5bec5f7eeb42560e3fdc24a4d904e9c1a45a6a3a6310a2059eae

Observation ed95fb85-610e-4847-a8c5-56b3496cd2ac · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Building and better understanding vision-language models: insights and future directions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:01.585192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:01.585192Z digest=sha256:c66b5e2cd0bc330a383524bf7b1178ae2fc394c182360e2f289522601e29af4b

Observation d39e0ca0-c515-4ee9-844a-d43e7d6201cd · inbound

LMM-Det: Make Large Multimodal Models Excel in Object Detection cites this paper.

LMM-Det: Make Large Multimodal Models Excel in Object Detection Building and better understanding vision-language models: insights and future directions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.196489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.196489Z digest=sha256:c1612323725ecb9d70935577395af56ca5a18e52f564fe77ba07de35f846108e

Observation e254639f-2e33-484c-87c7-28b6ff9e545a · inbound

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding cites this paper.

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding Building and better understanding vision-language models: insights and future directions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T22:10:08.145820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:10:08.145820Z digest=sha256:12af4a90baa0b139ec5f02df50409fccf83eeeb4d1e087e8ddace4e46cc0169f

Observation c3a56c70-296c-42e0-bd61-7193d84befbb · inbound

InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information cites this paper.

InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information Building and better understanding vision-language models: insights and future directions

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T00:26:56.175974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T00:23:52.475744Z digest=sha256:4bee99c6c915a72cef7e569fc8b360f9180651f7357db3472079b301f190a510

Observation 887a9bfd-dcdf-4a24-99f5-8af7eb53e7af · inbound

Measuring Epistemic Humility in Multimodal Large Language Models cites this paper.

Measuring Epistemic Humility in Multimodal Large Language Models Building and better understanding vision-language models: insights and future directions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:39.205111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:39.205111Z digest=sha256:c95dc65d303aa250555eb74293b5d945552c31d0d0298c12eaf40a3b5fc5c7d5

Observation 73d926d0-48cf-4725-bec6-a78b63643ab4 · inbound

MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal Models cites this paper.

MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal Models Building and better understanding vision-language models: insights and future directions

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:42:38.772215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T13:41:37.594782Z digest=sha256:5c1a6ea96215ee8d6a4df057ddf59ae35d3805c046e94619b0fefcddc2ad1ef4

Observation d11ee3f7-00d1-46f5-8818-e8275eecb668 · inbound

Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models cites this paper.

Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models Building and better understanding vision-language models: insights and future directions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:41:50.086157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:41:50.086157Z digest=sha256:34692c5756db54e1f3e9ae4a425cdb9d98b58567b4d4acb72470f0a936cfed83

Observation d1a1b7cc-4327-4d06-8e40-7aaa6efa2abc · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Building and better understanding vision-language models: insights and future directions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:54.248497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:54.248497Z digest=sha256:2ad80399cb850abf44b600613df7b4624810475446116fa10774a6e434503e92

Observation 8eae8d05-7347-43e8-ae0c-00d41ee2405e · inbound

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning cites this paper.

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning Building and better understanding vision-language models: insights and future directions

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:17:26.618597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T06:13:33.315525Z digest=sha256:a5710cb08a8cd0cbf34c7ed1751d2b423d979ba03ed5f3bc69176c297088eb8d

Observation cb705c73-6cf4-46f0-ab9c-b9f4c94fd856 · inbound

Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models cites this paper.

Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models Building and better understanding vision-language models: insights and future directions

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:20:59.281336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:07:13.367863Z digest=sha256:163db119d6359abba08bed9c5a196f5ac7c06a0c27fae8d554b9b0847f5a3045

Observation 8e96df26-dce8-45a1-946b-530100747880 · inbound

DenTab: A Dataset for Table Recognition and Visual QA on Real-World Dental Estimates cites this paper.

DenTab: A Dataset for Table Recognition and Visual QA on Real-World Dental Estimates Building and better understanding vision-language models: insights and future directions

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:53:04.999998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T08:33:00.988234Z digest=sha256:42da9afb0c387d99a3e963cd8840df5dd9943316b3e218ed733b2d9c544a9a74

Observation 1100dfe0-713e-42a0-aa91-4df2231482c6 · inbound

PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation cites this paper.

PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation Building and better understanding vision-language models: insights and future directions

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:56:11.157107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T05:53:34.007500Z digest=sha256:decf722df3037a42e3117c38b15be1e7cf8a4e86679c05b945b616f189e00c69

Observation 5357719c-62c0-442c-9aa0-19b21d897cfc · inbound

ZAYA1-VL-8B Technical Report cites this paper.

ZAYA1-VL-8B Technical Report Building and better understanding vision-language models: insights and future directions

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:23.510227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T01:15:16.607346Z digest=sha256:63b9cdc07133f5fcbdaf208e83852bf96ea587f54ae92efa33d198a0b0bbb142

Observation dec68707-2fe8-4e97-ac41-1169b7d858c8 · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone Building and better understanding vision-language models: insights and future directions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:57:09.560253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T02:52:43.674969Z digest=sha256:22ada1a7654d58b478b747d4c7433b390eb1c3d7633342b90777a049a75299e1

Observation 878d8f83-82c1-45cb-a155-b8526dfe88ec · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone Building and better understanding vision-language models: insights and future directions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:29:28.614421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T21:28:37.680681Z digest=sha256:577ba0c13539034bc2d9113c53b90dfa1011864dd2557f172676865139132127

Observation 5b374f96-111f-4955-a00a-e20b6dff8c51 · inbound

VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding cites this paper.

VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding Building and better understanding vision-language models: insights and future directions

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:34:02.594785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T22:24:20.787671Z digest=sha256:35d1023710b84ed751398264dd79d2a404a7dc8a1b3c7fb6fd1c420af035bf47

Observation 6c79169a-4d02-4fbe-8b40-9d751204ce42 · inbound

Zamba2-VL Technical Report cites this paper.

Zamba2-VL Technical Report Building and better understanding vision-language models: insights and future directions

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.658133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T22:34:20.970856Z digest=sha256:6cc863f733ec8010556098069c822ad312a5f29731372e48c7bdcef695fd28d2

Observation 8398975b-134e-490a-8000-37d2a02181e8 · inbound

Stellar: Scalable Multimodal Document Retrieval for Natural Language Queries cites this paper.

Stellar: Scalable Multimodal Document Retrieval for Natural Language Queries Building and better understanding vision-language models: insights and future directions

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:39.661854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T15:49:49.067128Z digest=sha256:bf68b0cea8cba4a1c88408510e1d69f01a3aa16f88fa70bda12ee92b46581674

Observation 13e6751d-735c-4c94-8cd0-5ab2e8566b25 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Building and better understanding vision-language models: insights and future directions

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.490405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:66dc70ce38163239e289f356639ebfe3266288f81ff0a6efbe0dcaac90be5401

Observation 38ab0453-c911-4761-86b7-c1cdacef7631 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Building and better understanding vision-language models: insights and future directions

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:23.854434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:2b4a2b378771389d599b3d81c34f2dd68541e1380b4d50b6c63a4f7355363c00

Observation c636bb41-f0a3-415f-b6c6-15f95f16eeb9 · inbound

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement cites this paper.

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement Building and better understanding vision-language models: insights and future directions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T08:37:42.020482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:37:42.020482Z digest=sha256:19c0016dd5d2ff0ead349f6587247acfdbfc5bf81441da2565c09ebd67fd8a15

Observation 0cfbcc96-ccc5-4fce-b4bb-fc008816a546 · inbound

MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents cites this paper.

MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents Building and better understanding vision-language models: insights and future directions

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-30T13:50:41.061598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T13:50:41.061598Z digest=sha256:4623d21894645fdce8fd3ab9c5d6702cb438b1624d7df01fc357ec028f1b2838

Observation e363be95-6a85-4242-aa43-7a3606fa7fdd · inbound

Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution cites this paper.

Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution Building and better understanding vision-language models: insights and future directions

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-12T00:43:45.298580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:43:45.298580Z digest=sha256:c02c25f27e84264956baab4e7b48defb2f3c886c872edf432173b8e2b7eb931c

Observation 22d9cf6a-c6dc-4053-90f6-c999d79303cd · inbound

DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation cites this paper.

DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation Building and better understanding vision-language models: insights and future directions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T20:26:25.681332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:26:25.681332Z digest=sha256:19aaf0cdbb2e87cb6747fe4759d0af89297811a96c0159b0e78047ce21af08be

Observation aa545273-8786-4b5c-bc0e-179698f21fa7 · inbound

Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ cites this paper.

Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ Building and better understanding vision-language models: insights and future directions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T10:59:51.884916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:59:51.884916Z digest=sha256:ac6e6ad816dc41b39b59740d7ce2ec7a62828f1ff234b694b0fdbd1318b27070