Pith. sign in

Paper Citation Record · LEDGER

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement

As of 18 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2507.01643.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01643 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:53:10.367278Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact3
  • verified fuzzy29
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87a779a4-5d37-4ccc-9100-ca4a1055b45e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.174120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.174120Z digest=sha256:54775bc84c951d0263e6a6f50c8d552a5a79960a5822b2e237d9dc9c263b19f6

Observation f3f8b829-d3b8-4c9e-b49c-7a2c08ed4b7a · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.289384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.289384Z digest=sha256:1abf6eef3f3e330dac640870a7cc77d3bef410a1a3c239690685c807a9740e6c

Observation 53466165-6d22-4a61-b980-3bf3b80cc3d9 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.354949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.354949Z digest=sha256:a772d21decfebaa134410cc9c081ad0b7017f1942b0d9380d9561bd1c24a55f5

Observation 50929551-2aa8-4e2f-ad9f-43148dab93d1 · outbound

This paper cites Qwen2.5-VL Technical Report.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.473093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.473093Z digest=sha256:33739467f534dcdf14bb0b49ecd297a90f0d972e1c58a6977e4393826712fedd

Observation 4edcd2fc-7665-47a6-9aa1-72c48046aeb5 · outbound

This paper cites OCR-IDL: OCR Annotations for Industry Document Library Dataset.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement OCR-IDL: OCR Annotations for Industry Document Library Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.530268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.530268Z digest=sha256:c42a691ab64e28e722bc6b75fba8e4df9f9e6c484a3f5740cdeeb45e3586b738

Observation c7f9a6f9-9427-4f31-81fb-2221bf4d07f5 · outbound

This paper cites Coyo-700m: Image-text pair dataset.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Coyo-700m: Image-text pair dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.606511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.606511Z digest=sha256:ed4bb3c17f5d972eb09ced243f9b90caeafe506319a6daadc726918ce06c0bf2

Observation 4efb48bd-318b-4f97-9e72-279d1f7ed98a · outbound

This paper cites Reversible Column Networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Reversible Column Networks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.746148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.746148Z digest=sha256:928e171b869fbce4f71e017c436d21f0951e70f4451c55e7254131176e921939

Observation df37130e-1b88-4449-b312-380fe9e3d592 · outbound

This paper cites InternLM2 Technical Report.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement InternLM2 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.890756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.890756Z digest=sha256:1a08a0780ce24aab577167300fcac0ac9c74ede6c1cfac2372dee475b75a7f6d

Observation 76b273cb-d5bd-4289-b241-8eed1c6b3670 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.992903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.992903Z digest=sha256:0fc42a38f2f054415eb7220d4b484ac943ef3568973e8f74ff0e99cab6aeeb6f

Observation 6603d20d-9e99-4651-bdeb-78240a99980d · outbound

This paper cites End-to-end object detection with transformers.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement End-to-end object detection with transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.082433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.082433Z digest=sha256:23cf1b8cfe2d718bb621692ebe3c747fa5d4addbc51dd3667f9b940085ccd9f5

Observation 846e9db1-cacd-4b1a-ac71-bd61cb2742e9 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Sharegpt4v: Improving large multi-modal models with better captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.202211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.202211Z digest=sha256:0090df4af690f6051ecccf9098d3bc5fd00384999d9c9c6ca5e64e64590cac20

Observation 36a94428-a52a-4253-9157-6208eaf65a54 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.365902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.365902Z digest=sha256:d745f7a018c9318abb11ca1da568435d29a6513978898a5cf2b337a4517f9cb3

Observation 190292f5-0b9c-450d-b19e-a5b00cc01660 · outbound

This paper cites Vision Transformer Adapter for Dense Predictions.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Vision Transformer Adapter for Dense Predictions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.512069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.512069Z digest=sha256:bcf7a51d9384a629a4e3a522baa098e3eb86961d216b1569d126be9958611132

Observation 3e4e7dab-6dad-4dce-a34a-f5ccd4100a79 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.696027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.696027Z digest=sha256:a2785d7a72e7760b6ad1eab598fe7d4b0fcb75b3b3a356e0637a02e4e0cf0b78

Observation ebb95258-c348-46e4-bc7a-d19ec1053c73 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.828141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.828141Z digest=sha256:7f3da4ea5153e6fa5082b33ac76fced93b18ffc73b7e1cffff4e63194e5faac6

Observation f3351db5-7b83-4ae9-a3f6-4568eec2f45a · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.984634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.984634Z digest=sha256:6e455458499455965d115bec4c20e5b98b9c07e15db5a492080186952228b4ea

Observation 2da8fe5c-44cd-4d40-8531-8ef120e6c04b · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Opencompass: A universal evaluation platform for foundation models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.083800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.083800Z digest=sha256:3ef69620225fa85dbba81cca33c8a8047e228ae14d22e64012718892933b7627

Observation 61c2ff98-d663-461d-8d30-7431fc90918e · outbound

This paper cites Grok-1.5 vision preview: Connecting the digital and physicalworlds with our first multimodal model, 2024.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Grok-1.5 vision preview: Connecting the digital and physicalworlds with our first multimodal model, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.213459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.213459Z digest=sha256:7d47c7cf465f65ea5210aab58efa1f5917927a7994216f5c4a4580ad49b338f6

Observation ed114bc2-240a-41d7-81a9-833564627524 · outbound

This paper cites Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.382480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.382480Z digest=sha256:67c759bf611b6aefbaa629196f039c7334e1aef28836ce31677bf5bd85b2067b

Observation c982de7d-2989-4d96-a4aa-8479303dc6d1 · outbound

This paper cites Deformable convolutional networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Deformable convolutional networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.460206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.460206Z digest=sha256:c51d16281f389b287e279c0f9e431bf33f239cd039c3829197ad89ac94592fda

Observation 9942ddee-4829-4229-80b5-987fe6c51805 · outbound

This paper cites Scaling vision transformers to 22 billion parameters.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling vision transformers to 22 billion parameters

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.556501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.556501Z digest=sha256:ff7d00025a773a76251a165fe7fe7e12cf155f8c0197da0ca036f761d0891db9

Observation 159447a5-105f-49d3-ae2b-f9316f254360 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.653350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.653350Z digest=sha256:cb786ac1634c726b3d31362f9ad44affe05711a3817b3ce3199298ea9c79b58d

Observation 04837be4-82b1-44a1-990f-3694e521845c · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Imagenet: A large- scale hierarchical image database

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.781183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.781183Z digest=sha256:2155c05238742ff5de26ad02b0e69a9c8992df75e47db3e9183a7ecfc10b845d

Observation bd0b9c54-068a-4ccd-8543-fb75e41684bd · outbound

This paper cites Scalable Vision Language Model Training via High Quality Data Curation.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scalable Vision Language Model Training via High Quality Data Curation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.977919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.977919Z digest=sha256:45b4a68041b21d6c2915c07f01b8a405b1088375502574344df25ca9926f6e20

Observation b9a8a1c8-dff9-438b-a88b-5f5ccc91fee6 · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Benchmarking and Improving Detail Image Caption

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.136534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.136534Z digest=sha256:faf68316026544b341f7dfa63accef631dc95b931662cad4cd7660997077d615

Observation dc0f21b2-c60c-4755-8734-e486f0e811b7 · outbound

This paper cites Adalrs: Loss- guided adaptive learning rate search for efficient foundation model pretraining.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Adalrs: Loss- guided adaptive learning rate search for efficient foundation model pretraining

Reference 26

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:53:11.795734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:03.306529Z digest=sha256:a719c91adcc52f7db0e477704f6b6686ea5a41a5222b0a202c8881a9fdc4e746

Observation de786191-9cd0-4b9f-ae4a-b2b735355d81 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.510387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.510387Z digest=sha256:eda1fb5a6afac44341d199ffabe59c93640e9e5a94e99d5e93c0de006dadec27

Observation de626c13-022e-4fee-8d32-926d2276b58d · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.669288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.669288Z digest=sha256:f877095457c43af3f6f02b5c06da11716d8876477fe43f15240a5504dd8fc418

Observation 4d8ec47f-687d-4fea-ab49-a7e1a0a11765 · outbound

This paper cites Scalable Pre-training of Large Autoregressive Image Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scalable Pre-training of Large Autoregressive Image Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.825159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.825159Z digest=sha256:0ba6333488abbf688202349b86fa49d6d722b70520df44e389cfda8a1e8aae63

Observation ece4340a-d08d-4993-894a-dab7d2cd9aaa · outbound

This paper cites Scaling Language-Free Visual Representation Learning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling Language-Free Visual Representation Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.943429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.943429Z digest=sha256:5be41d7410e9f9f439448d5b34bfa36410e9878ddb22de4c83c4ec274bad7e3e

Observation c8b7ce0f-21d6-4ae1-8aca-36773b1fb843 · outbound

This paper cites Data Filtering Networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Data Filtering Networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.057607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.057607Z digest=sha256:e49fd432baba9e97d5552416e6fabe52d155c64b426ce8fe8ff6db0a194b2257

Observation da4f14f6-bca2-4478-8c1e-597d7811d6e1 · outbound

This paper cites Slowfast networks for video recognition.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Slowfast networks for video recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.154213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.154213Z digest=sha256:536093f9bf510b4b3b56bac95c6b8fae2dabf53c3b35e78b8c6375af2793bd2c

Observation 420672e1-1066-4026-9620-448ba0bb8c5f · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.264739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.264739Z digest=sha256:1281ca65b39e1725d8f95b8506b6448136ccb5bfe81c9b8577c23ede8d218588

Observation 591ef97f-139f-478e-a5a2-38dfef806161 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.349270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.349270Z digest=sha256:36ed006de65c3769b82a2e89a37c59ed61cef86f60fb2461076f6e4920d408b4

Observation 974bae22-f802-40f6-a970-4f83a53ba6ec · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.453395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.453395Z digest=sha256:b3929a6ffa6faee577c82bde8ee10311b75ec746c87ed9b48739ff1d3236f670

Observation a1e2ecc9-2e50-4db3-95b5-b9ad56713b35 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.556702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.556702Z digest=sha256:5b40e1c0f6f13a32c78bee86001b496a949357f40642bf66a4919faa13d0d916

Observation 6db352bb-2fd3-4d10-a4e3-c8c098c71fa6 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.659113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.659113Z digest=sha256:c7af4b4e1b4f4aed38fc7d29bdc38cfab3d3b3b059ebc0fcb40d5996e56df655

Observation d9cc9776-df00-41ae-a9bf-2412b3bcc69f · outbound

This paper cites Mask r-cnn.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mask r-cnn

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.765807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.765807Z digest=sha256:d64e05aa2b5124ce8be451330d959d7989d66682f6ab8992038642a9573cb88e

Observation c9578615-533c-4f5f-9c72-b3e17aa48214 · outbound

This paper cites Deep residual learning for image recognition.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Deep residual learning for image recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.889813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.889813Z digest=sha256:ef8ecfc760d22e0281fed02e8a6d97fdb8260dfeca929f68fdb19776189c633a

Observation 7749181b-17cc-4cf8-844c-29f86273e8c5 · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization,.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The many faces of robustness: A critical analysis of out-of-distribution generalization,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.975078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.975078Z digest=sha256:965084992440c60be1481e34e45932f4cf095b643e571b2364fe99d662321dab

Observation fa1a1a87-8ac0-46cc-8df7-fc82cd529023 · outbound

This paper cites Natural adversarial examples, 2021.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Natural adversarial examples, 2021

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.049397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.049397Z digest=sha256:e8d63ffc8059fba8b9fd37eb77f77e1130fbc370590f9a7d38f568afe7460c9e

Observation f0de58a4-8897-4c39-adb4-2a867911ca71 · outbound

This paper cites Squeeze-and-excitation networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Squeeze-and-excitation networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.144799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.144799Z digest=sha256:bd0d5c77a67e30c7afa820d32b64bd7a7c50b29b7c2e5a0aa8e512f58b9979c4

Observation f611ea19-3a0c-48f2-88f7-795668443c75 · outbound

This paper cites Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.239887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.239887Z digest=sha256:93dcb3c43a461d8ccd4f10a99486f90a52f2476386ec76732fd514a9a166a812

Observation 51a87678-09b0-4b3d-83b4-a3a6103152c5 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling up visual and vision-language representation learning with noisy text supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.355096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.355096Z digest=sha256:e02d302449a64ef8b37e7bc8cd535baf14c0275d7cdc24649ea262dfc1c56cb3

Observation e5eab2a7-4141-450a-ad36-36ab8876a996 · outbound

This paper cites Deep visual-semantic alignments for generating image de- scriptions.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Deep visual-semantic alignments for generating image de- scriptions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.440840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.440840Z digest=sha256:20a19d1ca5ac7c8a5e7c4afa8bc4952fcaaa6828d1769c847de352049e3f6801

Observation 74f37a5d-2c5c-4385-b53d-f0865f2b4502 · outbound

This paper cites A diagram is worth a dozen images.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement A diagram is worth a dozen images

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.976874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:05.499872Z digest=sha256:847b4c0204643e2125ca6d07a3accafa6278ed460746a9277582fd06d607eb2c

Observation 1b36cab1-6f83-4bf8-8687-e77075edc616 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Lisa: Reasoning segmentation via large language model

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.815444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:05.571684Z digest=sha256:3a537262f963b4c1655806df72506c6dd0e1a5fb6ddfef3c7c111bdebbc02aef

Observation 827c4573-8d18-4ff2-bb34-487ed97466f8 · outbound

This paper cites Building and better under- standing vision-language models: insights and future directions.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Building and better under- standing vision-language models: insights and future directions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.676412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:05.672563Z digest=sha256:00d32fd9f057cfa8d1afd1ce350498fff30e23c3113ef296f074d2109c695979

Observation 0044f049-24e4-4e83-863d-0ba8b5f52189 · outbound

This paper cites What matters when building vision-language models? Advances in Neural Information Processing Systems, 37:87874– 87907, 2024.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement What matters when building vision-language models? Advances in Neural Information Processing Systems, 37:87874– 87907, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.515187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:05.761968Z digest=sha256:04c521183f1c1da2c72f5a523eee9ff23ad9a7b26d652b2c110a4d6695db889c

Observation 8e3d7bf0-8038-4b4e-9817-32ed004715e8 · outbound

This paper cites The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.867758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.867758Z digest=sha256:1d7bb5605de6f59c214e75b9714068232beee351195ac872609894a2723f75de

Observation fd89d39e-9699-41a9-9007-b31e3cb67f5c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement LLaVA-OneVision: Easy Visual Task Transfer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.966643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.966643Z digest=sha256:06673025641cef6fe3f909292c750194ef5de524a60fa5a85524c0b376f531cd

Observation 8edbc22b-7bf8-4438-9569-05985e23b1d5 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.041336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.041336Z digest=sha256:a88daf26ba1b3b077831adb9bacd7ea370362e43bb7e362f941c198fb8d7a145

Observation 59f4a2d1-fba9-48e7-99af-b69dbd5588ce · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.332980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:06.099402Z digest=sha256:96f235d89fe302efeca44ca396957689944c6ff259de6ddf31d090d8d549019e

Observation ebfdf35a-1998-4dcc-9cef-7e50202a6687 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.204273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.204273Z digest=sha256:53a32846cd4d6485bda863412bef3730866fe0379cc208dd7496efa90aa50998

Observation 333b4891-09aa-4b17-bf68-fae1118550be · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Evaluating Object Hallucination in Large Vision-Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.310118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.310118Z digest=sha256:74aea140502c7ba1996110e51aa0f09007a519a314de4fad4315a9e1f83800a8

Observation 0283545d-c372-412b-b1ce-0e4581157c35 · outbound

This paper cites Improved baselines with visual instruction tuning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Improved baselines with visual instruction tuning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.185468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:06.413431Z digest=sha256:99673e1c3cc0a350cd748628e2bafe6ed64f7765bfdcc898ed47e679c56d85a4

Observation 4a2fe0a4-b236-4c80-8df4-5eaf3dd23463 · outbound

This paper cites Visual instruction tuning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Visual instruction tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.961535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:06.494248Z digest=sha256:336a605bce315ee3bc00bcb3bec9737411de567ff40fd6f5d7725da98841e574

Observation a5a01265-787c-40d3-adc5-597f0b495cdd · outbound

This paper cites Muon is Scalable for LLM Training.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Muon is Scalable for LLM Training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.548354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.548354Z digest=sha256:f1bed9ac277fc61d16baa5e4aab251a7907ef7f152ea714b9578940299d32a45

Observation ae89d1a8-3bec-4090-ba17-80b3218be92e · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.741340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:06.649180Z digest=sha256:c6b3ff9b31f05e1b2a7673031e4399a51bbd2aa938fb72b44f023c34a5e08ec1

Observation 345b18e4-314a-4e36-8430-5d72da46ed77 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.552804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:06.750645Z digest=sha256:a9b55308359c52d6c479ff6669c36d072d5f44de5f28b9d862bb29f99593ecd7

Observation 2df5e4c3-e566-4e69-b955-7148d92b9383 · outbound

This paper cites Decoupled Weight Decay Regularization.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Decoupled Weight Decay Regularization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.868876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.868876Z digest=sha256:0d6c23ae7f57838cc04365ccae5ab5cb240de086afc0d8b9bb9d9e19b214a5d1

Observation e37a86bd-b7da-4470-9ad4-db79d3c5e662 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.917722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.917722Z digest=sha256:d110f6abc6359590ee0be434e061642a374444abf333f8f6181805c5bb788996

Observation a29d00c8-d994-499b-a20e-78726d285db5 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.020742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.020742Z digest=sha256:dfdba047ede2e74f713bb8227a8676588a64d32e672e999cd6c41d6ea7eefc3a

Observation bbbcbdee-0cf1-4a78-ac00-b63f7b284021 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.393314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:07.120075Z digest=sha256:b593e80c48a63b8d1230647872728f197f772fa4ca4cd5ce5a7bc0f2bcc36486

Observation ed625676-9ef6-457f-8452-fdcc4955ba73 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.221154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.221154Z digest=sha256:ce0b0f30ec9a5c3c8c2147266cb5931b76022f1308d4ceb3c207c5929ffe889b

Observation aea5201b-43bc-4e53-925d-97fcc6f8ebe8 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.312239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.312239Z digest=sha256:3ac8db06aec1b0f75d8d4e0deec506ac8945095e07f7bd94c061ac36fe0534df

Observation 831dbd81-8405-4b6e-8b64-aefbf16eef2b · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.233277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:07.419963Z digest=sha256:8ee8ba686ecf362055b0c9d68d06e8606d13685d4fd9a0d032e71f38c9809a3e

Observation 63e2f7dd-ad86-4244-a702-b01fb1db3d3a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.511190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.511190Z digest=sha256:4316637b791566af1d32ba2c27bdea30fac7950479ada739158e011695853859

Observation fde217ee-d772-4282-b2c7-64e7d15c94d0 · outbound

This paper cites Infographicvqa.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Infographicvqa

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.058601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:07.611803Z digest=sha256:c285566edbdb61dcc8875c8a1bce9062c85e75378df372d13257b2e1c2d93265

Observation 01ba0488-e905-41f9-87f4-b2895e1b06cd · outbound

This paper cites Docvqa: A dataset for vqa on document images.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Docvqa: A dataset for vqa on document images

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.878279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:07.684217Z digest=sha256:997698f30f3c04b76cec5bbc278f1490663aff916c2965f2bb6e1aba63078825

Observation 1afdc95a-9888-4be3-a8da-0751921568e7 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ocr-vqa: Visual question answering by reading text in images

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.694685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:07.797962Z digest=sha256:f3c5b2a7fc76e5e33b634da7c8609576336079e3937554bbcdcef2ba6a54bc0a

Observation d7e7995f-9a97-4699-b4fa-2833ed44db52 · outbound

This paper cites Introducing chatgpt, 2022.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Introducing chatgpt, 2022

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.509068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:07.848469Z digest=sha256:1b32dd532b1d40da37ef9156ac4899d12798c178ff85baefd9ef422e3ccab071

Observation 65d44a07-4e93-4451-bfc2-fb2aad7ba817 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Learning transferable visual models from natural language supervision

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.329003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:07.959684Z digest=sha256:43dd7590b5391dd6d7ff5691ee704db9fdbcdda2d120411bc0d54cc4108f6a23

Observation b61bca55-9d38-466c-a543-7801a7c5a1d9 · outbound

This paper cites Zero: Memory opti- mizations toward training trillion parameter models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Zero: Memory opti- mizations toward training trillion parameter models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.195330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:08.067283Z digest=sha256:19b505d3c1ce157855514d6552339bee4fd2a6db6174635e5997a755a3b8922c

Observation cbcbd198-0e40-451d-bb14-ea29fa46ad0a · outbound

This paper cites Do imagenet classifiers generalize to imagenet?, 2019.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Do imagenet classifiers generalize to imagenet?, 2019

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.040246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:08.176323Z digest=sha256:9118e9c4415942c59aa6f76e8b65ff8b4a74da48600aaf8679034eaefddfa71a

Observation 67b0689a-37ba-48e2-97be-68b05328c0e7 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:08.277308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:08.277308Z digest=sha256:a05bb27bbcd122e81367c01c0dc34c4ed0f63349d0b512410098962b60ce8fb1

Observation 9f4b5191-474d-4218-a490-47ba5b5a4763 · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:08.389462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:08.389462Z digest=sha256:3d4ac88d8c204120ab9deab277564fecdec55213e6f81c914a0648896ac654c2

Observation 129c05bc-c99d-480f-9f80-c2b7fbf1fc79 · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Hollywood in homes: Crowdsourcing data collection for activity understanding

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.840234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:08.498441Z digest=sha256:25e20aa2c94124e4d8ddf0799e846990dd3e357f404ed76cb37e87d05ee505e0

Observation d499bfdf-0ff9-40e4-b799-59f07d5dab77 · outbound

This paper cites Towards vqa models that can read.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Towards vqa models that can read

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.673354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:08.610364Z digest=sha256:ed6a2d657c247fb32e392393e5a02f231ca05fcf875afcbcbc50f90e80d9a259

Observation 446b6f94-7e19-4f2a-bf2f-8b0c05c62892 · outbound

This paper cites Kimi-VL Technical Report.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Kimi-VL Technical Report

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:08.679854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:08.679854Z digest=sha256:b2bb61fd0f062d2afc72b471f4543fdf9d0ccabdfa3743c513ea806a177f40da

Observation e03a2961-db65-42c8-b3da-05e24ef9ad89 · outbound

This paper cites Qwen3, April 2025.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen3, April 2025

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.554335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:08.740655Z digest=sha256:d22602dee5b8453d67f935dade7822431c142da8f3e2ea8e9398a41a93a5dc64

Observation fcb05f98-913a-4806-b5f8-fbd495febd22 · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:08.830124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:08.830124Z digest=sha256:a096ddbec3d44c2576cf8b202f5e48764543c41cdc8f9d266a92c6735bfe3ea1

Observation 366e44a7-412d-4023-a12b-e1e04d53176d · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.392724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:08.912731Z digest=sha256:b5b5492efda2cd99e81ab97d89ed0dd6a80ec1b0dd26fa5af0c732247f4bc0a3

Observation 01b2c095-0adc-41bd-b71b-925321588093 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.011049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.011049Z digest=sha256:40256c60b0ca69365019fd1b1c4d58eeb2b0137da20ed77dcac0020a23637e36

Observation 84feef7a-bfbd-46ec-adaa-a9571bf3c6ae · outbound

This paper cites VGR: Visual Grounded Reasoning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement VGR: Visual Grounded Reasoning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.114205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.114205Z digest=sha256:ede90b3d71b73a63fb5952b0262fa5c434f7d155b9736f7096834a21f09cf701

Observation 68de02d4-03e1-4ab7-b0da-48d7c9ffbfc4 · outbound

This paper cites World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.193598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.193598Z digest=sha256:b6e21566fa5bb660831e231ef0e039407dd93592facf4b23b486c7b48aaf22c8

Observation 99cb87b7-ae69-4d36-a51b-54a0499edfd3 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.284570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.284570Z digest=sha256:d9a585a40d88c3cf760b5dd3f39fd1a1315880170ea61c2aa6fb8986fec8953f

Observation 34bcd4d3-5a97-49c0-87ce-d15e2db08641 · outbound

This paper cites Pvt v2: Improved baselines with pyramid vision transformer.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Pvt v2: Improved baselines with pyramid vision transformer

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.252292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:09.384776Z digest=sha256:4094a6da7b616e1cee3506208be67f6c31877743c168bcc11d1aec68031acba3

Observation 275553a3-2d36-4427-becc-762e4e7daf42 · outbound

This paper cites Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.464295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.464295Z digest=sha256:7148b2c8e29000da7a60bfbe328d7aaed4157f26e1f483bac8924bf0121b9308

Observation df4f8766-c159-4818-9a53-645cf33e705d · outbound

This paper cites DocStruct: A Multimodal Method to Extract Hierarchy Structure in Document for General Form Understanding.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement DocStruct: A Multimodal Method to Extract Hierarchy Structure in Document for General Form Understanding

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:53:11.128260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:09.562208Z digest=sha256:e5cb41a6bed72b0109b586e5173c1e49af49891facc9eca284caa07a7304c506

Observation 42533c59-304e-48d5-b8b2-58c198e861a9 · outbound

This paper cites Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:53:11.031292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:09.643301Z digest=sha256:896387bc6a7278ac8a49e03be7b061bf65e2995df60a77e46ac84b4d2dfca683

Observation 21a80b97-285b-427d-8e2a-9042d2fba491 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Convnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.058543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:09.735785Z digest=sha256:e1a502f34e62f4753a4e3f2af76574764ed35ab895225fc9392c2c1b11dda465

Observation 665e9818-9536-480e-a7e3-16e73fe9a998 · outbound

This paper cites Seeing the image: Prioritizing visual correlation by contrastive alignment.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Seeing the image: Prioritizing visual correlation by contrastive alignment

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.875815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:09.815315Z digest=sha256:13a3377cabd88407ac65ac76f13c88c9aa4d19dd5f34d26b397f843ce2722274

Observation 6e00b62d-a9a6-4216-aa41-1eb5feb23ef1 · outbound

This paper cites Aggregated residual transformations for deep neural networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Aggregated residual transformations for deep neural networks

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.734487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:09.911243Z digest=sha256:97cbc9b17fb1fc99321789aba4afd74f101b3231a9f6233c8ef6f38fa4d7054c

Observation 3da9d2ac-8853-4031-9f32-69a7b568ddd4 · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.990394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.990394Z digest=sha256:dafedf20285010c0bc2f8e531412f729bf5045af31c125e31a45f60e5520673d

Observation b775d07d-d57c-4e11-ac09-ae766d755e34 · outbound

This paper cites Qwen2.5 Technical Report.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen2.5 Technical Report

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:10.097531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:10.097531Z digest=sha256:714293771078623f42e0509b43098e08e226763c5c797b8b5c12199edb3043aa

Observation 7de236ae-333d-48c8-ab3d-74acdbdccaac · outbound

This paper cites Pediatricsgpt: Large language models as chinese medical assistants for pediatric applications.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Pediatricsgpt: Large language models as chinese medical assistants for pediatric applications

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.585192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:10.152994Z digest=sha256:9d863fde9e25504e0adad172562305ac618a185863382367a379e72482396f22

Observation 431a2d12-4355-440a-81a5-9cd058174da8 · outbound

This paper cites Improving factuality in large language models via decoding-time hallucinatory and truthful comparators.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Improving factuality in large language models via decoding-time hallucinatory and truthful comparators

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.432758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:10.220443Z digest=sha256:da0246d40bf76a6e97929a5cf21f0d47ce305a033c1a055d61bd9b838696f65b

Observation 841289a6-f86c-493d-9105-fad874f1c0bb · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:10.285746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:10.285746Z digest=sha256:431d966f53c82c4af8b0abf9e4216fbaee3fd34e7269cea18b3dd78b8050c173

Observation 559193a3-2037-4663-b025-4d310ebbb974 · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.313580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T20:53:10.367278Z digest=sha256:28b26a1ab6c3f04cf1e13841c24a8b77c62ec51d29655747cf62e854157b5a67

Pith citing papers

No inbound Pith citation observations are available.