Pith. sign in

Paper Citation Record · LEDGER

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement

As of 7 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2507.01643.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01643 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:53:10.367278Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact3
  • verified fuzzy29
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87a779a4-5d37-4ccc-9100-ca4a1055b45e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.174120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.174120Z digest=sha256:bc792486d72bb816fbeb55770904bb9334ddd6483e2204c93ad9a7d301c11818

Observation f3f8b829-d3b8-4c9e-b49c-7a2c08ed4b7a · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.289384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.289384Z digest=sha256:5e445d0cb1a5130e580fe426d6cd34e02f96c2f9c0540a1d17b4fff5a66bc93f

Observation 53466165-6d22-4a61-b980-3bf3b80cc3d9 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.354949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.354949Z digest=sha256:b1b96125337f45d3a250326e7675741414ab21696eb6eb9dc92beb5619e07084

Observation 50929551-2aa8-4e2f-ad9f-43148dab93d1 · outbound

This paper cites Qwen2.5-VL Technical Report.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.473093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.473093Z digest=sha256:d90a0a75d04250425dbef8a9727888185bf9f22b3f26199e943ba08ac4a0ea6d

Observation 4edcd2fc-7665-47a6-9aa1-72c48046aeb5 · outbound

This paper cites OCR-IDL: OCR Annotations for Industry Document Library Dataset.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement OCR-IDL: OCR Annotations for Industry Document Library Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.530268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.530268Z digest=sha256:03de2b2e0fee93d63dc9aec2f26e0777c1ed46d02f4aa26a4fee9a7fbfcfb892

Observation c7f9a6f9-9427-4f31-81fb-2221bf4d07f5 · outbound

This paper cites Coyo-700m: Image-text pair dataset.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Coyo-700m: Image-text pair dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.606511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.606511Z digest=sha256:0e46d11f9ca5e607da9fc3407a0494c891df78ee8d2aa8003883e1cb775268a0

Observation 4efb48bd-318b-4f97-9e72-279d1f7ed98a · outbound

This paper cites Reversible Column Networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Reversible Column Networks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.746148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.746148Z digest=sha256:9b2b6abd344f4c6523bb52a2b32b0269bfd717b74a3aaf6c8807a8e273533897

Observation df37130e-1b88-4449-b312-380fe9e3d592 · outbound

This paper cites InternLM2 Technical Report.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement InternLM2 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.890756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.890756Z digest=sha256:55cfd9821d53f054c57e930203d9a82bdd4e801a07c3c04dbe42fdc1497590ca

Observation 76b273cb-d5bd-4289-b241-8eed1c6b3670 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:00.992903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:00.992903Z digest=sha256:3f70a16397aa21cdc551f13fff01a076516c361ea21797d1819be35fbc7117c0

Observation 6603d20d-9e99-4651-bdeb-78240a99980d · outbound

This paper cites End-to-end object detection with transformers.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement End-to-end object detection with transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.082433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.082433Z digest=sha256:a0491b2e4f6e20997908c3aeebaeb1816e876ffb0038c066426b69086ef20ce7

Observation 846e9db1-cacd-4b1a-ac71-bd61cb2742e9 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Sharegpt4v: Improving large multi-modal models with better captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.202211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.202211Z digest=sha256:f0da95328a2af3cf763c727b9c2a77dceed7193eace0a56709f527705d52a4ea

Observation 36a94428-a52a-4253-9157-6208eaf65a54 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.365902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.365902Z digest=sha256:068f49c8945f3c946207819dd1d190ae191da122cbbf07f2821a52c07c5edf9d

Observation 190292f5-0b9c-450d-b19e-a5b00cc01660 · outbound

This paper cites Vision Transformer Adapter for Dense Predictions.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Vision Transformer Adapter for Dense Predictions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.512069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.512069Z digest=sha256:c8df049c9f7aee330ad3436a519c8a273910f4dffc9474730cb78d7c0579dab5

Observation 3e4e7dab-6dad-4dce-a34a-f5ccd4100a79 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.696027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.696027Z digest=sha256:ecbf8ab7903eafb5d91b35c5fcc25fb9212b3c9d15906b324d677fa942a86e71

Observation ebb95258-c348-46e4-bc7a-d19ec1053c73 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.828141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.828141Z digest=sha256:3caefc81a99ca645ad01ba0ab120030b9e9019709d7366c164dc5a7072854d75

Observation f3351db5-7b83-4ae9-a3f6-4568eec2f45a · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:01.984634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:01.984634Z digest=sha256:69cf26308c597acc6f73601397e24b46d2b87400885777151f65741356b10b4a

Observation 2da8fe5c-44cd-4d40-8531-8ef120e6c04b · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Opencompass: A universal evaluation platform for foundation models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.083800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.083800Z digest=sha256:ff4d0cf167ddbe2bcc3cd09a9ffa2fdaa0e27ff2a35ae752acfbcef19cbb9788

Observation 61c2ff98-d663-461d-8d30-7431fc90918e · outbound

This paper cites Grok-1.5 vision preview: Connecting the digital and physicalworlds with our first multimodal model, 2024.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Grok-1.5 vision preview: Connecting the digital and physicalworlds with our first multimodal model, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.213459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.213459Z digest=sha256:9b4f18919e4005da3f120e6db32be983f400726ce486e971f72a6280a1911d81

Observation ed114bc2-240a-41d7-81a9-833564627524 · outbound

This paper cites Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.382480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.382480Z digest=sha256:d6e53d9188864500ce6597e41ac7b29552ae2f1bdfb5a56a7192e82d8570fe4b

Observation c982de7d-2989-4d96-a4aa-8479303dc6d1 · outbound

This paper cites Deformable convolutional networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Deformable convolutional networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.460206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.460206Z digest=sha256:2a6d923c2c81e44558994ce960fb08d322384f45ed94166c3f56911646edfbdb

Observation 9942ddee-4829-4229-80b5-987fe6c51805 · outbound

This paper cites Scaling vision transformers to 22 billion parameters.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling vision transformers to 22 billion parameters

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.556501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.556501Z digest=sha256:2e1ac770a6ad87d8e777772f27b49c7d53ecaa267bbbbea63d1201bce73d6363

Observation 159447a5-105f-49d3-ae2b-f9316f254360 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.653350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.653350Z digest=sha256:05c5ce80b07cb9932043096be953da15d14b96bace6a0a6e55bacd43be379292

Observation 04837be4-82b1-44a1-990f-3694e521845c · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Imagenet: A large- scale hierarchical image database

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.781183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.781183Z digest=sha256:6fe9c051e55b75c3fe85697cb721a4c21c2e7c2b78034bf38ac1c4430f1c876f

Observation bd0b9c54-068a-4ccd-8543-fb75e41684bd · outbound

This paper cites Scalable Vision Language Model Training via High Quality Data Curation.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scalable Vision Language Model Training via High Quality Data Curation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.977919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.977919Z digest=sha256:7cf6ef3b3ef37695300a508412074cd20a5add696f94583348060ded09440f1a

Observation b9a8a1c8-dff9-438b-a88b-5f5ccc91fee6 · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Benchmarking and Improving Detail Image Caption

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.136534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.136534Z digest=sha256:388c54877ed44424f59898fdd019c721532c03fc69b5bae3b7fd5a8e0793c145

Observation dc0f21b2-c60c-4755-8734-e486f0e811b7 · outbound

This paper cites Adalrs: Loss- guided adaptive learning rate search for efficient foundation model pretraining.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Adalrs: Loss- guided adaptive learning rate search for efficient foundation model pretraining

Reference 26

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:53:11.795734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:03.306529Z digest=sha256:0ee54ebac03c1b39fce82b199272f26a3876a0480a5ad6e518e703f93a7e75d0

Observation de786191-9cd0-4b9f-ae4a-b2b735355d81 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.510387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.510387Z digest=sha256:cbea842e691e7464aeacd07228d6a8c61cf4d0a9da82b57cc64f3d97c6f1345e

Observation de626c13-022e-4fee-8d32-926d2276b58d · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.669288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.669288Z digest=sha256:d0fab482eb303f70688d0a2dd457d45c3cdb6bd664cabb573317b31188dac3be

Observation 4d8ec47f-687d-4fea-ab49-a7e1a0a11765 · outbound

This paper cites Scalable Pre-training of Large Autoregressive Image Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scalable Pre-training of Large Autoregressive Image Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.825159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.825159Z digest=sha256:55dac2e81fda7379fd07e82fb3644c40d698c3cc803ef9c49b805b78bb3cc6bc

Observation ece4340a-d08d-4993-894a-dab7d2cd9aaa · outbound

This paper cites Scaling Language-Free Visual Representation Learning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling Language-Free Visual Representation Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.943429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.943429Z digest=sha256:6e5702ff4b86db0338c2ee89b15cbfd001c73d6a65325c131646329bd65a3d56

Observation c8b7ce0f-21d6-4ae1-8aca-36773b1fb843 · outbound

This paper cites Data Filtering Networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Data Filtering Networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.057607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.057607Z digest=sha256:6b7bbb856c25f7dd7bb17bc6991b9a78b5f9c3d2075b9fd0e483abab8b0f34e9

Observation da4f14f6-bca2-4478-8c1e-597d7811d6e1 · outbound

This paper cites Slowfast networks for video recognition.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Slowfast networks for video recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.154213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.154213Z digest=sha256:19fd8fda13fc5b0ef0523915623256c25344f3067d43488d408ff267254a4ae6

Observation 420672e1-1066-4026-9620-448ba0bb8c5f · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.264739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.264739Z digest=sha256:c7718c05dbbd7574cd6020273927e1f61e9ca5debbaa469a7b333cfa554c5fe9

Observation 591ef97f-139f-478e-a5a2-38dfef806161 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.349270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.349270Z digest=sha256:11ddee54c9b29823cb4b6d6903247e233d66c677b383a7b6e30fd162ad9de251

Observation 974bae22-f802-40f6-a970-4f83a53ba6ec · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.453395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.453395Z digest=sha256:3e6ec68ce2e46e1052a65ab873e05658f93ce60662efd1fdbb521a253ed2524f

Observation a1e2ecc9-2e50-4db3-95b5-b9ad56713b35 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.556702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.556702Z digest=sha256:d5f23c0a300c7702ba2b46222afe78233cb9dbbc46a0b24d4baf131c788ff9df

Observation 6db352bb-2fd3-4d10-a4e3-c8c098c71fa6 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.659113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.659113Z digest=sha256:5fdef093875a6a1c3c231234a50e7da092af98656bec44eb4283a24540aaadc5

Observation d9cc9776-df00-41ae-a9bf-2412b3bcc69f · outbound

This paper cites Mask r-cnn.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mask r-cnn

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.765807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.765807Z digest=sha256:2c2ef28e84dec68e772220bf777db1eb036dbeb53360bbbbb1f391fadaf76978

Observation c9578615-533c-4f5f-9c72-b3e17aa48214 · outbound

This paper cites Deep residual learning for image recognition.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Deep residual learning for image recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.889813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.889813Z digest=sha256:c705afc1caa5631d6feaba853295f0f95638348f8ef32cc842b39a617fb562f1

Observation 7749181b-17cc-4cf8-844c-29f86273e8c5 · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization,.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The many faces of robustness: A critical analysis of out-of-distribution generalization,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:04.975078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:04.975078Z digest=sha256:edab883edb9e4ea516db53c36df0d1ff6f663b374ff087839be423f968c3321a

Observation fa1a1a87-8ac0-46cc-8df7-fc82cd529023 · outbound

This paper cites Natural adversarial examples, 2021.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Natural adversarial examples, 2021

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.049397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.049397Z digest=sha256:92867f85b3bc7f23f050343bcf7387c2c58f13f1d0cccd77414ea856f3b9397e

Observation f0de58a4-8897-4c39-adb4-2a867911ca71 · outbound

This paper cites Squeeze-and-excitation networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Squeeze-and-excitation networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.144799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.144799Z digest=sha256:9e5def622e6c937406fa4df4c6426ec41b9e8d6ae2f9a60fd2fb5ae0b8280f45

Observation f611ea19-3a0c-48f2-88f7-795668443c75 · outbound

This paper cites Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.239887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.239887Z digest=sha256:c2841a06926013fd0e804e31baf3b2032cdca1a1ebb4040b836589c57c5c850d

Observation 51a87678-09b0-4b3d-83b4-a3a6103152c5 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling up visual and vision-language representation learning with noisy text supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.355096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.355096Z digest=sha256:aa5fa3aba99e121e1a5b1588888535cef66a4f87af850606f0902323183a3aa5

Observation e5eab2a7-4141-450a-ad36-36ab8876a996 · outbound

This paper cites Deep visual-semantic alignments for generating image de- scriptions.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Deep visual-semantic alignments for generating image de- scriptions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.440840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.440840Z digest=sha256:8b05731c3a5f5e13d13466a1b6ec7c05014fd14f7d93cc7eae0a371295483599

Observation 74f37a5d-2c5c-4385-b53d-f0865f2b4502 · outbound

This paper cites A diagram is worth a dozen images.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement A diagram is worth a dozen images

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.976874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:05.499872Z digest=sha256:c92f9c7243e872f3e78f48e0c84614e1f62f84396224ad9d6fa8fe0022219700

Observation 1b36cab1-6f83-4bf8-8687-e77075edc616 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Lisa: Reasoning segmentation via large language model

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.815444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:05.571684Z digest=sha256:2272863452e3c172190594da7efe2d9c7c4f58d212f8d4d60005475cfc381a44

Observation 827c4573-8d18-4ff2-bb34-487ed97466f8 · outbound

This paper cites Building and better under- standing vision-language models: insights and future directions.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Building and better under- standing vision-language models: insights and future directions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.676412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:05.672563Z digest=sha256:958867607a70a5ba76d397a70e394946799d96f810c5789a234d86b0dfcbb51b

Observation 0044f049-24e4-4e83-863d-0ba8b5f52189 · outbound

This paper cites What matters when building vision-language models? Advances in Neural Information Processing Systems, 37:87874– 87907, 2024.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement What matters when building vision-language models? Advances in Neural Information Processing Systems, 37:87874– 87907, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.515187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:05.761968Z digest=sha256:467f72b21a3effc8efa1d2ecbf690f99ff35dc7ef369f04c1dad97de85797fe4

Observation 8e3d7bf0-8038-4b4e-9817-32ed004715e8 · outbound

This paper cites The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.867758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.867758Z digest=sha256:ae5b81dad8234db2fcaf6267f286e0e77f480bff0812dfaf01ba344fa20d48ac

Observation fd89d39e-9699-41a9-9007-b31e3cb67f5c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement LLaVA-OneVision: Easy Visual Task Transfer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.966643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.966643Z digest=sha256:81baefc8651540bb8dc2cc8722376a520e8e93fd14ffc0f03b67d86fe9b90ba5

Observation 8edbc22b-7bf8-4438-9569-05985e23b1d5 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.041336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.041336Z digest=sha256:9b9c37994bce875fa37a53aa0f66bb44dde9b7dc9b911ad45b5cad7df049bb41

Observation 59f4a2d1-fba9-48e7-99af-b69dbd5588ce · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.332980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:06.099402Z digest=sha256:826857203327eb5bb0b941b52436ba8a0c09aa58969cb5fa6fdb70bed848cbc0

Observation ebfdf35a-1998-4dcc-9cef-7e50202a6687 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.204273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.204273Z digest=sha256:1c95ffe2f2c0f5a809d04a837d4678a67d06255064ea34608526190f4b8fff2c

Observation 333b4891-09aa-4b17-bf68-fae1118550be · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Evaluating Object Hallucination in Large Vision-Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.310118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.310118Z digest=sha256:3c9c67a3ac2c531c57e2810ff0b68dd8d4cc16e7427b04d55e014e3cb264fa78

Observation 0283545d-c372-412b-b1ce-0e4581157c35 · outbound

This paper cites Improved baselines with visual instruction tuning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Improved baselines with visual instruction tuning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:16.185468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:06.413431Z digest=sha256:0a7df55884159ec6cef7ee4a57848b6fa849187f39af2d2925c7c7f00a102201

Observation 4a2fe0a4-b236-4c80-8df4-5eaf3dd23463 · outbound

This paper cites Visual instruction tuning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Visual instruction tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.961535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:06.494248Z digest=sha256:a77c5da14ce7f5d99da92ec3788ec817a54dc7005338fd2777bded8c7f22132f

Observation a5a01265-787c-40d3-adc5-597f0b495cdd · outbound

This paper cites Muon is Scalable for LLM Training.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Muon is Scalable for LLM Training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.548354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.548354Z digest=sha256:bcdf640dbfd911df07e88d42e0f21f76fd95a89c7d698aabb761581d9055c17d

Observation ae89d1a8-3bec-4090-ba17-80b3218be92e · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.741340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:06.649180Z digest=sha256:45fa69eb774525452357b0f23bdcd3d740e058cfa3afb5422a1d16557085c799

Observation 345b18e4-314a-4e36-8430-5d72da46ed77 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.552804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:06.750645Z digest=sha256:73f2c1ef5269c9d09f57b6ee0d98cd3e995304d17b9f02477a512a5f9f022a2e

Observation 2df5e4c3-e566-4e69-b955-7148d92b9383 · outbound

This paper cites Decoupled Weight Decay Regularization.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Decoupled Weight Decay Regularization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.868876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.868876Z digest=sha256:f4431ac20a5bb7d460d684b2f60d2f3edaf49044b69d19f70bfdfdc37d65ac7a

Observation e37a86bd-b7da-4470-9ad4-db79d3c5e662 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:06.917722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:06.917722Z digest=sha256:d4fb0e67ffd228ba6bf46f0bdb8baf12d417645bcb201daaba4ebc7971a3d3de

Observation a29d00c8-d994-499b-a20e-78726d285db5 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.020742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.020742Z digest=sha256:3c6cb97703ac1e88ce8b72434857aa45ae197d802222c87c6eedae5894cb1c26

Observation bbbcbdee-0cf1-4a78-ac00-b63f7b284021 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.393314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:07.120075Z digest=sha256:b72380bbab825e22fbc01de15e25234e2c5afa5dbae7e6c03bc0855f3559c96c

Observation ed625676-9ef6-457f-8452-fdcc4955ba73 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.221154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.221154Z digest=sha256:f499757f2556c8f9be7206ee1090ea56d786219ef13877d747ce7369ef23c2db

Observation aea5201b-43bc-4e53-925d-97fcc6f8ebe8 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.312239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.312239Z digest=sha256:440bd495fb1f42535d52ceff005aeb4e26abe8bd0c36d7e938951294a751a71d

Observation 831dbd81-8405-4b6e-8b64-aefbf16eef2b · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.233277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:07.419963Z digest=sha256:e94f19a46a9e9ea8177e1c169c534ed84b4c6bcb35a410bcf93c81ddc3190fae

Observation 63e2f7dd-ad86-4244-a702-b01fb1db3d3a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:07.511190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:07.511190Z digest=sha256:8b2cead4d78292b7671d3e48f4119d379dfcc47ebbcde5bb730e083d3d43a184

Observation fde217ee-d772-4282-b2c7-64e7d15c94d0 · outbound

This paper cites Infographicvqa.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Infographicvqa

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:15.058601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:07.611803Z digest=sha256:c4464b712e52b161eb3a1c7288e605666c15ddd2cdf77c018bbaaea56658e991

Observation 01ba0488-e905-41f9-87f4-b2895e1b06cd · outbound

This paper cites Docvqa: A dataset for vqa on document images.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Docvqa: A dataset for vqa on document images

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.878279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:07.684217Z digest=sha256:8296e13cd272eac0015be2de1e52dde64b6f08cd503f77f9d2c5ce5a56d171e6

Observation 1afdc95a-9888-4be3-a8da-0751921568e7 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ocr-vqa: Visual question answering by reading text in images

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.694685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:07.797962Z digest=sha256:6713bd7807394dfb3e1bbc096750131425a69a02f85a492d5e8e4d0596c3ee89

Observation d7e7995f-9a97-4699-b4fa-2833ed44db52 · outbound

This paper cites Introducing chatgpt, 2022.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Introducing chatgpt, 2022

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.509068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:07.848469Z digest=sha256:1130398de9fe5b5cd8ddd9ecbb76e4ab1485c8a073e556000ed369a5c10b98f6

Observation 65d44a07-4e93-4451-bfc2-fb2aad7ba817 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Learning transferable visual models from natural language supervision

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.329003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:07.959684Z digest=sha256:d5e05e4ac5631dbc8778064154688be825bc518ddba19afb66c4f91c7e035df9

Observation b61bca55-9d38-466c-a543-7801a7c5a1d9 · outbound

This paper cites Zero: Memory opti- mizations toward training trillion parameter models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Zero: Memory opti- mizations toward training trillion parameter models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.195330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:08.067283Z digest=sha256:a5f115b81dd7732f77848085da0d7a392faafbdae73233a6aadd140e550ffe23

Observation cbcbd198-0e40-451d-bb14-ea29fa46ad0a · outbound

This paper cites Do imagenet classifiers generalize to imagenet?, 2019.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Do imagenet classifiers generalize to imagenet?, 2019

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:14.040246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:08.176323Z digest=sha256:ade7c4496f18956d6ac1302acc4c0a55bb7bfb6e000a848534f7b2d04389fb5f

Observation 67b0689a-37ba-48e2-97be-68b05328c0e7 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:08.277308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:08.277308Z digest=sha256:cc8ce6ed592b33a5e5ffd07eb86a2d5c42978f96f03260cf67074c2d1a8a0b80

Observation 9f4b5191-474d-4218-a490-47ba5b5a4763 · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:08.389462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:08.389462Z digest=sha256:ba37bc29c00d69585e588d205c0b3aa2ab31f5f4f31b9546abffe92fe3ad37e0

Observation 129c05bc-c99d-480f-9f80-c2b7fbf1fc79 · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Hollywood in homes: Crowdsourcing data collection for activity understanding

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.840234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:08.498441Z digest=sha256:b2bf97eb14807994349b0bc8abf9531d8789f2414259aa316d796954689a3d95

Observation d499bfdf-0ff9-40e4-b799-59f07d5dab77 · outbound

This paper cites Towards vqa models that can read.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Towards vqa models that can read

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.673354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:08.610364Z digest=sha256:12e3575c69769fc00b6036d742ce0d92f52c0653ade2606dbe9db193b110a664

Observation 446b6f94-7e19-4f2a-bf2f-8b0c05c62892 · outbound

This paper cites Kimi-VL Technical Report.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Kimi-VL Technical Report

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:08.679854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:08.679854Z digest=sha256:4f979a5e27bc7eaa27c84369cb046d37cda49c13abad21c496b417104610a9f4

Observation e03a2961-db65-42c8-b3da-05e24ef9ad89 · outbound

This paper cites Qwen3, April 2025.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen3, April 2025

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.554335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:08.740655Z digest=sha256:a30b6728b74d42cd5a3e715bbb4c9b062124ab06c5e0b8cc0405914cb76637f6

Observation fcb05f98-913a-4806-b5f8-fbd495febd22 · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:08.830124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:08.830124Z digest=sha256:472ead1285f2ee1fdd87c397fe8a5ec91335227c4383ade93f25e373c952b1d0

Observation 366e44a7-412d-4023-a12b-e1e04d53176d · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.392724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:08.912731Z digest=sha256:84a066a216b5816a42acbc42f0b9efcb843cb06c2929f548c9e0006a502a6e3b

Observation 01b2c095-0adc-41bd-b71b-925321588093 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.011049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.011049Z digest=sha256:2d08fe6ca5fdeb00e49ac588b7b8753656f8289d32b236a1abc47e77fa98810c

Observation 84feef7a-bfbd-46ec-adaa-a9571bf3c6ae · outbound

This paper cites VGR: Visual Grounded Reasoning.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement VGR: Visual Grounded Reasoning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.114205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.114205Z digest=sha256:1740b2f5edf3a673043dffc8fff3548d3beebf394cbb7fc17f95c95492fee95d

Observation 68de02d4-03e1-4ab7-b0da-48d7c9ffbfc4 · outbound

This paper cites World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.193598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.193598Z digest=sha256:21331a282c7ee29d5a8238bbcafb869dacad0545623e38587b8cde7e44ddf5f2

Observation 99cb87b7-ae69-4d36-a51b-54a0499edfd3 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.284570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.284570Z digest=sha256:38bc357446912459d8f8091d9fe2de43ff3d8c3f486c77f55fa9b50003d2973b

Observation 34bcd4d3-5a97-49c0-87ce-d15e2db08641 · outbound

This paper cites Pvt v2: Improved baselines with pyramid vision transformer.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Pvt v2: Improved baselines with pyramid vision transformer

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.252292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:09.384776Z digest=sha256:3c0c4bdd55dcece0b826aa62700e1820d62ec1f8680aab7271425226ad499d3f

Observation 275553a3-2d36-4427-becc-762e4e7daf42 · outbound

This paper cites Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.464295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.464295Z digest=sha256:f1029e4d8b99e07bd8fd47a0575190e401f715b9ec03058f062a92ba85bb152d

Observation df4f8766-c159-4818-9a53-645cf33e705d · outbound

This paper cites DocStruct: A Multimodal Method to Extract Hierarchy Structure in Document for General Form Understanding.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement DocStruct: A Multimodal Method to Extract Hierarchy Structure in Document for General Form Understanding

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:53:11.128260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:09.562208Z digest=sha256:4ad9da4f190c427ef06bd1679c0e7d641f9c96db227fe71ecdb85dfd6d1bc9b4

Observation 42533c59-304e-48d5-b8b2-58c198e861a9 · outbound

This paper cites Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:53:11.031292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:09.643301Z digest=sha256:00f1b778471a3abd2374150a846b018634bcb75e6a0efe8c1a22d9d0444c77f7

Observation 21a80b97-285b-427d-8e2a-9042d2fba491 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Convnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:13.058543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:09.735785Z digest=sha256:abd8df57b68328eaabc67fce336c3ce2712ab44a2cd183be9409e6a12483767e

Observation 665e9818-9536-480e-a7e3-16e73fe9a998 · outbound

This paper cites Seeing the image: Prioritizing visual correlation by contrastive alignment.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Seeing the image: Prioritizing visual correlation by contrastive alignment

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.875815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:09.815315Z digest=sha256:db90b3cf0fc1d3555691166b96af8497727e8aa2f84ff9a02b36b966c2f15714

Observation 6e00b62d-a9a6-4216-aa41-1eb5feb23ef1 · outbound

This paper cites Aggregated residual transformations for deep neural networks.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Aggregated residual transformations for deep neural networks

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.734487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:09.911243Z digest=sha256:b57bf8c1e8a88918ef481d0b845394019450d1adb5ae43ed7a9801ef7164fd86

Observation 3da9d2ac-8853-4031-9f32-69a7b568ddd4 · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.990394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.990394Z digest=sha256:69a54ed3ac9d2ff3ef963f40145d3b9895dfa484d503ea68074ba5249c81725c

Observation b775d07d-d57c-4e11-ac09-ae766d755e34 · outbound

This paper cites Qwen2.5 Technical Report.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen2.5 Technical Report

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:10.097531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:10.097531Z digest=sha256:8052b653a97618ef658ad07a66d16c7908c5a27d83ef5903ade770891754f88d

Observation 7de236ae-333d-48c8-ab3d-74acdbdccaac · outbound

This paper cites Pediatricsgpt: Large language models as chinese medical assistants for pediatric applications.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Pediatricsgpt: Large language models as chinese medical assistants for pediatric applications

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.585192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:10.152994Z digest=sha256:9a96d7a1f30d7064eef6c9212557370633ec222c5ee6fd2b74cc293f4e7b7cc6

Observation 431a2d12-4355-440a-81a5-9cd058174da8 · outbound

This paper cites Improving factuality in large language models via decoding-time hallucinatory and truthful comparators.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Improving factuality in large language models via decoding-time hallucinatory and truthful comparators

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.432758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:10.220443Z digest=sha256:1103c079d1879b77f4b73c02a99226a257f452f6f19b602dd76f013358a7fe76

Observation 841289a6-f86c-493d-9105-fad874f1c0bb · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:10.285746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:10.285746Z digest=sha256:1955b310b99bf96c9f10d30cb92d56aaef65af935df88fc4c2b9c08a126a0806

Observation 559193a3-2037-4663-b025-4d310ebbb974 · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:53:12.313580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:53:10.367278Z digest=sha256:84d3e23762906ac16081b60c3c9ec7c2f102e78c869ab0fc4ad31f16e8b9604c

Pith citing papers

No inbound Pith citation observations are available.