Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:53:10.367278Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2507.01643.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:53:10.367278Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 105 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 87a779a4-5d37-4ccc-9100-ca4a1055b45e · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3f8b829-d3b8-4c9e-b49c-7a2c08ed4b7a · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53466165-6d22-4a61-b980-3bf3b80cc3d9 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50929551-2aa8-4e2f-ad9f-43148dab93d1 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4edcd2fc-7665-47a6-9aa1-72c48046aeb5 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement OCR-IDL: OCR Annotations for Industry Document Library Dataset
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7f9a6f9-9427-4f31-81fb-2221bf4d07f5 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Coyo-700m: Image-text pair dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4efb48bd-318b-4f97-9e72-279d1f7ed98a · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Reversible Column Networks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df37130e-1b88-4449-b312-380fe9e3d592 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement InternLM2 Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76b273cb-d5bd-4289-b241-8eed1c6b3670 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6603d20d-9e99-4651-bdeb-78240a99980d · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement End-to-end object detection with transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 846e9db1-cacd-4b1a-ac71-bd61cb2742e9 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Sharegpt4v: Improving large multi-modal models with better captions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36a94428-a52a-4253-9157-6208eaf65a54 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 190292f5-0b9c-450d-b19e-a5b00cc01660 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Vision Transformer Adapter for Dense Predictions
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e4e7dab-6dad-4dce-a34a-f5ccd4100a79 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb95258-c348-46e4-bc7a-d19ec1053c73 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3351db5-7b83-4ae9-a3f6-4568eec2f45a · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da8fe5c-44cd-4d40-8531-8ef120e6c04b · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Opencompass: A universal evaluation platform for foundation models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61c2ff98-d663-461d-8d30-7431fc90918e · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Grok-1.5 vision preview: Connecting the digital and physicalworlds with our first multimodal model, 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed114bc2-240a-41d7-81a9-833564627524 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c982de7d-2989-4d96-a4aa-8479303dc6d1 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Deformable convolutional networks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9942ddee-4829-4229-80b5-987fe6c51805 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling vision transformers to 22 billion parameters
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 159447a5-105f-49d3-ae2b-f9316f254360 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04837be4-82b1-44a1-990f-3694e521845c · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Imagenet: A large- scale hierarchical image database
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd0b9c54-068a-4ccd-8543-fb75e41684bd · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scalable Vision Language Model Training via High Quality Data Curation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9a8a1c8-dff9-438b-a88b-5f5ccc91fee6 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Benchmarking and Improving Detail Image Caption
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0f21b2-c60c-4755-8734-e486f0e811b7 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Adalrs: Loss- guided adaptive learning rate search for efficient foundation model pretraining
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation de786191-9cd0-4b9f-ae4a-b2b735355d81 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de626c13-022e-4fee-8d32-926d2276b58d · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d8ec47f-687d-4fea-ab49-a7e1a0a11765 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scalable Pre-training of Large Autoregressive Image Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ece4340a-d08d-4993-894a-dab7d2cd9aaa · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling Language-Free Visual Representation Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b7ce0f-21d6-4ae1-8aca-36773b1fb843 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Data Filtering Networks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da4f14f6-bca2-4478-8c1e-597d7811d6e1 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Slowfast networks for video recognition
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 420672e1-1066-4026-9620-448ba0bb8c5f · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 591ef97f-139f-478e-a5a2-38dfef806161 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 974bae22-f802-40f6-a970-4f83a53ba6ec · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e2ecc9-2e50-4db3-95b5-b9ad56713b35 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db352bb-2fd3-4d10-a4e3-c8c098c71fa6 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9cc9776-df00-41ae-a9bf-2412b3bcc69f · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mask r-cnn
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9578615-533c-4f5f-9c72-b3e17aa48214 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Deep residual learning for image recognition
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7749181b-17cc-4cf8-844c-29f86273e8c5 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The many faces of robustness: A critical analysis of out-of-distribution generalization,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa1a1a87-8ac0-46cc-8df7-fc82cd529023 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Natural adversarial examples, 2021
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0de58a4-8897-4c39-adb4-2a867911ca71 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Squeeze-and-excitation networks
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f611ea19-3a0c-48f2-88f7-795668443c75 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51a87678-09b0-4b3d-83b4-a3a6103152c5 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scaling up visual and vision-language representation learning with noisy text supervision
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5eab2a7-4141-450a-ad36-36ab8876a996 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Deep visual-semantic alignments for generating image de- scriptions
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74f37a5d-2c5c-4385-b53d-f0865f2b4502 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement A diagram is worth a dozen images
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1b36cab1-6f83-4bf8-8687-e77075edc616 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Lisa: Reasoning segmentation via large language model
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 827c4573-8d18-4ff2-bb34-487ed97466f8 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Building and better under- standing vision-language models: insights and future directions
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0044f049-24e4-4e83-863d-0ba8b5f52189 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement What matters when building vision-language models? Advances in Neural Information Processing Systems, 37:87874– 87907, 2024
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e3d7bf0-8038-4b4e-9817-32ed004715e8 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd89d39e-9699-41a9-9007-b31e3cb67f5c · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement LLaVA-OneVision: Easy Visual Task Transfer
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8edbc22b-7bf8-4438-9569-05985e23b1d5 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59f4a2d1-fba9-48e7-99af-b69dbd5588ce · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ebfdf35a-1998-4dcc-9cef-7e50202a6687 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 333b4891-09aa-4b17-bf68-fae1118550be · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Evaluating Object Hallucination in Large Vision-Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0283545d-c372-412b-b1ce-0e4581157c35 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Improved baselines with visual instruction tuning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4a2fe0a4-b236-4c80-8df4-5eaf3dd23463 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Visual instruction tuning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5a01265-787c-40d3-adc5-597f0b495cdd · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Muon is Scalable for LLM Training
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae89d1a8-3bec-4090-ba17-80b3218be92e · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 345b18e4-314a-4e36-8430-5d72da46ed77 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2df5e4c3-e566-4e69-b955-7148d92b9383 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Decoupled Weight Decay Regularization
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e37a86bd-b7da-4470-9ad4-db79d3c5e662 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a29d00c8-d994-499b-a20e-78726d285db5 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbbcbdee-0cf1-4a78-ac00-b63f7b284021 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ed625676-9ef6-457f-8452-fdcc4955ba73 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aea5201b-43bc-4e53-925d-97fcc6f8ebe8 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 831dbd81-8405-4b6e-8b64-aefbf16eef2b · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63e2f7dd-ad86-4244-a702-b01fb1db3d3a · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde217ee-d772-4282-b2c7-64e7d15c94d0 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Infographicvqa
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01ba0488-e905-41f9-87f4-b2895e1b06cd · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Docvqa: A dataset for vqa on document images
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1afdc95a-9888-4be3-a8da-0751921568e7 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Ocr-vqa: Visual question answering by reading text in images
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d7e7995f-9a97-4699-b4fa-2833ed44db52 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Introducing chatgpt, 2022
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 65d44a07-4e93-4451-bfc2-fb2aad7ba817 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Learning transferable visual models from natural language supervision
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b61bca55-9d38-466c-a543-7801a7c5a1d9 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Zero: Memory opti- mizations toward training trillion parameter models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cbcbd198-0e40-451d-bb14-ea29fa46ad0a · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Do imagenet classifiers generalize to imagenet?, 2019
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 67b0689a-37ba-48e2-97be-68b05328c0e7 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f4b5191-474d-4218-a490-47ba5b5a4763 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The Curse of Recursion: Training on Generated Data Makes Models Forget
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 129c05bc-c99d-480f-9f80-c2b7fbf1fc79 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Hollywood in homes: Crowdsourcing data collection for activity understanding
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d499bfdf-0ff9-40e4-b799-59f07d5dab77 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Towards vqa models that can read
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 446b6f94-7e19-4f2a-bf2f-8b0c05c62892 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Kimi-VL Technical Report
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e03a2961-db65-42c8-b3da-05e24ef9ad89 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen3, April 2025
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fcb05f98-913a-4806-b5f8-fbd495febd22 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 366e44a7-412d-4023-a12b-e1e04d53176d · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01b2c095-0adc-41bd-b71b-925321588093 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84feef7a-bfbd-46ec-adaa-a9571bf3c6ae · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement VGR: Visual Grounded Reasoning
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68de02d4-03e1-4ab7-b0da-48d7c9ffbfc4 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement World to Code: Multi-modal Data Generation via Self-Instructed Compositional Captioning and Filtering
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99cb87b7-ae69-4d36-a51b-54a0499edfd3 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34bcd4d3-5a97-49c0-87ce-d15e2db08641 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Pvt v2: Improved baselines with pyramid vision transformer
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 275553a3-2d36-4427-becc-762e4e7daf42 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df4f8766-c159-4818-9a53-645cf33e705d · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement DocStruct: A Multimodal Method to Extract Hierarchy Structure in Document for General Form Understanding
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 42533c59-304e-48d5-b8b2-58c198e861a9 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 21a80b97-285b-427d-8e2a-9042d2fba491 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Convnext v2: Co-designing and scaling convnets with masked autoencoders
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 665e9818-9536-480e-a7e3-16e73fe9a998 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Seeing the image: Prioritizing visual correlation by contrastive alignment
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e00b62d-a9a6-4216-aa41-1eb5feb23ef1 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Aggregated residual transformations for deep neural networks
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3da9d2ac-8853-4031-9f32-69a7b568ddd4 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b775d07d-d57c-4e11-ac09-ae766d755e34 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Qwen2.5 Technical Report
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7de236ae-333d-48c8-ab3d-74acdbdccaac · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Pediatricsgpt: Large language models as chinese medical assistants for pediatric applications
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 431a2d12-4355-440a-81a5-9cd058174da8 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Improving factuality in large language models via decoding-time hallucinatory and truthful comparators
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 841289a6-f86c-493d-9105-fad874f1c0bb · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 559193a3-2037-4663-b025-4d310ebbb974 · outbound
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.