Pith. sign in

Paper Citation Record · LEDGER

Mordal: Automated Pretrained Model Selection for Vision Language Models

As of 11 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2502.00241.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00241 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:46:19.985312Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bec4c05d-551a-413f-b924-54def418d6b5 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Mordal: Automated Pretrained Model Selection for Vision Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.866429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.866429Z digest=sha256:a4d3d0024bbccb3216491f3a755e58efd716efeb8bb189f62011f6d14329789b

Observation 7b6c6e59-e266-4d53-90de-df79d193246a · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Mordal: Automated Pretrained Model Selection for Vision Language Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.888073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.888073Z digest=sha256:a81c002213e9a116871f3e7795d126dabbd5eaa2412e58a83dc5e3628e2f2d0e

Observation 1d1fd062-0332-49d1-9a53-85ea2494422e · outbound

This paper cites The Llama 3 Herd of Models.

Mordal: Automated Pretrained Model Selection for Vision Language Models The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.892394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.892394Z digest=sha256:70159f15d7008af79814042634b28c3f5674c8fde4ccbf6fefb6e15e3ecf8343

Observation 90d6fd78-725b-4db2-89b1-4f9d1d0c4589 · outbound

This paper cites Scaling Laws for Neural Language Models.

Mordal: Automated Pretrained Model Selection for Vision Language Models Scaling Laws for Neural Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.911043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.911043Z digest=sha256:1430e81112578fa1a5a104b7338622fb32cb3ccfd25dac5ef0b94349929a4bc5

Observation 35fecdfa-ef75-40bd-b19e-979db3f42ba3 · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

Mordal: Automated Pretrained Model Selection for Vision Language Models Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.914910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.914910Z digest=sha256:5d8977124972fb5710a560384385a95b249419eb990d38ccf6d9b01e3569a24f

Observation 9b077f7b-fb24-485c-a199-e9f957af083f · outbound

This paper cites Selecting Large Language Model to Fine-tune via Rectified Scaling Law.

Mordal: Automated Pretrained Model Selection for Vision Language Models Selecting Large Language Model to Fine-tune via Rectified Scaling Law

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.922978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.922978Z digest=sha256:1cf2730a5daff96996923e93a6dbdf961611ce11731bb80b3a41c35c79171ac8

Observation 4671bf80-a271-4821-81f0-af4a69bbef7f · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Mordal: Automated Pretrained Model Selection for Vision Language Models Improved Baselines with Visual Instruction Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.926793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.926793Z digest=sha256:4daf03efc1534d92500e21be1c3414213b6ffdbb3f7a0bdaf294b955c4645796

Observation 77851c57-13ca-415b-83dd-d0019802b5d0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Mordal: Automated Pretrained Model Selection for Vision Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.930509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.930509Z digest=sha256:e7f84501e7dd809ba2e09bc6c59a349aadc5078188b5bd4d669c3e9772d088d1

Observation eb5187f8-49e3-4e22-bc0c-43f3a93ba8af · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Mordal: Automated Pretrained Model Selection for Vision Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.938708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.938708Z digest=sha256:bdade985468af683102c670f44297359a697735cb692f87d845fcb7d9e9cd95a

Observation e4b300bb-5e98-47f1-a0a1-b36ddc1f8eaf · outbound

This paper cites Observational Scaling Laws and the Predictability of Language Model Performance.

Mordal: Automated Pretrained Model Selection for Vision Language Models Observational Scaling Laws and the Predictability of Language Model Performance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.942383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.942383Z digest=sha256:aa7d413497e9d9362dc57933ca9a3f0eeebd5e94134f54daedd06c4c83ed0a60

Observation bf4dcb7f-4fef-44e1-a6c9-553d63dc8f8a · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Mordal: Automated Pretrained Model Selection for Vision Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.946327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.946327Z digest=sha256:a8c9854cdee3cf88f4538b2ad2fb713443250222dcb437cde237f2697904b2eb

Observation bdd6982c-a46c-401f-bc37-e4730079990d · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Mordal: Automated Pretrained Model Selection for Vision Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.950152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.950152Z digest=sha256:cfe0ce1c719c08632f29d89c9e637ddad9f6a790c8248d3b3f9dea7e98089fce

Observation 3af2ca88-3707-44b7-92dc-fe475830cd35 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Mordal: Automated Pretrained Model Selection for Vision Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.954305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.954305Z digest=sha256:aff181805fc5d39d546dcf96902499a6bf09875ce972a38384336bed9c51a051

Observation 8ff81fba-60df-44f2-b28e-a69734d8a0e9 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mordal: Automated Pretrained Model Selection for Vision Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.958360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.958360Z digest=sha256:5f6ac3487039155bda207d8410f00b841a7be27265d08e85c03ec6af63457095

Observation eb120aa1-bdd7-46e1-8e7b-ae0afff127fa · outbound

This paper cites Exploring and Predicting Transferability across NLP Tasks.

Mordal: Automated Pretrained Model Selection for Vision Language Models Exploring and Predicting Transferability across NLP Tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.962287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.962287Z digest=sha256:352104cc70d59b8efe3c940e3f163c94f90bd3743617a87fcdb9cfd4f7339d88

Observation af5bed81-1076-4e7a-a157-0d6db8b74234 · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

Mordal: Automated Pretrained Model Selection for Vision Language Models LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.966277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.966277Z digest=sha256:f90f9fb868295cf5a07ccf5e52ed72845f0a8b66a00481194f4ff61ea58cba63

Observation e791b7bd-c42c-4b90-a7b5-3d88b0c99dec · outbound

This paper cites Qwen2 Technical Report.

Mordal: Automated Pretrained Model Selection for Vision Language Models Qwen2 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.970562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.970562Z digest=sha256:06f3b6d2372f68b4afa7277410586d059b5ed8eed20527444021eac94c8fdb69

Observation 22a8c01b-60d0-4d79-9f99-e02db2c23309 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Mordal: Automated Pretrained Model Selection for Vision Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.974377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.974377Z digest=sha256:ad0174eab78ef57d323cb23c98b11f086a615469eebc37f79accb9ec9f7219ac

Observation 142f0774-44e9-44b9-819d-b2bfd02d3b8c · outbound

This paper cites In LLM clustering, we use the last hidden state from LLM as the sentence representation for CKA computation since it produces the best clustering performance.

Mordal: Automated Pretrained Model Selection for Vision Language Models In LLM clustering, we use the last hidden state from LLM as the sentence representation for CKA computation since it produces the best clustering performance

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:46:20.307624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T19:46:19.978129Z digest=sha256:430c4e4ff816899206981b3cc73636b261fe9864b80af37064ceb50ca55df84a

Observation 28ff9466-7f85-479b-8cdd-27d94d7cb194 · outbound

This paper cites For ConvNeXt, we interpolate the output embeddings to 16x16 patches following Cambrian-1 (Tong et al., 2024).

Mordal: Automated Pretrained Model Selection for Vision Language Models For ConvNeXt, we interpolate the output embeddings to 16x16 patches following Cambrian-1 (Tong et al., 2024)

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:46:20.295597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T19:46:19.981642Z digest=sha256:8d43382aac1f65c11698f48cbe539747c6aebbb52a04f2dbc005d66e90877eca

Observation c4af7bd7-4481-45a7-b670-e16f107c9c8a · outbound

This paper cites Recently, some proprietary models have employed end-to-end training without using any pretrained models (Bai et al., 2023), but it is not common due to the excessive training cost.

Mordal: Automated Pretrained Model Selection for Vision Language Models Recently, some proprietary models have employed end-to-end training without using any pretrained models (Bai et al., 2023), but it is not common due to the excessive training cost

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:46:20.283292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T19:46:19.985312Z digest=sha256:059b0805363032e238a39a7170c3b9e2e8180953a9bfe1edda25828542eae16e

Observation 120bcc3f-5689-4384-bdf4-85e971de888b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Mordal: Automated Pretrained Model Selection for Vision Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.871514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.871514Z digest=sha256:dc6f0808312d6704b61469608ecd5248e929894b94755a6e041e177e287cc794

Observation 668d5e7c-d887-45ec-9f43-82533af7dd21 · outbound

This paper cites Mistral 7B.

Mordal: Automated Pretrained Model Selection for Vision Language Models Mistral 7B

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.907772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.907772Z digest=sha256:249252bd455bc26a73e26aac5012a049f11cc9753a72b62f64cd924c11b8835e

Observation 151d75c1-6967-4e2e-9896-c2c3ce5e5c38 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Mordal: Automated Pretrained Model Selection for Vision Language Models Training Compute-Optimal Large Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.899686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.899686Z digest=sha256:e5b51404c806351dee2ab6224af28dfafe0da9f679d8351d8100b91f1628b8d4

Observation 2159f985-2c93-4955-b790-4ab8a1a02c52 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

Mordal: Automated Pretrained Model Selection for Vision Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:46:20.338125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T19:46:19.880318Z digest=sha256:bae9a61b2e812c905bbbdebac0b8a5defa6a5fb529f19520af5af3eff3ffcf75

Observation 900e0c72-6566-4462-ac90-c465f7acfa5a · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Mordal: Automated Pretrained Model Selection for Vision Language Models MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.884215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.884215Z digest=sha256:e365161c2779fe1a84eec6bf940b8238dbf8cda3cb529a6a40405f0df98d8454

Observation e41a09d0-393d-49cd-9928-cbc6189312e0 · outbound

This paper cites Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and Depth.

Mordal: Automated Pretrained Model Selection for Vision Language Models Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and Depth

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.934437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.934437Z digest=sha256:6887d752b21617b7881bbb42d9be1734369364b4893329ca3ceef056bdbb0570

Observation f96f9348-e38f-4fa3-a8ad-cc9a3a380e92 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Mordal: Automated Pretrained Model Selection for Vision Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.903358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.903358Z digest=sha256:9bd6afca8cbf8010aa0bb5f1e6337839b8c98f4586a517ec955f927e876bee71

Observation a080b3f2-6cf6-4886-8c19-e4d76c9cfa33 · outbound

This paper cites PLATO: Pre-trained Dialogue Generation Model with Discrete Latent Variable.

Mordal: Automated Pretrained Model Selection for Vision Language Models PLATO: Pre-trained Dialogue Generation Model with Discrete Latent Variable

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.876108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.876108Z digest=sha256:93e9dfbd36886bfd8a4fc03006afc87c53f4d840dd3d4e492903de03243ecac0

Observation 0ef3e349-2caf-46ea-955a-15e8c592f249 · outbound

This paper cites A diagram is worth a dozen images.

Mordal: Automated Pretrained Model Selection for Vision Language Models A diagram is worth a dozen images

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:19.918707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:19.918707Z digest=sha256:9e09beb3f1b1b3a9200f8fb282edc39d09ed0459b43316dc732aaa4edc122f66

Observation aff61c96-329e-49e3-806c-52e4cebaa9cd · outbound

This paper cites Accessed: 2025-01-30.

Mordal: Automated Pretrained Model Selection for Vision Language Models Accessed: 2025-01-30

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:46:20.326767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-09T19:46:19.896211Z digest=sha256:dec9e00bfe2e11c4aa12b9eb89cff906390c7e65cfabd78410ca85756ffef8f5

Pith citing papers

No inbound Pith citation observations are available.