Pith. sign in

Paper Citation Record · LEDGER

Scalable Vision Language Model Training via High Quality Data Curation

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2501.05952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05952 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:12:30.679583Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.276573Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 221fb2f8-0767-4ecc-b0ff-add9d8e1fe3e · inbound

Scaling Pre-training to One Hundred Billion Data for Vision Language Models cites this paper.

Scaling Pre-training to One Hundred Billion Data for Vision Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T12:12:30.679583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:12:30.679583Z digest=sha256:be6791a98ba4ffa480d5242911c1f4ca2b05671dec473a9096a0eefc156d96f7

Observation 99891245-cc99-4ab7-9a17-510f3a59453f · inbound

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs cites this paper.

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs Scalable Vision Language Model Training via High Quality Data Curation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:25:16.423936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T01:23:01.892612Z digest=sha256:5ee05c67adc2fa66dfb129e1a9854b6ead0b09fe4c09e850b3866afb5ad92483

Observation b2028d01-687a-41f2-a5b2-7f76926d2d2b · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.726967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:fa46779473bee5d52915ccdeb0f5f38a2ce5140dcb220153fb786e1d02b5fbbb

Observation 4b18d18c-0f14-454d-8fe6-7f82253a753c · inbound

Affordance Benchmark for MLLMs cites this paper.

Affordance Benchmark for MLLMs Scalable Vision Language Model Training via High Quality Data Curation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:52.599631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:52.599631Z digest=sha256:390f4fac1bdc195a7521910a30c773e4f39408294bf68610fd43c6903ea4b098

Observation 53938199-ee10-439f-95ff-bb1581701a47 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning Scalable Vision Language Model Training via High Quality Data Curation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:12:14.410055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:24baabca23cbf96087dafe1469d704d5c7bd27fd0de1e54b06222de2e7db20b0

Observation e768521e-f731-4390-ba9b-14a09ac76714 · inbound

MMSearch-R1: Incentivizing LMMs to Search cites this paper.

MMSearch-R1: Incentivizing LMMs to Search Scalable Vision Language Model Training via High Quality Data Curation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:27:04.398232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T15:27:04.228144Z digest=sha256:56af8a8a3d2d4295b6c12397c1237b47f957711e8d6249422b373f6119231844

Observation cda295d1-e7e2-4768-962b-fe867de61575 · inbound

Ovis-U1 Technical Report cites this paper.

Ovis-U1 Technical Report Scalable Vision Language Model Training via High Quality Data Curation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:19.113047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:19.113047Z digest=sha256:d72df0849f19a229475d4c41365be1e30b2175666ff47ce2db941c32b874bde9

Observation bd0b9c54-068a-4ccd-8543-fb75e41684bd · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scalable Vision Language Model Training via High Quality Data Curation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:02.977919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:02.977919Z digest=sha256:4db1fcbfee174db3292e01e08a291d0210c4e62b43b2a57a0a0fcc25e8b56c18

Observation e726fda4-961a-49f9-9f94-11b812202f0d · inbound

Describe Anything Model for Visual Question Answering on Text-rich Images cites this paper.

Describe Anything Model for Visual Question Answering on Text-rich Images Scalable Vision Language Model Training via High Quality Data Curation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:23.196657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:23.196657Z digest=sha256:126a1080f2ae5596eb13b09dd78594698bd228e522288895032c942b2c52f7fa

Observation b047048d-af23-46fa-a388-0d7e3eb87d20 · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Scalable Vision Language Model Training via High Quality Data Curation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:54:20.305212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T19:51:04.983299Z digest=sha256:2671975a51ed09271af929a6780742b3d8c7e086fe10409f8c4926e76261adef

Observation ffa74552-aa60-46b4-9ef2-edd80efcf464 · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Scalable Vision Language Model Training via High Quality Data Curation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T21:43:59.233138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:43:59.233138Z digest=sha256:fb1ecdefaead95eab56824edd00117a4198a2b9c0a1bbdf94cc2234b46f8f8dc

Observation d7a1f2b6-c17e-4c71-8291-0ba6d6dbdf98 · inbound

S-GRPO: Unified Post-Training for Large Vision-Language Models cites this paper.

S-GRPO: Unified Post-Training for Large Vision-Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:22:37.500851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T08:18:59.432486Z digest=sha256:84145f37f83bbdcdf7cf81066d32556bf67951a3e269dd538bc90dd87a785839

Observation b7b0dfae-f74c-4dfc-90af-709908f83d66 · inbound

S-GRPO: Unified Post-Training for Large Vision-Language Models cites this paper.

S-GRPO: Unified Post-Training for Large Vision-Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:54.128765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:54.128765Z digest=sha256:319d72a12b72371611d1c5cb53afe421450fd2cda4432fc343098571f897f79c

Observation 41f30715-eb9e-4cd9-a640-6df4c649d18f · inbound

Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models cites this paper.

Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:11:18.015960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-07T17:58:17.879345Z digest=sha256:fcf726c14caf7b3e2d530dd1377380c245fa3babd4d60a6f80e0a84acc8e3717

Observation 44e3a5ec-46ba-42eb-af05-5908e7a706a5 · inbound

Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding cites this paper.

Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding Scalable Vision Language Model Training via High Quality Data Curation

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:27.331648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T05:01:38.118237Z digest=sha256:39aa8ccb09c3a8e3e0f4600189841fcde30460c704600031a158d371f688de3f

Observation daf12f4d-9e6c-4abf-8a53-bfab3084d48e · inbound

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval cites this paper.

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval Scalable Vision Language Model Training via High Quality Data Curation

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:49:36.502408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T15:34:55.062016Z digest=sha256:28892cd4874440c90a8092e58eaebab670ae414696eea919b0059bd27378aefa

Observation 6cf7d13e-e151-4f27-9c49-c120b048f620 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Scalable Vision Language Model Training via High Quality Data Curation

Reference 231

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.278258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:e7e9c730bbb4bac8d292b4c4a64c7abc2fa6a63715c25db9812df2972dbcb349