Pith. sign in

Paper Citation Record · LEDGER

SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2402.05935.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.05935 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:10:28.973202Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.058979Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3bd4fa02-295b-4fc1-a701-98eb8693ee29 · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:07:42.641601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:52ef06af28301f9be4d89b47656f72a7387692170e105c02ed5fc260226ba3da

Observation 8f99e238-87b7-449b-a9b6-b90bdbdbe51d · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.755206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:4047cfba759f66a57410419caa41b532b122bde16fb5965e594de132024750e9

Observation af8ae0c2-4493-4ead-8c91-db74540d08e0 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.109213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:e3b0c2893bee9f3e5b89331e3e38ce49085b50f0582bdc114fb701c31818c12a

Observation c0f9db05-ce6e-42d4-a7e3-872c866349bf · inbound

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? cites this paper.

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:29:30.108729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T01:29:30.032408Z digest=sha256:7da0ea3c03e273645eafda95ec3a116b36fb2d91267dd7cf3877a6c927b2fcb3

Observation 29366462-c276-43d9-aec2-81dab9c59aa7 · inbound

Are We on the Right Way for Evaluating Large Vision-Language Models? cites this paper.

Are We on the Right Way for Evaluating Large Vision-Language Models? SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.501498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T19:41:44.263663Z digest=sha256:a714ab2d9606ce76e9089264a0905e04731fe051935fe752bce48a900f836dfe

Observation 55871160-18e5-499d-9433-355b10e849f1 · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:05:03.779425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:de90db42c490745774e962e348452fc29bbaad60dc6537b8b7e3fd57359fe3db

Observation 7be3e54b-3137-42fd-b0d5-754a9281ea62 · inbound

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models cites this paper.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.007023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:76da272101c904357a1ec3b20be09d12a84a80bd7e8296691901179ad4bb0543

Observation 289ae9cd-18e5-408e-a7a0-98ad1fee55ee · inbound

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining cites this paper.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.973202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.973202Z digest=sha256:3a553fdc904d9cd96dee59d98b4e4071bf60cef6ae679062bc1e87036ac7c2d1

Observation 5f912bcb-b483-4bd9-8eef-d73b36f1e3da · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.501091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.501091Z digest=sha256:7977d4b623a5a6cec2c352330d5c9e1ae05f4cf02e03d630e457374996462eb3

Observation 63b36243-8350-4724-85a8-431ab9f2679a · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 175

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.313149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:7ce933da67db5b25c987d46f31c0068dc86ab0be610b36e1968252645fdec03a

Observation 2cbadcd9-caee-4e34-9511-f6bea7f32226 · inbound

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding cites this paper.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.622849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.622849Z digest=sha256:cc14580118b2335b7d6881bacc73b0a5c55316db8757b95a24edc4b770e40722

Observation 7b675fb5-e226-4e93-853a-ee05c37b0b8b · inbound

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark cites this paper.

MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T11:26:29.178954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:26:29.178954Z digest=sha256:7d90f2522441948b653e7057cf66ab963b8d6aa8fe9bd7ef3e50bdb15f44fb60

Observation 617b104a-6ce5-43d2-8271-dfbca4478006 · inbound

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems cites this paper.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:57:13.216249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T22:55:34.238427Z digest=sha256:9b2b09cc547fcada10ad46f52b9f687b649077f22cdbbc5bcc7fe31e6f414bb9

Observation a4ca0389-5ce2-4125-898c-fd9dac8f4f8c · inbound

Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts cites this paper.

Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:25:06.518053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:25:06.518053Z digest=sha256:1587ed621058a732c73ce6a0d6ced5173469b09e6bd6f6ab9f914a4c6139607a

Observation 9626591a-cf6b-4f0e-85b6-962987fec048 · inbound

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos cites this paper.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.917393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.917393Z digest=sha256:bd634e08d1ab0b2c26e2c8761124bf04a1b6b5fd8a77032db0c39de499580f70

Observation dc258f49-a426-4b23-bc29-558dde71b26f · inbound

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation cites this paper.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:39.543588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:39.543588Z digest=sha256:58b5134e41bc9bb1243c058ab629d3821c3e9ea9f499c91aa3b9ccbf68d48c0b

Observation ce0f5d7b-3421-45af-b0a9-5c6c2810d65c · inbound

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation cites this paper.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.746024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.746024Z digest=sha256:70fbbb88530ee46cd22e6b521cc5604bd5c888d8b75d7bba8d65e54e99996ae9

Observation 019be010-9436-40f4-8fa9-182185e4c18a · inbound

KnowDR-REC: A Benchmark for Referring Expression Comprehension with Real-World Knowledge cites this paper.

KnowDR-REC: A Benchmark for Referring Expression Comprehension with Real-World Knowledge SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T21:11:51.363684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:11:51.363684Z digest=sha256:6dcf4308ccb27b3ebfba45edbeff11d88cde5c6b5a66de301923b83764de9a2e

Observation 11692ed6-db86-4248-a5f4-961a1285f591 · inbound

NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation cites this paper.

NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:40:52.962013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T04:39:58.296388Z digest=sha256:e55056585332477c68daf07da4018d7c62cde4c22b0aa60c20ee85b6e18eaf9a

Observation 9038ad2b-7b01-4e72-b0ee-82aab3cd54fa · inbound

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs cites this paper.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:05.200466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:05.200466Z digest=sha256:1766ed487bb03149842a5bfaaf477373998d2d6d27f4a8d35a6450a3cdbbf4b0

Observation 7ed6f792-9916-4963-8016-0e0cb6490cef · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:16.978097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:16.978097Z digest=sha256:419aba44628a84d28be95fecbb44b6869cc59d63a3620a80c2f7fc38a1b632c9

Observation c2171398-8dff-454d-8f16-48ea58911e64 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:15:50.606098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:891e1551c6a99ea3f9d2a6e0f448654a8c13f4d2135cfa48e672a698bf8cbe6f

Observation da617aea-d05e-4d36-83ed-82ac78656311 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:f7d0de46eea5e9f1aee2c9bf8eda0f455e9c3734e51f51f973230f8e800bbbf2

Observation 507128b0-d48a-44cc-bd47-42f6a73732b9 · inbound

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving cites this paper.

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:31:01.255346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:36:38.627415Z digest=sha256:27618379648189863d897d620caaa3ab016879641ce723c6a99dc351323ff734

Observation 65867ad8-9a18-4f36-bbb4-ad9155a03ef7 · inbound

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models cites this paper.

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:25:18.565430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T11:23:46.371799Z digest=sha256:e133b055372f1e01368bae460156dceb889125ae47ccc0517e98299d49dc2d13

Observation cccfad59-f34e-4e61-bdfd-7697f1a14d00 · inbound

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models cites this paper.

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.087658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T22:43:33.929871Z digest=sha256:4493cc77c70c647fec0b6ed0e2e4b6a99d57c9d0e115e13fbc30e30580cad2e0

Observation 076ef308-5182-43f4-9b8e-d6bfde10de36 · inbound

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams cites this paper.

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:37.911591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T14:24:19.702588Z digest=sha256:f507e114e304dd1c66b98af39f119eae5578123b7e8b126ff5e02b06daf3af3d

Observation 964e3653-86d9-4bd9-a086-fc91ddd3fe2d · inbound

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams cites this paper.

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:38.306016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:38.306016Z digest=sha256:276f6a37930553eef3983cda5ee559adad15d0b0662941228c3539be3c7eed9a

Observation 3ca1043b-cd7d-4edf-994b-34da34f2cbe3 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 289

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.649753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:8e2f2373a31b7e0eea7ecbdb924d7f91e6cf199164e72491ea284341632d0c75

Observation 6c1e1df4-7704-4278-b672-c2868a1fb8b6 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.060578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:d99183e16198e603867dc816e1612b8d1dadf9e27b450fc9fffbe1684483ecb4