Pith. sign in

Paper Citation Record · LEDGER

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2412.05237.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05237 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:08:10.041348Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:07.154168Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c46585a1-cbc9-40e3-9432-9fe4c0339963 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 294

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:10.041348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:10.041348Z digest=sha256:407f1752d650ae0a24e960ba9711af6f89f7fe480db2792c788fe073807c2b83

Observation f02c1b66-f08c-445d-a481-2881b210d20a · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.147822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.147822Z digest=sha256:c3190340656637c96d020a44b1e18b7e549c5249dc706369f8f7242f03731baa

Observation b93518cb-dc78-4422-81fb-50abce96ca5d · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.077389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:726d737b26e9330bfa740a12f57d9b13eb3b97fd94c16e0d672ef7b0ca02911c

Observation d37b096c-0952-4d52-aa29-37be6b437a11 · inbound

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization cites this paper.

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:19:20.527581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T00:19:20.462455Z digest=sha256:5c8b5d977270fddaca05d616b73c43d7153136cf85764874ff38af63eee824b6

Observation 0ccbae95-297b-469a-988f-6eacb9334bb8 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 256

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.580628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:3a78d75af33dc9c9ef99004350a8683b0a33c1b937dd9c28f7fb876b43218636

Observation 7df89b55-a208-4c96-8197-b601b33a9e54 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:59:03.395929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:9450e3a86d327996908ac15e3d8f5298a8d15bebb6faacbb81988969cc865485

Observation 8ef85ad1-efd7-4ce8-b69a-92e49a222d70 · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:23:51.709217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:4a8fb868e4659c326e1927cba04658347ff4e3dbda56e84d83b063eac299d9f5

Observation 954eea10-c87b-43cf-9ec2-6b1c8ad09be5 · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:23:42.186601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:0a30a8270c7651d35ca60184bc107566adbdf93eb0937361b9fdd3f52bd3759f

Observation 703e8ac5-584c-48aa-bd45-2f71d209e61b · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.771573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:5e7dae2f99a9096ab301fc818acf24aa86f4bdfb1afac5292aca16fb1587fe69

Observation 74420e64-e9bf-490e-9bad-8518905b5848 · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.232126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:3d3ce59d493f3824cb9e848a2ab6b05603dfaca19211ca6b57505e4ca483f61c

Observation 535af07f-bec3-4081-8631-57ba21dd4d66 · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:03.775772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:03.775772Z digest=sha256:1a5283b02da547ea4022eab5cb191ce147b59a54fcb1af8c4088ca2f9bae541b

Observation f099e45b-584b-4429-90f8-7bfef4ee4b09 · inbound

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start cites this paper.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.708467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.708467Z digest=sha256:5eb40cc45acce5c9f1e197d6ea53961df1b4d34acf0210aec33a61ef6c01acbc

Observation 4cc616f9-4037-4362-bdce-c9897d81045c · inbound

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding cites this paper.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:44.530844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:44.530844Z digest=sha256:69215bc77927f5de1fafd2d187495bd3e65285d43091b8e0013e60ca7e3e33e9

Observation 2e5b393c-95b4-44a1-8b11-cd619c3751bb · inbound

Multimodal Tabular Reasoning with Privileged Structured Information cites this paper.

Multimodal Tabular Reasoning with Privileged Structured Information MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:59.200727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:59.200727Z digest=sha256:70cd89c0d25c9d273b260ad1cb661b7f17a082bcee9da2375fd4e3a5500b877d

Observation e5b51c26-c286-46df-9016-544557ae8b24 · inbound

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification cites this paper.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.600959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.600959Z digest=sha256:25045f5eb16355c9f8719921b27a2b3cac0939b965c2ebf619394b412b65aca8

Observation 545bdbd5-21ae-42ef-983b-af63ca588c00 · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:18.908261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:18.908261Z digest=sha256:f0920ca06ba9af6f0381d3c3a5963ebca4616431bab3ada710864ca1c6013afd

Observation 456f258d-cdda-45e3-a431-c54fb8de5b8a · inbound

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training cites this paper.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:03.216795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:03.216795Z digest=sha256:97cf13889e18d0831c91696ca7402e64253561be165caac7fcb324931360754a

Observation de573936-c557-4175-a5a9-a506d29443bb · inbound

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings cites this paper.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.687226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.687226Z digest=sha256:8556f5a171444c859268ebc0385703a85b2dcc2d9ff79298e32f7f497ade51a6

Observation 4a6c5b47-e72a-42bd-911d-b726c8c2e9cd · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.353274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.353274Z digest=sha256:d79eab3903143f9d463891aaacd5cb50af793f5b7023538c3ae35e36dd46bd2b

Observation 4d4fd369-42ff-48dd-aa16-8a3b2a588772 · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.742706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.742706Z digest=sha256:ee9582347855ccbdec4575329d25a0de077ec93675c516d7c76f33898b4f7d62

Observation 48f42d94-ef92-4dc4-9138-292463bcf64b · inbound

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs cites this paper.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:12.605646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:12.605646Z digest=sha256:d9784e424d7ec9e1ae69395cc1f04e579746decb1557cfa67273197f1914d5f6

Observation 6fb502f8-5f62-4b9c-91fe-4b1e13d4cbbe · inbound

A Survey on Diffusion Language Models cites this paper.

A Survey on Diffusion Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-05T20:15:26.877320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:15:26.877320Z digest=sha256:4aeba26cb53abe3771245be8458a15f38a953786030115387021e74a1b604196

Observation b685855f-c7de-430b-a29e-3eb89d776736 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.399950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.399950Z digest=sha256:42c811dbba95dca6bc4677c078048a28b776b2c5d0370b3d57bd98af21ea4ac8

Observation 262f9d64-69a0-4f8f-bde8-be33d8710e31 · inbound

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation cites this paper.

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T15:39:35.481243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:39:35.481243Z digest=sha256:dc728940f3ac9c71bbb9c887a163cc295bea67347cec4476b79ff7cb9aa3aa86

Observation 0cdb59a1-27ed-4549-b01c-685666eba19b · inbound

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model cites this paper.

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:19:00.503308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T04:17:07.534291Z digest=sha256:163b62a739cf547fa06064881b1fea04c18a85815e2be4ce5ad99e77d9a162fd

Observation 422dbeb9-5e7f-4fe7-90ca-88f2e6ae6291 · inbound

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models cites this paper.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.491974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.491974Z digest=sha256:01e1d6afcdcabafc6cbd812b133920681ebd88cb6bdcb3695131fcfac5fec65c

Observation ddbf47d8-68f0-498c-9092-dccac2e046ba · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.394620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:260483cf2f837ed0d2cd56fa4d3a31357ee2b5530b9178249ddfcb9e8e37d95a

Observation 0f5eeb8b-5a5a-402f-8a76-d3fcc9128ac7 · inbound

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model cites this paper.

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:57.788020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:29:12.064221Z digest=sha256:72eae6c69ffe3ab5b5566d4e473551d9ce668f40c90b70b67e2ea856db5bb835

Observation 7cd4543e-76bc-4aef-82f1-3d065e399684 · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.313335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:e05a4b09bab7afe6129fcd1372e9285c6a407c043dfccd31afa42079302702e5

Observation 76be5cd6-6450-49f2-94c9-332626a6307e · inbound

ZAYA1-VL-8B Technical Report cites this paper.

ZAYA1-VL-8B Technical Report MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:23.557358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:15:16.607346Z digest=sha256:8d9a61130f3bea3e9e87fda7c09d68d5217fa868ab7093845c00016f86e710eb

Observation 6e757b67-7e18-4a2b-886d-e421665c2846 · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:57:09.462443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T02:52:43.674969Z digest=sha256:83729b64c71fe086756ff6e4e0c22141581b8601269a51c801efcdc9f32bb767

Observation 17807cb5-4354-418f-bf11-a74de5b5aafc · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:29:28.594707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T21:28:37.680681Z digest=sha256:b53bb6a572e19d1d4b948ac9465170a23dd1d39ed7d37e88e572b9dc7b023e48

Observation d4359dae-d5c6-4551-9dd7-9321babad716 · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:48:23.449898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:21f8ba9a6fb810370f558bc907eb88e97f7b1bddd5b6fe1e49bf2393ec53464a

Observation 6fce37bc-1f7d-4864-8b6a-863019af4870 · inbound

Zamba2-VL Technical Report cites this paper.

Zamba2-VL Technical Report MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.687479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T22:34:20.970856Z digest=sha256:b828d4b78a199172d374f4c9e1d82a57284200ce8efdf9b20e07b71bf4b2fee6

Observation f408383a-4775-42a0-928c-b751d250b4af · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 165

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.896690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:4cbc73afc2cf7b477ee015005b3bc877b8c5afea7be8ddcb5525dbcc1d5f31d6

Observation 53f7f1b4-06df-4382-85c3-14791267f844 · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.156125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-25T21:05:36.836361Z digest=sha256:c2282952119942f7ca837fb9d9b3d9b3b479f319ee3959c72613799440cff515

Observation a48d31d2-7767-4a25-86a5-01e2722079ac · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.583796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T06:30:27.178950Z digest=sha256:5a0da22b3051319d6ab4d47e703184bd22b42ec043fd38d7ad273cc9d840d72b