Pith. sign in

Paper Citation Record · LEDGER

BabyVision: Visual Reasoning Beyond Language

As of 21 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 31 inbound Pith citation observations for arXiv:2601.06521.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.06521 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:25:37.168330Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:15:39.766849Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T09:37:00.824668Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2de732fb-c481-4680-b05a-50591f5ae38a · outbound

This paper cites Accessed: 2025-01-09.

BabyVision: Visual Reasoning Beyond Language Accessed: 2025-01-09

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.254842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.254842Z digest=sha256:7ed6b29a9cfe65c49afb9d5f69efc6f0d950a5c401d3f4519ad952f48d1ccfc4

Observation bfb5f875-3479-4275-8e03-d259b1939353 · outbound

This paper cites Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, and Feng Zhao.

BabyVision: Visual Reasoning Beyond Language Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, and Feng Zhao

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.324358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.324358Z digest=sha256:5eee330c73d24fa3da1e7019f99ca984c24837f79e28b28d4eea2932f0c3a815

Observation 801ef178-dbe7-43fd-9ae7-febbc2b9b640 · outbound

This paper cites Accessed: 2025-01-09.

BabyVision: Visual Reasoning Beyond Language Accessed: 2025-01-09

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.583244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.583244Z digest=sha256:56399d28aafb42e362f41171f8a2f8d4c798e870b4b277a971e3c7cb3a829426

Observation 4af5246c-d98b-474a-b118-4a42e8a91ecb · outbound

This paper cites MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation.

BabyVision: Visual Reasoning Beyond Language MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.655028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.655028Z digest=sha256:ffd690ceb8adb55d7a4708f684075a04d1b4e560ba743eccaa32fc1b7833fac8

Observation d0e1e774-3d1e-47c6-846e-c86c5eabbe7f · outbound

This paper cites Accessed: 2025-01-09.

BabyVision: Visual Reasoning Beyond Language Accessed: 2025-01-09

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.738689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.738689Z digest=sha256:2b5cb4e542c9ccc64529ffa3ac433ecc3231611d2f5144e4a0faabce1cdd7772

Observation d644983c-2a3e-4d98-8be0-1c4e66419cdc · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

BabyVision: Visual Reasoning Beyond Language HybridFlow: A Flexible and Efficient RLHF Framework

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.780815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.780815Z digest=sha256:63b5c360534b8e18b2284f6c52c9916c0fdba1bed09d06b173a0dde647dbac41

Observation bdaee791-7db0-4c97-bf96-ee1f340191ed · outbound

This paper cites Core Team, Zihao Yue, Zhenru Lin, Yifan Song, Weikun Wang, Shuhuai Ren, Shuhao Gu, Shicheng Li, et al.

BabyVision: Visual Reasoning Beyond Language Core Team, Zihao Yue, Zhenru Lin, Yifan Song, Weikun Wang, Shuhuai Ren, Shuhao Gu, Shicheng Li, et al

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.862991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.862991Z digest=sha256:e9dfa4409c11d4d0532d64afafbb4518a0a95bcb4a03b9c745a329b8311f6889

Observation fc4343a4-c7c2-4d2a-a26e-0760b32da638 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

BabyVision: Visual Reasoning Beyond Language Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:37.014871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:37.014871Z digest=sha256:5359f01e3d2548ccaa5b2bbf7b9a5f8ccaac036f6428ab3a5aa0d81b02f37c94

Observation f743877f-9b8b-4328-956c-c8e6acb7a2e5 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

BabyVision: Visual Reasoning Beyond Language InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:37.086374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:37.086374Z digest=sha256:50463ebc42a7fef9e0542fc4b423c835a2172f2707d2b017188b0714b7a35d11

Observation e54c4f02-70dd-455d-bed4-3f1f0daa5138 · outbound

This paper cites DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?.

BabyVision: Visual Reasoning Beyond Language DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:37.168330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:37.168330Z digest=sha256:3667ef977c95e4cd2b81852a29ad844241fa11a26884f41af6e628cba80decf8

Observation 241372c7-3494-44a1-a7be-4c5019d3e86e · outbound

This paper cites Humanity's Last Exam.

BabyVision: Visual Reasoning Beyond Language Humanity's Last Exam

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.947259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.947259Z digest=sha256:09e0860d9e347a185f33e17ed66e75cbe781b9880ac1a0a49d694621091c301b

Observation 4882c10a-7f71-48fa-8989-421a225fdd51 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

BabyVision: Visual Reasoning Beyond Language MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.510939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.510939Z digest=sha256:64086fa615d4e1ed17df0a786dd1597682413aae38cee47c55aa94185ba4df69

Observation 768ad5cf-84db-4e70-a90b-7928f392fe03 · outbound

This paper cites Qwen3-VL Technical Report.

BabyVision: Visual Reasoning Beyond Language Qwen3-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.210179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.210179Z digest=sha256:37915db2fc01ec611bf2b2bcdc0de2b743bd56457ee3f8d826f4d9d0e6e0efd9

Pith citing papers

Observation ecbae974-0770-40b4-b6da-45b8cb86e6d2 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence BabyVision: Visual Reasoning Beyond Language

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:a4ad722022aff9e0aac33d146eaa694d25c92c28e4ddf618a90195a635aeee22

Observation 3be3dc69-02b2-43c8-ba62-ad5c117f3a99 · inbound

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors cites this paper.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors BabyVision: Visual Reasoning Beyond Language

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:e354223d12cc5d0f16f34fd6de253ccd82b3db9d9a729b0d62519cf99dd98092

Observation 9846bd0a-5d1f-42ee-ae57-920cd691042d · inbound

EXAONE 4.5 Technical Report cites this paper.

EXAONE 4.5 Technical Report BabyVision: Visual Reasoning Beyond Language

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:47:34.692414Z digest=sha256:804044405c40b19a0201e6aaa3584e813bfc2d6738a94c4fd1584b623456e5ee

Observation 37bffe5c-46df-406f-b6b1-b7718587d055 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm BabyVision: Visual Reasoning Beyond Language

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T23:56:02.856878Z digest=sha256:f3787ee83398559fe54d8018a8455bcf847689360c2b4be16f60c67521f3d13f

Observation 0eb02fec-f321-4728-bf03-db7c1a25d8e8 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm BabyVision: Visual Reasoning Beyond Language

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:d97d518a6a727f7edc27a8c55d2801ee0a064c273ad092f51dd1f2172fa3d719

Observation 5cbe58c9-2535-4f0f-b992-fe8167643885 · inbound

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression cites this paper.

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression BabyVision: Visual Reasoning Beyond Language

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-09T14:36:29.666730Z digest=sha256:18ed0c82c6b6291d40369616ae7295b32ed47583ff1c8a32a7574dc25762268f

Observation f8c77675-e600-4f98-9826-bd26e24b018d · inbound

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression cites this paper.

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression BabyVision: Visual Reasoning Beyond Language

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T04:52:09.685243Z digest=sha256:ad7dbc9af54db55944c5f057db73ff76121f751c5a603d7225ec06eb3e8bb63c

Observation 51fcbd8a-30ee-41a5-9c1b-2a117f04ae4d · inbound

Do multimodal models imagine electric sheep? cites this paper.

Do multimodal models imagine electric sheep? BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T03:24:01.933339Z digest=sha256:1757765e5c5aa335f0507e13ead6da493cd8804ce7df55697da53c6d77347c64

Observation e69986e0-e556-419c-b975-eceda920f3c1 · inbound

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space cites this paper.

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:30:54.053958Z digest=sha256:2eb5e94792050f1a518d2865be1ad4220b8dbca8aa1fe9f297dfb22575190896

Observation 17c6eedd-5c6b-498e-9742-17acee3eae00 · inbound

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space cites this paper.

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T22:58:04.574536Z digest=sha256:b07dfef7b05b327e968e5e665b079f0082bb3608e94762f39c84925cc7634297

Observation 2eecc1b5-7e55-4014-b21c-e4c0d30ef38b · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture BabyVision: Visual Reasoning Beyond Language

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:6d00f1c440ca71e99494b664c91a7ef5bd08ce53b3e036b9a732c7555314a0ea

Observation b2c596a6-d638-446c-a9c5-1e75c5c40fd1 · inbound

VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following cites this paper.

VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following BabyVision: Visual Reasoning Beyond Language

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T20:14:34.739062Z digest=sha256:be98325a2a3f9939989e8e1ec47831475060ddbcfb75828cb530673f33960427

Observation 5851d4f6-769c-4b6f-92e9-0d896e34dc6a · inbound

Step-wise Rubric Rewards for LLM Reasoning cites this paper.

Step-wise Rubric Rewards for LLM Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:c7996b41bac19ffb5b90d2bcdb2aef9a40bfc8922898165aac0f07d2e7505e8a

Observation 2127c17d-1a86-413a-9736-bdbbf64a8041 · inbound

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data cites this paper.

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T12:12:45.251924Z digest=sha256:58a42acb5ca2d03972b5edf5877cbc27d7e85a5f6bfbc6f762883155fd23fbb9

Observation 04dbbbd5-7c41-42c4-8c80-707d31f8d328 · inbound

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding cites this paper.

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding BabyVision: Visual Reasoning Beyond Language

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T05:25:20.511933Z digest=sha256:01e725ec1b6411291971abb8bc798cb6f6ac57353d3f8b2b0e692e1c52f66c63

Observation 3ebebce8-606e-4713-9819-be895d60852c · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 108

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:e8cfef8168501fd34ae77f4b8023f10150ba2df73ef47138822ec06bfbbf1ea0

Observation 187c584c-1f2d-406b-94cd-aa8f92e5307c · inbound

ATLAS: Agentic Test-time Learning-to-Allocate Scaling cites this paper.

ATLAS: Agentic Test-time Learning-to-Allocate Scaling BabyVision: Visual Reasoning Beyond Language

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T15:27:28.290178Z digest=sha256:1c03f01fbb3d558f8fe848f5428c5ce2f920cfe48c9a74cfab66663cd3b0f80e

Observation abdb3657-0c8e-497f-a05f-882bf39451c9 · inbound

LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?") cites this paper.

LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?") BabyVision: Visual Reasoning Beyond Language

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T06:39:08.246174Z digest=sha256:c6fb8770193df197ddc2af044bbde13d3bae76a96a9c1165a6c26cf190ba738f

Observation 381440d4-1561-4548-a217-0477d06c4de3 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients BabyVision: Visual Reasoning Beyond Language

Reference 131

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:27476ef6bb985c1465e055d69096fc2d00caa437a77367e0983b2ce4d6495e65

Observation 27ad2c73-da42-4e3e-a21a-860c8bceef53 · inbound

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning cites this paper.

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-25T21:23:10.051805Z digest=sha256:2916e31838ff5c2225219df57d6bafa035f3d4575c6a3a00632980ea32841774

Observation dba8cbd6-253b-42d7-a467-25bf97e808a5 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity BabyVision: Visual Reasoning Beyond Language

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:c3410611b642cf23cb3ad5de71aeaa9b2cf662bc5197cc262091f680872f1e75

Observation 71f17e32-ee9b-4509-b9e9-2e4ec44e263e · inbound

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models cites this paper.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models BabyVision: Visual Reasoning Beyond Language

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.825956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:2393208bd36dd7d79575b30ffad63827487547d22633aacebd0227bf8dcd7e3f

Observation 5f823fbd-a930-48f6-a4e3-720357328807 · inbound

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning cites this paper.

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T05:37:19.632869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T05:37:19.632869Z digest=sha256:9a7d831e33e27bb957b3152ef4ccefb0925acbf5d49699cabe4c6b3962a4705f

Observation 17b1b511-f44e-4f07-b535-a58fa7242f42 · inbound

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning cites this paper.

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T07:01:19.711542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:01:19.711542Z digest=sha256:6a81ca9c441559f6369e77b07f1619d752d90336fbfefc2c7991d553b535b71a

Observation c4384486-1fbb-4686-9187-89c420c32b05 · inbound

An Exam for Active Observers cites this paper.

An Exam for Active Observers BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:03.005752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:03.005752Z digest=sha256:e7b001ddb9d496874b003eca5b0c5dea2b1b29bb2972263d1e8524540072bc66

Observation 0986a07a-a4f6-43b1-a070-8b88c1d6301c · inbound

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests cites this paper.

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests BabyVision: Visual Reasoning Beyond Language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:32:27.377351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:32:27.377351Z digest=sha256:a8b53bb0f26f95b97a115d43c46eaa4651e63d3d992066c01574a647f7610eb9

Observation 747b163b-6bef-4176-97fe-c820f1d48fc4 · inbound

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection cites this paper.

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection BabyVision: Visual Reasoning Beyond Language

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:20.183933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:20.183933Z digest=sha256:41ecb447fbd01a4ede4fae6c012a78522a2fd3ccad03282781631b6f2d26dc12

Observation d51a5525-1455-4600-a670-736d17b93e09 · inbound

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models cites this paper.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models BabyVision: Visual Reasoning Beyond Language

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T05:01:28.838169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T05:01:28.838169Z digest=sha256:5f3587c084aad4940dd3debe42c8ff07babe6efa01ffb9bb1a10bd1d7a03a555

Observation 52f1f3d1-c9a0-418c-9ea2-43d2dd02a6fd · inbound

Beacon: Knowing When and How to Perform Agentic Visual Reasoning cites this paper.

Beacon: Knowing When and How to Perform Agentic Visual Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T02:45:28.580613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:45:28.580613Z digest=sha256:0b4dc6c9d5b9117a366e1d13272a408f5df5110d05f41cc0f032f557bada25a9

Observation 16cf6ca3-e2f6-4e97-8a37-0cf641434036 · inbound

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making cites this paper.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making BabyVision: Visual Reasoning Beyond Language

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.121965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.121965Z digest=sha256:196bc7a64740c5458296255adfdfdcbc92ea1c491fe472373d45287606086b0b

Observation 4069ec4b-08de-4012-854e-4aafb89148dd · inbound

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams cites this paper.

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams BabyVision: Visual Reasoning Beyond Language

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:39.766849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:39.766849Z digest=sha256:a7338b42e1b48e46bb56f8c95b75724d1ee0c0a456251018658c6fafaff587c0