Pith. sign in

Paper Citation Record · LEDGER

BabyVision: Visual Reasoning Beyond Language

As of 4 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 28 inbound Pith citation observations for arXiv:2601.06521.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.06521 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:25:37.168330Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:50:20.183933Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T09:37:00.824668Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2de732fb-c481-4680-b05a-50591f5ae38a · outbound

This paper cites Accessed: 2025-01-09.

BabyVision: Visual Reasoning Beyond Language Accessed: 2025-01-09

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.254842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.254842Z digest=sha256:823b2e114a62489c50a4bea59cd06860f8ebc9d7f094b094b878d8029df9e883

Observation bfb5f875-3479-4275-8e03-d259b1939353 · outbound

This paper cites Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, and Feng Zhao.

BabyVision: Visual Reasoning Beyond Language Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, and Feng Zhao

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.324358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.324358Z digest=sha256:b31bfe788282e17ea03d7a7bc10517be821242452a2b6975cdaa217de8acb02e

Observation 801ef178-dbe7-43fd-9ae7-febbc2b9b640 · outbound

This paper cites Accessed: 2025-01-09.

BabyVision: Visual Reasoning Beyond Language Accessed: 2025-01-09

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.583244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.583244Z digest=sha256:988da4f26d36f3fd350be792d6174b194c167119e24c4a847fd448b5878be9d2

Observation 4af5246c-d98b-474a-b118-4a42e8a91ecb · outbound

This paper cites MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation.

BabyVision: Visual Reasoning Beyond Language MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.655028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.655028Z digest=sha256:da2235fb69ee5d02e803dfc4be05c15d82bbd072051308cfca993b3905eab68c

Observation d0e1e774-3d1e-47c6-846e-c86c5eabbe7f · outbound

This paper cites Accessed: 2025-01-09.

BabyVision: Visual Reasoning Beyond Language Accessed: 2025-01-09

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.738689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.738689Z digest=sha256:942059b56171fb9f1144be03ee2208600405043cef050e297fba2a01b4b60e42

Observation d644983c-2a3e-4d98-8be0-1c4e66419cdc · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

BabyVision: Visual Reasoning Beyond Language HybridFlow: A Flexible and Efficient RLHF Framework

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.780815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.780815Z digest=sha256:3966c89a72697d151d3bf11c628600cb48c4ba0b98e7e0bcedfb4934b4a1678c

Observation bdaee791-7db0-4c97-bf96-ee1f340191ed · outbound

This paper cites Core Team, Zihao Yue, Zhenru Lin, Yifan Song, Weikun Wang, Shuhuai Ren, Shuhao Gu, Shicheng Li, et al.

BabyVision: Visual Reasoning Beyond Language Core Team, Zihao Yue, Zhenru Lin, Yifan Song, Weikun Wang, Shuhuai Ren, Shuhao Gu, Shicheng Li, et al

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.862991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.862991Z digest=sha256:1a6e82bc208d1ac482a68dbe1b383ea412023dd08c5f5ee2d64b37c39fdb2473

Observation fc4343a4-c7c2-4d2a-a26e-0760b32da638 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

BabyVision: Visual Reasoning Beyond Language Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:37.014871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:37.014871Z digest=sha256:09b1407283b4978260a374980f278bb82d75d7bb2d571a040ed20f2b29b462d7

Observation f743877f-9b8b-4328-956c-c8e6acb7a2e5 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

BabyVision: Visual Reasoning Beyond Language InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:37.086374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:37.086374Z digest=sha256:3065687fba2bb788fa3232664a741e1e504a3071b8bb58d7c667fd1b227ecd9d

Observation e54c4f02-70dd-455d-bed4-3f1f0daa5138 · outbound

This paper cites DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?.

BabyVision: Visual Reasoning Beyond Language DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:37.168330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:37.168330Z digest=sha256:020dd20e820970dfa6c925a1fd48502d6f696cf55f5ebc5d7cfcf70a6e10bb4f

Observation 241372c7-3494-44a1-a7be-4c5019d3e86e · outbound

This paper cites Humanity's Last Exam.

BabyVision: Visual Reasoning Beyond Language Humanity's Last Exam

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.947259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.947259Z digest=sha256:bd768df15efd18d52595a9f63515f55ae14d90a9d03bcaa48dcfe00ffdfa98c7

Observation 4882c10a-7f71-48fa-8989-421a225fdd51 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

BabyVision: Visual Reasoning Beyond Language MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.510939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.510939Z digest=sha256:24cfce6b60c0eb8e63e0722cceb0a5b76e29f3fb46aba6ee503a3177dfe0bec1

Observation 768ad5cf-84db-4e70-a90b-7928f392fe03 · outbound

This paper cites Qwen3-VL Technical Report.

BabyVision: Visual Reasoning Beyond Language Qwen3-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.210179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.210179Z digest=sha256:84161fe1993a71d5ce55952114e17d6351dc2efff674069fe28dadb0377dc9c0

Pith citing papers

Observation ecbae974-0770-40b4-b6da-45b8cb86e6d2 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence BabyVision: Visual Reasoning Beyond Language

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:b28215ef18d82d0e20598e1cfeaf9053a42ff97a2ea3d2902107d597e82a3ae3

Observation 3be3dc69-02b2-43c8-ba62-ad5c117f3a99 · inbound

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors cites this paper.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors BabyVision: Visual Reasoning Beyond Language

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:0949db1860017d3589b9afbeccfd315d6aca377b67028b21a7cee5823ca7b32f

Observation 9846bd0a-5d1f-42ee-ae57-920cd691042d · inbound

EXAONE 4.5 Technical Report cites this paper.

EXAONE 4.5 Technical Report BabyVision: Visual Reasoning Beyond Language

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:47:34.692414Z digest=sha256:d818a3e3784a5f16c9a1c5112ab934d945bf6c29f182c0ed50233b60941fc614

Observation 37bffe5c-46df-406f-b6b1-b7718587d055 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm BabyVision: Visual Reasoning Beyond Language

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T23:56:02.856878Z digest=sha256:e28f03c4adbb6dfdb6d1c20ffc922ed89ff1f40b3c9cde6abc17d51941078bb7

Observation 0eb02fec-f321-4728-bf03-db7c1a25d8e8 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm BabyVision: Visual Reasoning Beyond Language

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:2ca5c3dea9dcaf2585e7cd025480ce7fc18df13e7600ada7c7a4550b408a279c

Observation 5cbe58c9-2535-4f0f-b992-fe8167643885 · inbound

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression cites this paper.

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression BabyVision: Visual Reasoning Beyond Language

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T14:36:29.666730Z digest=sha256:e7a7002784feed85910d569b2d38dd92b4275b070249f9b6e600b53b1bc36690

Observation f8c77675-e600-4f98-9826-bd26e24b018d · inbound

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression cites this paper.

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression BabyVision: Visual Reasoning Beyond Language

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T04:52:09.685243Z digest=sha256:b602116f0b2dbf9980522308eeec1a7bbe5612eddcb29d68be7790294e272972

Observation 51fcbd8a-30ee-41a5-9c1b-2a117f04ae4d · inbound

Do multimodal models imagine electric sheep? cites this paper.

Do multimodal models imagine electric sheep? BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:24:01.933339Z digest=sha256:34a2f65322aad4e7b89945e2d772844a82aa663790dd1dd624ce8d8b5dd77c05

Observation e69986e0-e556-419c-b975-eceda920f3c1 · inbound

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space cites this paper.

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:30:54.053958Z digest=sha256:47b38e0d587bc77dcbd12c0828a7588bdc9481dd42300d1df8d9cab27099d75f

Observation 17c6eedd-5c6b-498e-9742-17acee3eae00 · inbound

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space cites this paper.

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T22:58:04.574536Z digest=sha256:e7e18259cec7ff9de722ed54b146493dc11f53af6a603f00821f5296284a1d4a

Observation 2eecc1b5-7e55-4014-b21c-e4c0d30ef38b · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture BabyVision: Visual Reasoning Beyond Language

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:fed688e2205d608cf9582848b90fff641ec995c154ced57872c883a7611e1653

Observation b2c596a6-d638-446c-a9c5-1e75c5c40fd1 · inbound

VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following cites this paper.

VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following BabyVision: Visual Reasoning Beyond Language

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:14:34.739062Z digest=sha256:c7edd2622d8651b20b396711f3408236e562759ceddc1b75d3cdef0f399579d9

Observation 5851d4f6-769c-4b6f-92e9-0d896e34dc6a · inbound

Step-wise Rubric Rewards for LLM Reasoning cites this paper.

Step-wise Rubric Rewards for LLM Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:4497c50daed27f160e6fc4f6945920a9f13728ff5b8fc06d7055aec42b9e2551

Observation 2127c17d-1a86-413a-9736-bdbbf64a8041 · inbound

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data cites this paper.

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T12:12:45.251924Z digest=sha256:7f13c5694f6d0ccb5792c7ebc1fb182b4449fc11baa1b3c8060fa900157041dd

Observation 04dbbbd5-7c41-42c4-8c80-707d31f8d328 · inbound

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding cites this paper.

Bridging Structure and Language: Graph-Based Visual Reasoning for Autonomous Road Understanding BabyVision: Visual Reasoning Beyond Language

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T05:25:20.511933Z digest=sha256:cf3bbdca006216c041af708efd9e72eb3488ea0de673acb6dd740cd57e69e265

Observation 3ebebce8-606e-4713-9819-be895d60852c · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 108

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:efa23cc0938a0d4ed1a0bc4ca82bc8fc5b4f1bbcfec3ba6412f1c34ac8a05395

Observation 187c584c-1f2d-406b-94cd-aa8f92e5307c · inbound

ATLAS: Agentic Test-time Learning-to-Allocate Scaling cites this paper.

ATLAS: Agentic Test-time Learning-to-Allocate Scaling BabyVision: Visual Reasoning Beyond Language

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T15:27:28.290178Z digest=sha256:42067a590f03e25e784e9f8b7fbd799a0bfede98db37eda55162413bffd493b8

Observation abdb3657-0c8e-497f-a05f-882bf39451c9 · inbound

LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?") cites this paper.

LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?") BabyVision: Visual Reasoning Beyond Language

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T06:39:08.246174Z digest=sha256:1cf78c238f4ca67ea4324d1147fbbbec760b3370be27cc0ee40714956d61edd1

Observation 381440d4-1561-4548-a217-0477d06c4de3 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients BabyVision: Visual Reasoning Beyond Language

Reference 131

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:077a50570c4a71775eec5a3a1819f9d0212c7bced9d868c5cff80848de95a661

Observation 27ad2c73-da42-4e3e-a21a-860c8bceef53 · inbound

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning cites this paper.

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-25T21:23:10.051805Z digest=sha256:685ac564fee298d97689581473ad075ea01909ce17df06d5419361e62e4ebe58

Observation dba8cbd6-253b-42d7-a467-25bf97e808a5 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity BabyVision: Visual Reasoning Beyond Language

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:565cf4ee33ab27914fcabf658dd1d727d2c82eebb375f20c8831ec00aa806568

Observation 71f17e32-ee9b-4509-b9e9-2e4ec44e263e · inbound

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models cites this paper.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models BabyVision: Visual Reasoning Beyond Language

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T09:37:00.825956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T09:30:46.564013Z digest=sha256:aac5088645b310ce3a93e6000e16ea37cd9afc49afba23542dde517b6de27e83

Observation 5f823fbd-a930-48f6-a4e3-720357328807 · inbound

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning cites this paper.

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T05:37:19.632869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T05:37:19.632869Z digest=sha256:3b81d7934474b60808798163e85636f276695533e17a52bf46e27fda5d4dabab

Observation 17b1b511-f44e-4f07-b535-a58fa7242f42 · inbound

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning cites this paper.

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T07:01:19.711542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:01:19.711542Z digest=sha256:73dc0436d7fa5ff291381c76918fd98287f022c2ad0f8ca48ae1f980a567f989

Observation c4384486-1fbb-4686-9187-89c420c32b05 · inbound

An Exam for Active Observers cites this paper.

An Exam for Active Observers BabyVision: Visual Reasoning Beyond Language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:03.005752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:03.005752Z digest=sha256:087241a0f416cbad563b5f93ea584985ca8c992753dbeac1519ded601f8b2e3f

Observation 747b163b-6bef-4176-97fe-c820f1d48fc4 · inbound

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection cites this paper.

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection BabyVision: Visual Reasoning Beyond Language

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:20.183933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:20.183933Z digest=sha256:7614a71479bd70e30d30bcb3d52ef3c5def7ea45697306c73cb93f59d08052f7

Observation d51a5525-1455-4600-a670-736d17b93e09 · inbound

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models cites this paper.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models BabyVision: Visual Reasoning Beyond Language

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T05:01:28.838169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T05:01:28.838169Z digest=sha256:a727aa70f6266a1a4873c36544e5c216ec30f8a7245a73e77eb485a3b0018ac7

Observation 52f1f3d1-c9a0-418c-9ea2-43d2dd02a6fd · inbound

Beacon: Knowing When and How to Perform Agentic Visual Reasoning cites this paper.

Beacon: Knowing When and How to Perform Agentic Visual Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T02:45:28.580613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:45:28.580613Z digest=sha256:dfd7cf91381a2566ed674024ccee701d5a41446f0919143b6c4154a59b9641a9