Pith. sign in

Paper Citation Record · LEDGER

RADIO1D: Elastic Representations for Condensed Vision Modeling

As of 8 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2607.03624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.03624 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T01:07:20.766474Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66caf685-be41-48be-a201-9ebaf35dc24a · outbound

This paper cites Learning transferable visual models from natural language supervision.

RADIO1D: Elastic Representations for Condensed Vision Modeling Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:bdfb172b9948b7db44770a259b368aacd99f9b7434dee2c7bb3b3e8154ecbb93

Observation 649677a9-321a-42cc-99eb-bb259d835013 · outbound

This paper cites BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

RADIO1D: Elastic Representations for Condensed Vision Modeling BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:08bfd200acc47db01f581ebc2254e4f3c3fe33a64669df5b7faf99a744577ee0

Observation 83d05092-8465-40ae-a363-34acb78c5246 · outbound

This paper cites Language is not all you need: Aligning perception with language models.

RADIO1D: Elastic Representations for Condensed Vision Modeling Language is not all you need: Aligning perception with language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:e6d0ee3af1d0a052fb952708c573bd3ee579b5673bc2cb54ee1683db6b201d64

Observation 2729808d-b74d-4414-852e-f4de603e2957 · outbound

This paper cites Visual instruction tuning.

RADIO1D: Elastic Representations for Condensed Vision Modeling Visual instruction tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:b5c4b1b7f60c36c440bb5422cc682653564f82c621f8de7f14622fddaa081407

Observation cc0c6dda-8b77-4550-afd5-0a09274a61e1 · outbound

This paper cites Sigmoid loss for language image pre-training.

RADIO1D: Elastic Representations for Condensed Vision Modeling Sigmoid loss for language image pre-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:526f5323ed294cd828bba12e35ee3be46dc125a96947dbf77007dbd1c1f4a52c

Observation b4e8b263-c779-45a5-8665-88c5eeb5aa94 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

RADIO1D: Elastic Representations for Condensed Vision Modeling SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:61ad69ab9d01356527baaf92975ace171c62b179cd608a82ab1f21bf367f2086

Observation ff0e86db-0c8c-4716-aee0-47e46c9bb751 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

RADIO1D: Elastic Representations for Condensed Vision Modeling PaliGemma: A versatile 3B VLM for transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:0e3bebb6305d6c6410d77f859ff50031b02656caba4869e019365f01ad5ec291

Observation 3425b0fc-02ef-4160-84f0-d0889b1b3e97 · outbound

This paper cites What matters when building vision-language models?.

RADIO1D: Elastic Representations for Condensed Vision Modeling What matters when building vision-language models?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:08b2a2691d6df769deae2a99400a832f2081c9331a58cf607d0d51cd920d397f

Observation 6f740f26-d7de-4d0d-aeb2-0c7793c251cd · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

RADIO1D: Elastic Representations for Condensed Vision Modeling Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:b9ebdbcdc034382349eba88262b6a77f48f4307008b7be79e61c051586ce90cc

Observation 89826ebe-5819-4ffd-a570-9dbf1521d10b · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

RADIO1D: Elastic Representations for Condensed Vision Modeling NVILA: Efficient Frontier Visual Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:3a56c4c107c89a4b9c7906e714c5710267df7bdac6fdf4be26a7d33fb15f4a99

Observation 981597fb-7f4b-4ae1-89d9-0c02a279d687 · outbound

This paper cites Qwen3-VL Technical Report.

RADIO1D: Elastic Representations for Condensed Vision Modeling Qwen3-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:35c2dd158f1dd8ba5471467d4d77ec957fa0d76690be369b3a763b565045d040

Observation 7330b42e-8738-4433-8747-a026a6ef14cc · outbound

This paper cites Radiov2.5: Improved baselines for agglomerative vision foun- dation models.

RADIO1D: Elastic Representations for Condensed Vision Modeling Radiov2.5: Improved baselines for agglomerative vision foun- dation models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:d2693d0c74f3ee7401d421d5d0bcdf228d1c0203a1ff0acedff4cb12cb01f1fc

Observation b66eb5b0-fed4-4dbe-8022-4187bf71d53e · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:f07e9841efe99dcc6b6ecb4cb6192dabca4723f8598dfda43f536511dff7381b

Observation 753be084-d9a9-4903-8007-a0d6e3da9b50 · outbound

This paper cites DINOv3.

RADIO1D: Elastic Representations for Condensed Vision Modeling DINOv3

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:bb89fa81ffae0566039a41f0281363c94bb344fc1dfec6d154d019fcae41d2a0

Observation c59013d7-ec72-4d46-9bfc-5f41d561a554 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

RADIO1D: Elastic Representations for Condensed Vision Modeling SAM 3: Segment Anything with Concepts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:9c72f502752b454a209320f1e9dc349f41dba6de2b9ada1cb5cbef458a215bdb

Observation 7486a3ee-c219-41a7-b09e-48ad9bd30107 · outbound

This paper cites Eagle: Exploring the design space for multimodal LLMs with mixture of encoders.

RADIO1D: Elastic Representations for Condensed Vision Modeling Eagle: Exploring the design space for multimodal LLMs with mixture of encoders

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:4f397e261bc0cc9d65f9f0290b5b6887386e1edccc4dd80d267577b15e7df898

Observation 6f75ea9e-ab5a-4abf-9b66-1dda9a5b9129 · outbound

This paper cites VILA-u: a unified foundation model integrating visual understanding and generation.

RADIO1D: Elastic Representations for Condensed Vision Modeling VILA-u: a unified foundation model integrating visual understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:49407577d703562891fdbc9b2e083436123e40ec8f42c913b2935b85cf087454

Observation d816267e-b8c2-4385-b7de-6804e3fa1cc6 · outbound

This paper cites Qwen-VL: A versatile vision-language model for under- standing, localization, text reading, and beyond,.

RADIO1D: Elastic Representations for Condensed Vision Modeling Qwen-VL: A versatile vision-language model for under- standing, localization, text reading, and beyond,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:cec2ca414e8b36887c1bfe445c9ec9b4ec8900bb423c820d75f8f0b4d5cda4bd

Observation b529dc4a-ac26-4eb9-8b0a-1dc38d6b038e · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:b686aef0ae2633dfdd1348c84fdd62f09bde76a0b9802ce56077bbbcabe50204

Observation cdb03c63-ac96-41cd-b557-54e2890f3a6c · outbound

This paper cites LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer.

RADIO1D: Elastic Representations for Condensed Vision Modeling LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:90c3fbe1c60ac71e04506d41ae515426179b803d0e7ff6e341af8969f7b759ee

Observation eec94014-4c85-4388-864c-9077dff19806 · outbound

This paper cites Scene parsing through ADE20K dataset.

RADIO1D: Elastic Representations for Condensed Vision Modeling Scene parsing through ADE20K dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:6050573cfc178532e50d5528247d48ece631e87576d43364ea1d42dff5e08379

Observation 0b670147-5789-4d89-9d6b-49701cf1cbbd · outbound

This paper cites Schwing, Alexander Kirillov, and Rohit Girdhar.

RADIO1D: Elastic Representations for Condensed Vision Modeling Schwing, Alexander Kirillov, and Rohit Girdhar

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:15d8412c10cdf457800eba584704d3bde2a49a925cbe3b632de3d03df9b3bc0f

Observation 3f39cf81-c855-4dfe-81c8-23bd62c7077d · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:51855def333b5243e6d6c1da258e95cca1788ea0ca4284672987db6a2a3edeb1

Observation e3efc387-ff03-40e9-b7d9-0c425878f774 · outbound

This paper cites DINOv3-driven se- mantic segmentation for landslide mapping in mountainous regions.Sensors, 26(2), 2026.

RADIO1D: Elastic Representations for Condensed Vision Modeling DINOv3-driven se- mantic segmentation for landslide mapping in mountainous regions.Sensors, 26(2), 2026

Reference 24

Resolution
verified exact
doi, observed 2026-07-12T01:08:22.975093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:f5a2cb1529910545f9a939035c3e269affe2478c14679a02ef6c1b8628e57ea6

Observation 7b208968-cf6b-45b3-80ee-1b1a2699b973 · outbound

This paper cites Similarity of neural network representations revisited.

RADIO1D: Elastic Representations for Condensed Vision Modeling Similarity of neural network representations revisited

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:5654259ecdabb87ea61f59ab4a84d882133fbfdb8b57aba488b88bb73387e3d3

Observation 3fcef4d6-2d43-484b-907e-78114ff4c0f3 · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:3ec94802c309ccca4816d3f8772cb45d15eb6a2ccfb105db507acb1827025b67

Observation c4fa02a4-4d50-4f62-b8fc-00247176752f · outbound

This paper cites Lawrence Zit- nick, and Piotr Dollár.

RADIO1D: Elastic Representations for Condensed Vision Modeling Lawrence Zit- nick, and Piotr Dollár

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:ca8bca5e47a57b6986265e2e41bc8317b674994879bcca972e2c675fb57273bb

Observation ce04983f-4c12-4b34-9514-9d3c258db8ee · outbound

This paper cites C-radiov4 (tech report), 2026.

RADIO1D: Elastic Representations for Condensed Vision Modeling C-radiov4 (tech report), 2026

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:b7070820a3218ba0f54822028aaf157444f2fd6d16d4312350edafd5a21fe6ee

Observation 5e653ed3-47fd-4bf1-b489-5dd587627d0e · outbound

This paper cites Pereira, and William Bialek.

RADIO1D: Elastic Representations for Condensed Vision Modeling Pereira, and William Bialek

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:51234734363a5a487db8e8fcd56fbf44ce51ed30d98515972e83cb8b611b2cf9

Observation 0b70eefb-7a76-4dc7-9032-7eb2101de28b · outbound

This paper cites Modeling by shortest data description.Automatica, 14(5):465–471, 1978.

RADIO1D: Elastic Representations for Condensed Vision Modeling Modeling by shortest data description.Automatica, 14(5):465–471, 1978

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:6878efd1a4ca3bcf8fd396e73c1a9923649d0e03fa3f8f524dc41c94384ba770

Observation 4f332976-3e04-4e2e-ac85-88a8e7db861a · outbound

This paper cites AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into One.

RADIO1D: Elastic Representations for Condensed Vision Modeling AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into One

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:7e8bbece07dafa13ce4e4e0ed453510e549144e321bbc7146f2c30d9460b7b3b

Observation 2e28156b-c771-4b87-8db6-2829b584e89f · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

RADIO1D: Elastic Representations for Condensed Vision Modeling An image is worth 16x16 words: Transformers for image recognition at scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:e00b78db5fee84030f3166c9d72b65e0967b7cb95b6677ebca199938a4e488b4

Observation 69362b1e-f22f-4eb7-b66c-e219cf945c0f · outbound

This paper cites Flextok: Resam- pling images into 1d token sequences of flexible length.

RADIO1D: Elastic Representations for Condensed Vision Modeling Flextok: Resam- pling images into 1d token sequences of flexible length

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:57558971f4996596a1e59835b90034d395e63c39be0ae2f6958b61d393e1cd54

Observation d0a3f5da-d4fe-4108-a26a-803d177fc429 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

RADIO1D: Elastic Representations for Condensed Vision Modeling Swin transformer: Hierarchical vision transformer using shifted windows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:6f9e2e6a08f27b2633e8591f196353e55d10d950dde547ac9766fa554b422023

Observation 64f9987d-9b2e-495e-97c9-e375d3f97e63 · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

RADIO1D: Elastic Representations for Condensed Vision Modeling DataComp: In search of the next generation of multimodal datasets

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:5ae25fe4e0fb9cd62efc366457e3ebc0b1f044f34621fc95b8d4659e30efcac1

Observation f4d1825f-7d6f-49c2-a2e0-48f4e84aeeed · outbound

This paper cites Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design.

RADIO1D: Elastic Representations for Condensed Vision Modeling Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:986e2546dcf343c1edff0732eb8418b51a41e3b80d3fc6204309e3b10fe5d0d2

Observation 29f51107-8494-4970-9be4-cb42825b3884 · outbound

This paper cites Cover and Joy A.

RADIO1D: Elastic Representations for Condensed Vision Modeling Cover and Joy A

Reference 37

Resolution
verified exact
doi, observed 2026-07-12T01:08:22.982569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:4024809f9a7ee924cd43453df912d22ea5e3bb5eca8207a5333d3c2898613376

Observation 804dbc10-94b1-439c-a10b-d3979b0f1d76 · outbound

This paper cites NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model.

RADIO1D: Elastic Representations for Condensed Vision Modeling NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:a5f8cd94ef611ba0b44cfad8f05019f7caace8f410e4a86c053d9ecf1d681613

Observation dbefd43c-09a1-43fa-af84-f742e1ce83e9 · outbound

This paper cites Towards vqa models that can read.

RADIO1D: Elastic Representations for Condensed Vision Modeling Towards vqa models that can read

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:e78a88d3cd9b9f59e5770fcd0b91a9549d63301d7d453a48a523280c1b880614

Observation d0a1ea24-90f5-47ea-a7ef-7e51dccc1a17 · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

RADIO1D: Elastic Representations for Condensed Vision Modeling DocVQA: A Dataset for VQA on Document Images

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:e057774a4fe6d059ee0468ea21a6288aa20f9b6fd23ad91d68184ef72b774765

Observation 55a4d441-fd07-464d-8ab9-848e44d22386 · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:157f84684613eeecc040b49f99cdfaed8453e83d776baa6c78034e5ec69528a6

Observation ceabeb51-c265-4956-8071-29fcef4cb21f · outbound

This paper cites OCRBench: On the hidden mys- tery of OCR in large multimodal models.Sci- ence China Information Sciences, 67(12):220102, dec 2024.

RADIO1D: Elastic Representations for Condensed Vision Modeling OCRBench: On the hidden mys- tery of OCR in large multimodal models.Sci- ence China Information Sciences, 67(12):220102, dec 2024

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:b72fd2bc934f3ac3c4c52c113f06c2002b288afda00137d19933a016fff8fd6e

Observation 0b898045-7a72-44b1-a8b8-ee2d60cf0b6a · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

RADIO1D: Elastic Representations for Condensed Vision Modeling OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:4a57bdceb7a2d7d0192b5dad45788adcce9a042d7ec773a28e05db335077a657

Observation 5b8448dd-d5af-458c-b582-ca231587385e · outbound

This paper cites A diagram is worth a dozen images.

RADIO1D: Elastic Representations for Condensed Vision Modeling A diagram is worth a dozen images

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:2e21e3396f2983e8cdbc591d233eb3821b61d0d3d9b4e03d5eca76696a14590a

Observation 461920ee-dd50-4318-ab0d-74fcaaaea9c2 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

RADIO1D: Elastic Representations for Condensed Vision Modeling ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:25d084bac11d5c330262dda071dcc0db8890af7f5bd5e63a9fabeaac49dc29a1

Observation f69973c4-41e2-4e3e-8c2f-91e7df19fd66 · outbound

This paper cites findings-acl.177.

RADIO1D: Elastic Representations for Condensed Vision Modeling findings-acl.177

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:22c9ea883b6655ec1a211d254db89e93cb83c24cfbf3e91c96a9d36829a13633

Observation f364fbe3-0860-454d-82c8-afdf54fae839 · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

RADIO1D: Elastic Representations for Condensed Vision Modeling Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:969099d6eaa2ce5a9767ae1a59dcaa927150c1f9ab132a29d31b71ece8efc26c

Observation 4d5edd92-9330-46c6-83b4-312464a78755 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

RADIO1D: Elastic Representations for Condensed Vision Modeling Seed-bench: Benchmarking multimodal large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:8492516d093e10d341e958edd99434283e23982e71efedbc223e4d39bd4e751c

Observation 9bee5ee8-9c47-4b2c-83dc-8bdf3f6797f9 · outbound

This paper cites Longvideobench: A benchmark for long- context interleaved video-language understand- ing, 2024.

RADIO1D: Elastic Representations for Condensed Vision Modeling Longvideobench: A benchmark for long- context interleaved video-language understand- ing, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:272f2cba6087db347a1cfa0430752b78cc6082e1c7476fc91b93ad813a296f19

Observation 8e10099c-2ee2-4bde-9bd8-18fad7fbfe66 · outbound

This paper cites Token merging: Your ViT but faster.

RADIO1D: Elastic Representations for Condensed Vision Modeling Token merging: Your ViT but faster

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:97dffd12105fa1c59eca363f42643fd5d5617386517ad51c3756ad9c251925fc

Observation af6ec779-02cc-4cd9-9823-eff642eca4cb · outbound

This paper cites Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs.

RADIO1D: Elastic Representations for Condensed Vision Modeling Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:8b045ea078bb86345993681599af507852a0fbcf04af1b363518558dda0db468

Observation f6521bc4-61be-4709-bdfd-a16a87087341 · outbound

This paper cites Generalized intersection over union.

RADIO1D: Elastic Representations for Condensed Vision Modeling Generalized intersection over union

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:e2b196dae4536609fd716c36208253d7be845b0e88f9fe066e1575c95dcedaec

Observation 75a16e6d-d832-405c-ba20-5780c54e871b · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:75d2322120d28910f6b05b5eac3ba4ec37468929facc686b9b817e7de5cf89a2

Observation 54a38a24-d414-4cbe-95ab-73f04a3552b7 · outbound

This paper cites nuScenes: A multimodal dataset for autonomous driving.

RADIO1D: Elastic Representations for Condensed Vision Modeling nuScenes: A multimodal dataset for autonomous driving

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:99541ab6d0f527e7a7c483417360343eb5a8ad01f3bd7c890852b87fd32e0547

Observation 9bd40bc9-03a8-4e3d-a9f8-a92971d1603a · outbound

This paper cites Vision transform- ers need registers.

RADIO1D: Elastic Representations for Condensed Vision Modeling Vision transform- ers need registers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:e9548855e0f1ff5ffa220a695f4dfdb746a02098f0b09fec5548eb659f9b5be6

Observation aeb46822-da92-4a36-b465-b232079d691e · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:930ac1e5cfd5ca9a4169d2f61a6235335368939e7877eb9b66db172d57d33c89

Observation 55116568-fc65-4b8a-a22b-30387427ee3b · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

RADIO1D: Elastic Representations for Condensed Vision Modeling An image is worth 32 tokens for reconstruction and generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:9caf9331aa924e1fd7e08d9d0a5afe475104a84c951277e542142f32d7eac746

Observation af783732-6037-44ee-ba36-c46a55a80944 · outbound

This paper cites Net2Net: Accelerating Learning via Knowledge Transfer.

RADIO1D: Elastic Representations for Condensed Vision Modeling Net2Net: Accelerating Learning via Knowledge Transfer

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:8d8fd3cf4fdc80686c8d06a82cf537008ea7b29a725a48058155e2f32f313cd6

Observation 75e8dad6-1e16-439d-8faa-08ce07560d8d · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:3c8d358bf3060a82c1d774e8a3e1d241a767b8def2f0aa34af3934992332a946

Observation a5a0d2a4-68a1-4e6e-a731-74058dca946f · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

RADIO1D: Elastic Representations for Condensed Vision Modeling InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 60

Resolution
malformed identifier
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:b70a0f1c946d247db14bb1299b113130767a321f190e8cd7e8e2a1cb4459402b

Pith citing papers

No inbound Pith citation observations are available.