Pith. sign in

Paper Citation Record · LEDGER

Unveiling Encoder-Free Vision-Language Models

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2406.11832.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11832 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:10:38.185094Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b1462b5b-21a8-4213-ba81-c3c614305d62 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Unveiling Encoder-Free Vision-Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:20.790119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:0139ce8c292431588fe63940b8e439c882175138f9fcfc9e0fa4cc0d9b027a91

Observation 66838fd0-dab7-475e-b7dd-0ad5dc33381b · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:56:07.501418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:5084c2f578597fa109cd13b8a0c75bc977339445e763bf5c337d317708425393

Observation 81d5e71b-6d5f-4eef-9a8a-2c6473d63b89 · inbound

Interleaved-Modal Chain-of-Thought cites this paper.

Interleaved-Modal Chain-of-Thought Unveiling Encoder-Free Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:38.185094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:38.185094Z digest=sha256:1d1fded969945ae7ec9b9b4206d83a9fb149360c6c3db5289d35df3cdcd43ca2

Observation 16ee3840-c0b6-4cfc-b6cd-a81b11080e4d · inbound

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding cites this paper.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.661560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.661560Z digest=sha256:bce1d53acd0effa3b9d43ad40b0820cddcfacd743437469cf691154b8a2189fc

Observation de7930ef-f9ba-4cdf-8795-e67183fa8da3 · inbound

Optimizing Vision-Language Interactions Through Decoder-Only Models cites this paper.

Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.405342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.405342Z digest=sha256:6ef4a1cb003eb305e62a01fc70f7edb0e5c3c43bccbba4c9719a7252559b6355

Observation d9efcc2f-8978-4d23-97a8-97fad21a7f40 · inbound

Optimizing Vision-Language Interactions Through Decoder-Only Models cites this paper.

Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.412311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.412311Z digest=sha256:5cf0b098c88a21b0ec35dd4e43e59e07d86104e06aa55e0acc076bce8d186ed6

Observation b77644af-4a32-4f32-9ec4-3fc6f4dbf302 · inbound

LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering cites this paper.

LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering Unveiling Encoder-Free Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:13:24.448179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:13:24.448179Z digest=sha256:8edb5735ea6e30d5e8516b1b04384f8208a1ff13f520aeeb3a91a667dffe2b42

Observation f2eaaa8b-b60f-4334-adcf-1a09a07d5bea · inbound

FastVLM: Efficient Vision Encoding for Vision Language Models cites this paper.

FastVLM: Efficient Vision Encoding for Vision Language Models Unveiling Encoder-Free Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:19:23.160965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:19:23.160965Z digest=sha256:2d5ad0baed214e5f9fd5703231a9d9b51a24eec8a9629b63afed591d26fe7f8b

Observation 7d57e557-7d78-405d-89fe-b6177df03839 · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization Unveiling Encoder-Free Vision-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.775119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:71ed6f0a6834a911d06d31cf96692d9d993590f7c0d2e318b622c24d1d34f238

Observation 7e1bbfec-005f-4607-b11a-17934885a95c · inbound

ECG-Byte: A Tokenizer for End-to-End Generative Electrocardiogram Language Modeling cites this paper.

ECG-Byte: A Tokenizer for End-to-End Generative Electrocardiogram Language Modeling Unveiling Encoder-Free Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:20:25.637117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:20:25.637117Z digest=sha256:a76e3e6fd47928493c48246678cfb414e2549925bd3f648b6ee3d2790860fefa

Observation 09609b9d-e3e4-4981-bc3b-3ab8bd6810fa · inbound

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding cites this paper.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Unveiling Encoder-Free Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.009009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.009009Z digest=sha256:7345b833b2c56e0d3ecd0d66d2f096335c283ad4dd0758f050de46fdd95ea1a4

Observation 45336d13-7523-4ecb-a115-66b85f1dd4c4 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unveiling Encoder-Free Vision-Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.509004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.509004Z digest=sha256:b204c5228d1af7acae7b63212f16960737c358f3134a8640ce296e18d9aa3995

Observation 353cc42d-074c-4cc4-a546-351ab013e0d6 · inbound

KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification cites this paper.

KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification Unveiling Encoder-Free Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:30.501449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:30.501449Z digest=sha256:595b07742453dd5a036261614059f81def9949237d1b145282896a891085465f

Observation ae6bffc1-ee39-4e33-abcb-314b493cca86 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.616853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.616853Z digest=sha256:a0da34aaa3ce911ac7e0f37130e8ad7f48933dbbad7f78229a8c51583894b9b3

Observation 5c81fc57-ac1e-4e04-81bb-c0024add02cb · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Unveiling Encoder-Free Vision-Language Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:05.462877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:542082a0b1e1441500b95da8186fc811d910271b1f36d08e8df856a9f2881879

Observation 7778795f-6ccf-4f52-8b8e-cc97d9ed29f7 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Unveiling Encoder-Free Vision-Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:22:37.309970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:e183b65936d7228063a561cbffa144a4097babbd2e720ea3fca0c3c247467b31

Observation 9713abf4-f732-4f4c-9f1b-cdedf42b3fbf · inbound

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation cites this paper.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unveiling Encoder-Free Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:15.700721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:15.700721Z digest=sha256:41f27db4c9ca1ffac34dad3fc6c0ac1dcbcd1a471b0efc5227f1a4a37aa457ca

Observation 6a10af03-a77f-46dc-9977-4ebbbfb433b1 · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas Unveiling Encoder-Free Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:47.368477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:47.368477Z digest=sha256:2dc9d49a5694549b8d597f04a35bd76845bafae6df5870b7e73fe99065a91556

Observation 077632fb-0668-46c5-a881-19e25638e747 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Unveiling Encoder-Free Vision-Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:51:16.044495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:1addfbea73e0cf0d40873720c12300647e79d186ad96463406f843f9f6993ed1

Observation e7305b86-daf9-40fe-b4aa-958b5de44d65 · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unveiling Encoder-Free Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.483409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.483409Z digest=sha256:b528a074f20edeea2f8cc42caf09826eb68e4e1f71d34f4d74ec5ef6e30cda8e

Observation 9946d252-2a16-4bd1-a9e0-50e93c74d681 · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation Unveiling Encoder-Free Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.468458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.468458Z digest=sha256:ad638e0a0bdec37dda721c8ba42a434634c034d3a9b84deaf3f97b1edc875b93

Observation 31e4b70e-3842-4761-9ffd-58b30ea63fba · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Unveiling Encoder-Free Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.218152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.218152Z digest=sha256:daf4fce1c828890ed30f7a67ab66241288c870f5314f367d9c1a97c69bf3cad7

Observation 3412c648-d966-4653-9424-1378125f9dd4 · inbound

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation cites this paper.

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:15:58.922895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:45:36.306400Z digest=sha256:2ccb72c5f928f049547a8716aaa81bf85a7beab34be896270a35c658564f74f7

Observation 6820083f-b286-49e2-9af7-96923cc49188 · inbound

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning cites this paper.

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning Unveiling Encoder-Free Vision-Language Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:13.486294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T17:19:59.247074Z digest=sha256:f77e20a2c6c9712b8e6f88a98cf5db8ba11e09298b1ceebff804e02fa81f01b9

Observation 3ff67812-984b-4a7f-8fc2-d12182f1d980 · inbound

From Pixels to Words -- Towards Native One-Vision Models at Scale cites this paper.

From Pixels to Words -- Towards Native One-Vision Models at Scale Unveiling Encoder-Free Vision-Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:28.011099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T13:29:53.006064Z digest=sha256:e91a1e9e27b69a9176ae94193313fdca2f51917adb98a5a15d464b5692d7986f

Observation ba1b7639-e74c-4049-9c4a-0414db2ff596 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Unveiling Encoder-Free Vision-Language Models

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:32:46.758323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:b233858a4dab2d30f3956929158bf39ae216154cf2be725ae326fd8ece421bdf

Observation cda596e6-7081-46cb-ba46-9774f1433403 · inbound

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning cites this paper.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Unveiling Encoder-Free Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:081353c7cc8d7404086f415ab4f4b327d49a1a8bca2bf2ca7912f83603510c3c