Pith. sign in

Paper Citation Record · LEDGER

Unveiling Encoder-Free Vision-Language Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2406.11832.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11832 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:10:38.185094Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b1462b5b-21a8-4213-ba81-c3c614305d62 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Unveiling Encoder-Free Vision-Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:20.790119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:6172f701837375299723508dbc0f32577350f7e3cdd483eabcc0ef3580bde330

Observation 66838fd0-dab7-475e-b7dd-0ad5dc33381b · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:56:07.501418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:74fdf0b792dc4ebd3398578129b3879d183c65bbe7b57e508089531c4c3c39f0

Observation 81d5e71b-6d5f-4eef-9a8a-2c6473d63b89 · inbound

Interleaved-Modal Chain-of-Thought cites this paper.

Interleaved-Modal Chain-of-Thought Unveiling Encoder-Free Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:38.185094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:38.185094Z digest=sha256:ff0a7e9af777d0dd08807e7b50721d0671034a3a8391082eef4735129486e13b

Observation 16ee3840-c0b6-4cfc-b6cd-a81b11080e4d · inbound

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding cites this paper.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.661560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.661560Z digest=sha256:780ee1f4d7a130b08d128c8a555aa9ce81bea5e9f18573ac633f06705e20d5a2

Observation de7930ef-f9ba-4cdf-8795-e67183fa8da3 · inbound

Optimizing Vision-Language Interactions Through Decoder-Only Models cites this paper.

Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.405342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.405342Z digest=sha256:0783450406706fc62773cefbe1bbc8d47123673cbfc1a84482b4233d12edae78

Observation d9efcc2f-8978-4d23-97a8-97fad21a7f40 · inbound

Optimizing Vision-Language Interactions Through Decoder-Only Models cites this paper.

Optimizing Vision-Language Interactions Through Decoder-Only Models Unveiling Encoder-Free Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.412311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.412311Z digest=sha256:4d5d21e283b44e06dede65cc745bc8b106711ba97c8aa195850514f8f8c9635c

Observation b77644af-4a32-4f32-9ec4-3fc6f4dbf302 · inbound

LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering cites this paper.

LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering Unveiling Encoder-Free Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:13:24.448179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:13:24.448179Z digest=sha256:ab1761ee038d787467808b1b08b8c42209dab27e2f2fa4c5788b2be2cbe15a83

Observation f2eaaa8b-b60f-4334-adcf-1a09a07d5bea · inbound

FastVLM: Efficient Vision Encoding for Vision Language Models cites this paper.

FastVLM: Efficient Vision Encoding for Vision Language Models Unveiling Encoder-Free Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:19:23.160965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:19:23.160965Z digest=sha256:8e17b48d856f024e792c3eca251a59fe76eeeca11b36fc5ca9825631af42fd76

Observation 7d57e557-7d78-405d-89fe-b6177df03839 · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization Unveiling Encoder-Free Vision-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.775119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:7738bfb6021dc0cec44d9438f5a87d203138741d73dd2d634f9f48932918d68c

Observation 7e1bbfec-005f-4607-b11a-17934885a95c · inbound

ECG-Byte: A Tokenizer for End-to-End Generative Electrocardiogram Language Modeling cites this paper.

ECG-Byte: A Tokenizer for End-to-End Generative Electrocardiogram Language Modeling Unveiling Encoder-Free Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:20:25.637117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:20:25.637117Z digest=sha256:fbbb392ae7625fc53131ebe28306248a1c1a56f4f596cda9f3013c69c9fa5eb8

Observation 09609b9d-e3e4-4981-bc3b-3ab8bd6810fa · inbound

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding cites this paper.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Unveiling Encoder-Free Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.009009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.009009Z digest=sha256:6e6cf8cb6af29689b4c44c8eda4fc3385c3a278cecad31ca3ed42bee6d9cfe87

Observation 45336d13-7523-4ecb-a115-66b85f1dd4c4 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unveiling Encoder-Free Vision-Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.509004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.509004Z digest=sha256:1e75d0ae0e0287f3236249087915fdbe740ff9da6140d6107720b76d5de0cbdb

Observation 353cc42d-074c-4cc4-a546-351ab013e0d6 · inbound

KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification cites this paper.

KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification Unveiling Encoder-Free Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:30.501449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:30.501449Z digest=sha256:90ca23d34dae7a2a181e9293e63f45be5f092d4469a26e831e84ead40a9204d2

Observation ae6bffc1-ee39-4e33-abcb-314b493cca86 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.616853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.616853Z digest=sha256:14416b649922e1bad81acba9e87a3fd19cc8ce9fbcc93625d3c0e6a2372753dc

Observation 5c81fc57-ac1e-4e04-81bb-c0024add02cb · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Unveiling Encoder-Free Vision-Language Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:05.462877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:246f921048e370239821b1c0160dd43620d78d58bdfef291acba7856fc3cd633

Observation 7778795f-6ccf-4f52-8b8e-cc97d9ed29f7 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Unveiling Encoder-Free Vision-Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:22:37.309970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:5f81515d18dad7c607411d72aed033e9ee27b38cbbdc958fdc2bf9c33ca44b08

Observation 9713abf4-f732-4f4c-9f1b-cdedf42b3fbf · inbound

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation cites this paper.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unveiling Encoder-Free Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:15.700721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:15.700721Z digest=sha256:ca563449879bbde21d598aa55407df2f6b29532fd3399c772087fc948d1af624

Observation 6a10af03-a77f-46dc-9977-4ebbbfb433b1 · inbound

Dense360: Dense Understanding from Omnidirectional Panoramas cites this paper.

Dense360: Dense Understanding from Omnidirectional Panoramas Unveiling Encoder-Free Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:47.368477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:47.368477Z digest=sha256:bead5cf75467e5d032d9502ee35143cefaeff8c21da6a894438a140d3692e87f

Observation 077632fb-0668-46c5-a881-19e25638e747 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Unveiling Encoder-Free Vision-Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:51:16.044495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:4d17a7200dc5379ee77c8b80f10a4939ce2f64eba925bd0f3156bbcaace6adbc

Observation e7305b86-daf9-40fe-b4aa-958b5de44d65 · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs Unveiling Encoder-Free Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:44.483409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:44.483409Z digest=sha256:603abc6ebe0e874bbfc6f7bc473fc7e2b7369557b385dd7cc66a3ec543a2f1ef

Observation 9946d252-2a16-4bd1-a9e0-50e93c74d681 · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation Unveiling Encoder-Free Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.468458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.468458Z digest=sha256:c6d6129e3b7c475082e7e6086b6cbd292f3ef51623fba6cae8537f0a7492bb54

Observation 31e4b70e-3842-4761-9ffd-58b30ea63fba · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Unveiling Encoder-Free Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:56.218152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:56.218152Z digest=sha256:986bb52f0ed48baee061a6f5a4519bde3cd7fe21cc430592e8bc9b93c0f88d48

Observation 3412c648-d966-4653-9424-1378125f9dd4 · inbound

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation cites this paper.

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:15:58.922895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:45:36.306400Z digest=sha256:a3e7715a7c4b9eb5a74a9a148ab1305c594d6fb2062544b4d8c00296c4d775cf

Observation 6820083f-b286-49e2-9af7-96923cc49188 · inbound

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning cites this paper.

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning Unveiling Encoder-Free Vision-Language Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:13.486294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T17:19:59.247074Z digest=sha256:c357d266aa72aa226ca1c3a86a5ca5ae81b8a769749da36679faccf455f654f3

Observation 3ff67812-984b-4a7f-8fc2-d12182f1d980 · inbound

From Pixels to Words -- Towards Native One-Vision Models at Scale cites this paper.

From Pixels to Words -- Towards Native One-Vision Models at Scale Unveiling Encoder-Free Vision-Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:28.011099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T13:29:53.006064Z digest=sha256:a61c6774398040b85c61f347283b93ff6e7a51afba2d73065e2624b22d614989

Observation ba1b7639-e74c-4049-9c4a-0414db2ff596 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Unveiling Encoder-Free Vision-Language Models

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:32:46.758323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:d57dd2718f4a454a4f2ef52f3f1d7f7387c75c4cb371d4f3743a42c4d6908097

Observation cda596e6-7081-46cb-ba46-9774f1433403 · inbound

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning cites this paper.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Unveiling Encoder-Free Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:0bc213ce1b060f35a65820de6e3ccc04267e60eb83e4412b1e10c65962c3fd9b