Pith. sign in

Paper Citation Record · LEDGER

From Pixels to Words -- Towards Native One-Vision Models at Scale

As of 10 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 1 inbound Pith citation observation for arXiv:2605.28820.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.28820 v1

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T13:29:53.006064Z

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T05:44:12.758733Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T10:15:44.322072Z

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ff67812-984b-4a7f-8fc2-d12182f1d980 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

From Pixels to Words -- Towards Native One-Vision Models at Scale Unveiling Encoder-Free Vision-Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:28.011099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T13:29:53.006064Z digest=sha256:5d8d6ff9ff0bc1aaf3f2dc0361b1d528b2d9e83be67c4867bf18d426f494ccbc

Observation c332b767-a910-4eec-8996-2c315bf9a5be · outbound

This paper cites InEuro- pean Conference on Computer Vision, volume 9908, pages 235–251, Amsterdam, The Netherlands.

From Pixels to Words -- Towards Native One-Vision Models at Scale InEuro- pean Conference on Computer Vision, volume 9908, pages 235–251, Amsterdam, The Netherlands

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T13:29:53.006064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:29:53.006064Z digest=sha256:adaaf9f3473bc1dcf5d547e24b0dad7e227cbedc9e2397a59113c9360dbc0108

Observation d3dcd580-27df-4fee-9729-3b3de8ca34c9 · outbound

This paper cites The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer.

From Pixels to Words -- Towards Native One-Vision Models at Scale The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:28.016525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T13:29:53.006064Z digest=sha256:f388e04e9bd72dd2690202779eb94518acd3d838221f1b3c31c10783bc2824a4

Observation a983fc5c-43b3-4e05-9891-3768fc8af4fb · outbound

This paper cites Langbridge: Interpreting image as a combination of language embeddings.

From Pixels to Words -- Towards Native One-Vision Models at Scale Langbridge: Interpreting image as a combination of language embeddings

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:28.013987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T13:29:53.006064Z digest=sha256:b547947f374318b3399a4105a9158f7569b4e3c9fdf3e930531e2553a794434d

Observation a2a6460c-55b5-4cea-9c7c-64f554b43d5e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

From Pixels to Words -- Towards Native One-Vision Models at Scale Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:33:28.008391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T13:29:53.006064Z digest=sha256:6ad94f0ee218ca90ce89df34fb1d567f7edffd933ca8003453306d1c41443672

Observation 91f6ad7a-cd68-46cc-a942-2e48006fb6d1 · outbound

This paper cites HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding.

From Pixels to Words -- Towards Native One-Vision Models at Scale HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:28.005804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T13:29:53.006064Z digest=sha256:ddc323144e3606588246332bf2fb899475e9e14707a006b09bcc60d2c7494c9c

Pith citing papers

Observation c1487200-b3a3-4620-be98-4bbd4f3ce553 · inbound

ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs cites this paper.

ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs From Pixels to Words -- Towards Native One-Vision Models at Scale

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:15:44.323340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T05:44:12.758733Z digest=sha256:bdd9f3247017343d06753e1dbc82d1985baa691aab66ccbfb3857b8470b9f1bc