Pith. sign in

Paper Citation Record · LEDGER

From Text to Pixel: Advancing Long-Context Understanding in MLLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2405.14213.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.14213 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:47:38.551116Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T11:44:37.993936Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e0de8095-6b5e-4156-a81e-d353d6f4f0ef · inbound

SFNet: Fusion of Spatial and Frequency-Domain Features for Remote Sensing Image Forgery Detection cites this paper.

SFNet: Fusion of Spatial and Frequency-Domain Features for Remote Sensing Image Forgery Detection From Text to Pixel: Advancing Long-Context Understanding in MLLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T22:50:06.414052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:50:06.414052Z digest=sha256:f5251ca70ce607c36f8294eb826f5e0be9a6c5d893cfbb393d576629740b6d22

Observation d6017804-7895-448c-a87a-29dde3b08fba · inbound

Structured Attention Matters to Multimodal LLMs in Document Understanding cites this paper.

Structured Attention Matters to Multimodal LLMs in Document Understanding From Text to Pixel: Advancing Long-Context Understanding in MLLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.551116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.551116Z digest=sha256:45e157150f0715beb2043b0b4f7aac4406a1787611d99d5d2bd6b779495ce0a7

Observation 4599e2b0-48c9-43fc-a656-216372864089 · inbound

Visual Text Compression as Measure Transport cites this paper.

Visual Text Compression as Measure Transport From Text to Pixel: Advancing Long-Context Understanding in MLLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:57.787128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:17:48.814483Z digest=sha256:494c6ec578ab182c9a9fce151830556c2731f2ee008ba5bf9fa74aff49dc490b

Observation e146111f-92ea-46cb-b125-8ac708e82180 · inbound

Memory Shot for Long-Term Dialogue cites this paper.

Memory Shot for Long-Term Dialogue From Text to Pixel: Advancing Long-Context Understanding in MLLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:44:37.997937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T11:41:00.909055Z digest=sha256:a69a80c5a1391aea40ce10e4da020d26a6acc0d7fdec349bb48bb29305322c24

Observation db775473-8b0a-4b59-9fed-53ac5b5815c5 · inbound

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation cites this paper.

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation From Text to Pixel: Advancing Long-Context Understanding in MLLMs

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:34:37.688285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T11:31:56.598221Z digest=sha256:ac300d2be82fad780e1c74b81e8833548f07e8cc33707639670320ed1170324a