Pith. sign in

Paper Citation Record · LEDGER

DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2406.04334.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.04334 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:03:32.186910Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4d65746e-6007-420e-885f-691aa1ddb11d · inbound

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models cites this paper.

TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:32.186910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:32.186910Z digest=sha256:a72d533e0503b04f7ad6d0584e591218144b137443fc189c8d3a0489b8b780e3

Observation b123146d-63d4-4d92-8c87-a48a13b683dc · inbound

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs cites this paper.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.626572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.626572Z digest=sha256:50c2f4b8895bdd1a87e1e76d1e5312e61733ac4d30ec3457bd692e8056a4ea38

Observation 939c4584-e4f0-4317-8a4b-07b74f717717 · inbound

A Study on Context Length and Efficient Transformers for Biomedical Image Analysis cites this paper.

A Study on Context Length and Efficient Transformers for Biomedical Image Analysis DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:50:08.461274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:50:08.461274Z digest=sha256:7dd21b43b337be049721f4d99f13282fc4b808cc41c0ea3067d8a2c2e8e32675

Observation c6747194-ffa5-4fad-a468-8e298098a2ef · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.529236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.529236Z digest=sha256:6426a156493f3ed52405e597e28534e3096e253ce58754f29ed2d17f509596ca

Observation 8a8f9037-a292-480c-a907-1bb878ebd52c · inbound

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models cites this paper.

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:50:14.761238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T14:48:21.787919Z digest=sha256:fb16bf24ae2d52c5ee096a83797d55e8b63c5fba4dcda7e174d25969a458a7df

Observation 9461c75d-775d-4938-a173-9a6c9be1ee97 · inbound

LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents cites this paper.

LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:56.184492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:56.184492Z digest=sha256:42b3f7021c600da0e749d1189ed68d6728f25c35fe4d26fe9a31e148e25c7d92