Pith. sign in

Paper Citation Record · LEDGER

EVLM: An Efficient Vision-Language Model for Visual Understanding

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2407.14177.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.14177 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T21:39:21.630533Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T06:20:36.339231Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 95ac0b8e-6af4-4f6c-862b-3910245f4064 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 202

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.342787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:488194887d27652d2b3f88226b899988f9a24072b79c3380bbee511bf7ad64ac

Observation 0c4271e4-c97b-4810-9b10-48639d805bf4 · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.572829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:ba7a083f205994780e7b3c76089092cf9df678b1aeab5f4fa9dc58273188928b

Observation bd44a04c-86c7-44f6-a822-f14fb55347d5 · inbound

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs cites this paper.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.630533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.630533Z digest=sha256:ac3f40bae85d6a608f2659f949fc05fc9b96e632e90216e673276fbe91e1c981

Observation aa14a1d1-29be-4aeb-89cf-76468b19416b · inbound

CoMemo: LVLMs Need Image Context with Image Memory cites this paper.

CoMemo: LVLMs Need Image Context with Image Memory EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.175851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.175851Z digest=sha256:21fd82b769896b8e6493056d7126c8d84b93549147d4b97e357eaa54f4f91f03

Observation ff8c83f7-2708-46ab-b3b5-8589b3cf3601 · inbound

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models cites this paper.

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:27.089083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:27.089083Z digest=sha256:acc7fd100397e61cc3caaa53f04028022827c2e06a8873f73af40a4680d8f8cb

Observation 865f3073-7922-4280-8fb8-51adb5bb41d0 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.808793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.808793Z digest=sha256:9cde5dbe19caac17b6623b8293a81923a3d1a89f5ef0a3d95ba845f3b9a7c4c7

Observation 56b4e346-6a64-4462-8f01-3ed46df19c48 · inbound

From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion cites this paper.

From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T10:18:31.583291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:18:31.583291Z digest=sha256:62c1093b0c5db0d7c043007fcf608e6da3e840edca0d3c317c57ed773371f74e