Pith. sign in

Paper Citation Record · LEDGER

ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2405.15738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.15738 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:35:52.235865Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:59:57.250489Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 98d09c32-aef5-431c-8245-a3c823a180fd · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.729673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:1ea6359136e5144b8268c2bce6f644da874b91b124857ee01550788afe89699e

Observation 432abb0b-d338-4279-8de6-313248a04d2b · inbound

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion cites this paper.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.235865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.235865Z digest=sha256:fb457e1b2bdce81d4f3db41eac55c87e04582ec45c6550f0aabbd47010d659d6

Observation 0a835e97-59ee-4fa8-bca9-a090ddfd0bd2 · inbound

CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception cites this paper.

CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:21.306292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T04:46:52.071071Z digest=sha256:c74c8f3411494d9fb94f8aa66caa9fe352a28fba457607f7d5767f2c76d52f22

Observation df9e0a50-bc5e-4a15-a92b-adcae9858e33 · inbound

ActiveScope: Actively Seeking and Correcting Perception for MLLMs cites this paper.

ActiveScope: Actively Seeking and Correcting Perception for MLLMs ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:59:57.252044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:06:40.481169Z digest=sha256:6dc8ebf827185f63e2e22536e0382bc954c82f47d2d56f293853c3c5db142660

Observation 33c0eaca-233f-4245-a0ec-acf2299cc9d2 · inbound

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration cites this paper.

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:07:02.344333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T14:00:31.479090Z digest=sha256:47d32d68462959bfa9122a9309030ee6e813ea3de9ac0da236fe4b62f64ab1b3

Observation 6da9a312-d3e2-4246-aa58-688f1c359446 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 145

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:dd1eba9ebfc4344ab45e21d8f843fa0b7280075a4ae03a228204317a953a24a0