Pith. sign in

Paper Citation Record · LEDGER

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

As of 19 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.24424.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.24424 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T15:16:36.659461Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 265850aa-de40-408b-a3a6-99916840bc9e · outbound

This paper cites Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.452784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.452784Z digest=sha256:5d5aa471438906a97b89a823d9c8c561122735f45180f64bf7471a3fc769206f

Observation a40eecd5-d6ad-41c4-a083-8102051bed19 · outbound

This paper cites Yingen Liu, Fan Wu, Ruihui Li, Zhuo Tang, and Kenli Li.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Yingen Liu, Fan Wu, Ruihui Li, Zhuo Tang, and Kenli Li

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.666344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.666344Z digest=sha256:e18681341ac127296e85a57fab033f741cf2f1842200907cac019c055b9dc62e

Observation 8d1b3508-6f4c-4e2b-8fb4-820f22c94017 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning MMBench: Is Your Multi-modal Model an All-around Player?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.713332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.713332Z digest=sha256:0014c02b0a3ddf81523062898c347a82ca424c52434e16a086fad078a086fbaa

Observation 5c263949-b6bc-42f6-ba5f-88e8d9d2d287 · outbound

This paper cites LLM-CoT enhanced graph neural recommendation with harmonized group policy optimization.arXiv preprint arXiv:2505.12396,.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLM-CoT enhanced graph neural recommendation with harmonized group policy optimization.arXiv preprint arXiv:2505.12396,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.776370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.776370Z digest=sha256:825e2977381027eeedb2f9913c3c886f620262f795b619cb5e5fe5e5370fb240

Observation 076e2494-6d70-4159-b50e-7c46c16574ee · outbound

This paper cites Synthetic Lung X-ray Generation through Cross-Attention and Affinity Transformation.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Synthetic Lung X-ray Generation through Cross-Attention and Affinity Transformation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.827292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.827292Z digest=sha256:48932558c22b123fa8ecc8f0e17b5b58aa13dac791696857111e462bb6a04f86

Observation 2f5f2ab3-e589-4d00-a86b-9f6df844c650 · outbound

This paper cites Boosting General Trimap-free Matting in the Real-World Image.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Boosting General Trimap-free Matting in the Real-World Image

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.885573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.885573Z digest=sha256:d58000c3e39b63306ba1710acfca7fe89147eed5b82aa7b5f22dbf9e826aecf6

Observation 2860f3ab-5cea-41f0-840f-f2d73964f7e7 · outbound

This paper cites Edge-guided and Class-balanced Active Learning for Semantic Segmentation of Aerial Images.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Edge-guided and Class-balanced Active Learning for Semantic Segmentation of Aerial Images

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.961854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.961854Z digest=sha256:8d6e78fe1bfc82bbe124256cb820dc67624cab0432c52742cc05c906defb82d1

Observation 5b72cc80-e18a-441d-bb84-bd50cac27d7f · outbound

This paper cites LLaV A-PruMerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLaV A-PruMerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.036325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.036325Z digest=sha256:ea3fa69ed11ee91dc00eafc87d2bd8cfe92952ff437d2b1ec03acc3c9db4b739

Observation 615aaa49-0563-4914-a4ef-9059c1386717 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.166786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.166786Z digest=sha256:9073d6c3df67b70f1b64fced86852ed6f1732a8d4ef25a2b3671e0c4e83b69d6

Observation aff33c9b-f53e-4232-b2ac-e8a09bce3b99 · outbound

This paper cites GDGS: 3D Gaussian Splatting Via Geometry-Guided Initialization And Dynamic Density Control.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning GDGS: 3D Gaussian Splatting Via Geometry-Guided Initialization And Dynamic Density Control

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.231359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.231359Z digest=sha256:4404649b82292336a7e89061fabbbed8ab5cf085ea796bc8a91dfa98fc43d203

Observation 256a2bc7-e757-49fd-98e6-8408da8096ee · outbound

This paper cites RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.284072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.284072Z digest=sha256:0f16a6fad48098b34efffa94bc5f2c32d16fd36d7e95275e366e1ee69a339729

Observation c955fe02-0de6-4385-83ac-b075f505bd63 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.341288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.341288Z digest=sha256:baf40c60316c1c4da2d8cfd4b2991ad8d112c3a3c3e7ae92e2df63daf5daf5b6

Observation 7df803d8-1062-441e-9490-50f8a674284e · outbound

This paper cites KV-Efficient VLA: A method to speed up vision language models with RNN-gated chunked KV cache.arXiv preprint arXiv:2509.21354,.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning KV-Efficient VLA: A method to speed up vision language models with RNN-gated chunked KV cache.arXiv preprint arXiv:2509.21354,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.403095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.403095Z digest=sha256:c4b5906c88015988064acea882cdfab3d2fb5f523a804ed099fe789fe768c889

Observation aa5ec07b-db4a-4d1c-b66a-904a0e21d061 · outbound

This paper cites A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.496400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.496400Z digest=sha256:670031474571a58a3fc585d73630ae20ca9d019fd586c3d1336db4c4fb0bc972

Observation b13f99e4-8ea2-4f11-961c-0038e2117e76 · outbound

This paper cites Asymmetric Mamba–CNN collaborative architecture for large-size remote sensing image semantic segmentation.IEEE Transactions on Geoscience and Remote Sensing, 63:2002419, 2025a.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Asymmetric Mamba–CNN collaborative architecture for large-size remote sensing image semantic segmentation.IEEE Transactions on Geoscience and Remote Sensing, 63:2002419, 2025a

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.589098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.589098Z digest=sha256:3818aef395dfd84485f6d5e2ec7592a4ca67267bc8ce1a9887bd5934cf67dd12

Observation 8485d1c8-2336-4c7e-ade7-65d6837eeaea · outbound

This paper cites DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.659461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.659461Z digest=sha256:91bda49d24f393f38823df2c02558d6ebc764e26b43e87c4ebf846e30dcd84da

Observation a33bd187-fc65-4803-86b4-289a43d4f062 · outbound

This paper cites NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.591649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.591649Z digest=sha256:343640bdd9766984d8771ad741e73faf3518fc02a5d24887a370a590d398fa97

Observation 1504ed36-23b4-4c05-9f33-67006ed22c3b · outbound

This paper cites GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.098226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.098226Z digest=sha256:81e3fa6e56eadb46670bf1dd578942ee78567316cbe4ce7c0f71dc6cd9f25d07

Observation 0c31cf5c-e39d-475f-a4b4-9f95d8461d36 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.257881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.257881Z digest=sha256:96cb330f1985b6ecca2b0ec9045ead2cade62790cb6fa3e9ba7060ee48e00492

Observation f86ea83f-0f57-4ac1-88c1-8fa8535132cb · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.069570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.069570Z digest=sha256:3e51ef7f2f680592a3650ed234c8bece5eda7030e53ff497f1a9178dc1e661a7

Observation 0fca96e3-fc26-49c9-a068-05d487e41777 · outbound

This paper cites The Binary Quantized Neural Network for Dense Prediction via Specially Designed Upsampling and Attention.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning The Binary Quantized Neural Network for Dense Prediction via Specially Designed Upsampling and Attention

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.157744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.157744Z digest=sha256:6c2a71333f3cca33ddbd3b3e2186cd85f7742fe07f70c08b362cb3f635943472

Observation 37815ebb-7506-4f76-af7f-0b0ad3507636 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.320995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.320995Z digest=sha256:0cc6131bb6ab7fae80daae52fe70f0ff41c94db6a215aa94fef266de5b17d703

Observation 8f953d7f-c32b-4170-910c-9dc353963321 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.532712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.532712Z digest=sha256:d0ddd1f209eb20a6e691272bbf080b3942eaa071526b214bdd04367fdd6a9a15

Pith citing papers

No inbound Pith citation observations are available.