Pith. sign in

Paper Citation Record · LEDGER

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

As of 7 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.24424.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.24424 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T15:16:36.659461Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 265850aa-de40-408b-a3a6-99916840bc9e · outbound

This paper cites Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.452784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.452784Z digest=sha256:39a79df20dde00541ceca2e1988fedaee8caf4b3b57cc5f337c800d9dbee4e1f

Observation a40eecd5-d6ad-41c4-a083-8102051bed19 · outbound

This paper cites Yingen Liu, Fan Wu, Ruihui Li, Zhuo Tang, and Kenli Li.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Yingen Liu, Fan Wu, Ruihui Li, Zhuo Tang, and Kenli Li

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.666344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.666344Z digest=sha256:e003b7b6557200e3a02bf57967bff27879e083643ba9bbc42f21eb6fa00c5aa3

Observation 8d1b3508-6f4c-4e2b-8fb4-820f22c94017 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning MMBench: Is Your Multi-modal Model an All-around Player?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.713332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.713332Z digest=sha256:b220d01a7a86425dbe511ff23b74c8e502ee1341acfe10d80fb051ca4c459076

Observation 5c263949-b6bc-42f6-ba5f-88e8d9d2d287 · outbound

This paper cites LLM-CoT enhanced graph neural recommendation with harmonized group policy optimization.arXiv preprint arXiv:2505.12396,.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLM-CoT enhanced graph neural recommendation with harmonized group policy optimization.arXiv preprint arXiv:2505.12396,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.776370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.776370Z digest=sha256:0c1e7a2fb9768381dd183e1ee2b059e544bb841c8beb826ef9f9d7998abe1674

Observation 076e2494-6d70-4159-b50e-7c46c16574ee · outbound

This paper cites Synthetic Lung X-ray Generation through Cross-Attention and Affinity Transformation.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Synthetic Lung X-ray Generation through Cross-Attention and Affinity Transformation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.827292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.827292Z digest=sha256:fdddd67da59bb90cfa6fed1f4da7d24507b29826ec95f94e3e56487ffc872d92

Observation 2f5f2ab3-e589-4d00-a86b-9f6df844c650 · outbound

This paper cites Boosting General Trimap-free Matting in the Real-World Image.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Boosting General Trimap-free Matting in the Real-World Image

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.885573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.885573Z digest=sha256:d5352739ca76f494a8d77890f8fb498b1738f32eb11bd5d59033e6a26fe6c79e

Observation 2860f3ab-5cea-41f0-840f-f2d73964f7e7 · outbound

This paper cites Edge-guided and Class-balanced Active Learning for Semantic Segmentation of Aerial Images.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Edge-guided and Class-balanced Active Learning for Semantic Segmentation of Aerial Images

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.961854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.961854Z digest=sha256:b53ac9d6c7701888cc0632c101d5663047caa303a0db89788119b2a91a91d7ef

Observation 5b72cc80-e18a-441d-bb84-bd50cac27d7f · outbound

This paper cites LLaV A-PruMerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLaV A-PruMerge: Adaptive token reduction for efficient large multimodal models.arXiv preprint arXiv:2403.15388,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.036325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.036325Z digest=sha256:66d887c90e61b0e8ba2d1160a418586e5b744ec914d802c880adae2a71ad57f0

Observation 615aaa49-0563-4914-a4ef-9059c1386717 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.166786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.166786Z digest=sha256:48b5882c1545768b6e53f149bb950c03b10f119df8b888f53f973dc3f39428de

Observation aff33c9b-f53e-4232-b2ac-e8a09bce3b99 · outbound

This paper cites GDGS: 3D Gaussian Splatting Via Geometry-Guided Initialization And Dynamic Density Control.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning GDGS: 3D Gaussian Splatting Via Geometry-Guided Initialization And Dynamic Density Control

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.231359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.231359Z digest=sha256:a935814c28cc3788d8a0778e86821d81fa4a69d4fcaa33e72199bd1653031685

Observation 256a2bc7-e757-49fd-98e6-8408da8096ee · outbound

This paper cites RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.284072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.284072Z digest=sha256:a65a3d431cacd91040c50495dd640cb3b3cfe0134b99c2a38933e1bbaa7dccd9

Observation c955fe02-0de6-4385-83ac-b075f505bd63 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.341288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.341288Z digest=sha256:53178ebda906117c4750aec44be5402cec59a607c95f353f2d6040e429f4665a

Observation 7df803d8-1062-441e-9490-50f8a674284e · outbound

This paper cites KV-Efficient VLA: A method to speed up vision language models with RNN-gated chunked KV cache.arXiv preprint arXiv:2509.21354,.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning KV-Efficient VLA: A method to speed up vision language models with RNN-gated chunked KV cache.arXiv preprint arXiv:2509.21354,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.403095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.403095Z digest=sha256:6267d4661d5f90803d38240d528f076f6e0eefff37524a6e6297aa2461235e5d

Observation aa5ec07b-db4a-4d1c-b66a-904a0e21d061 · outbound

This paper cites A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.496400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.496400Z digest=sha256:93a27d8cc9bb6ea5400b9651f9be6026c8026ac6a7623c876c1d48ff7beddb12

Observation b13f99e4-8ea2-4f11-961c-0038e2117e76 · outbound

This paper cites Asymmetric Mamba–CNN collaborative architecture for large-size remote sensing image semantic segmentation.IEEE Transactions on Geoscience and Remote Sensing, 63:2002419, 2025a.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Asymmetric Mamba–CNN collaborative architecture for large-size remote sensing image semantic segmentation.IEEE Transactions on Geoscience and Remote Sensing, 63:2002419, 2025a

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.589098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.589098Z digest=sha256:50b186214bc6c402635e626e7f9519f388ff63cefaf5fa5ce6b779915b63ce59

Observation 8485d1c8-2336-4c7e-ade7-65d6837eeaea · outbound

This paper cites DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.659461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.659461Z digest=sha256:97475042aa6cf954ad19ab4dc0a86cc5c5cee3a53fd7dc7d3e51ea92d2877ff4

Observation a33bd187-fc65-4803-86b4-289a43d4f062 · outbound

This paper cites NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.591649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.591649Z digest=sha256:25eba8818ad44e6cc29d090082fd955ef32070db65541079073dd7dc8c582a35

Observation 1504ed36-23b4-4c05-9f33-67006ed22c3b · outbound

This paper cites GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:36.098226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:36.098226Z digest=sha256:0e35f5a6f582ca1aa5aa633c66bfa45df35fa1c8ad29bcb90ad38ab12fa10e1b

Observation 0c31cf5c-e39d-475f-a4b4-9f95d8461d36 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.257881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.257881Z digest=sha256:70363bb6e9a1f3511be063dab98e216890cbac7e8aedc4673825b286186b3ccd

Observation f86ea83f-0f57-4ac1-88c1-8fa8535132cb · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.069570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.069570Z digest=sha256:d2500f78eb8b201643e6c0893bcad1a1f72de8796028ae95736538761149304e

Observation 0fca96e3-fc26-49c9-a068-05d487e41777 · outbound

This paper cites The Binary Quantized Neural Network for Dense Prediction via Specially Designed Upsampling and Attention.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning The Binary Quantized Neural Network for Dense Prediction via Specially Designed Upsampling and Attention

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.157744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.157744Z digest=sha256:9693b9e6784a915c86c2450bee580f9f99c2d4a26cfb9b6d71fe1d53e042916c

Observation 37815ebb-7506-4f76-af7f-0b0ad3507636 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.320995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.320995Z digest=sha256:344249ff7029d38167fb13ef6ebc109eb4e1119ec7cd304a555e8ba1cd90c633

Observation 8f953d7f-c32b-4170-910c-9dc353963321 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T15:16:35.532712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:16:35.532712Z digest=sha256:d8ce17af8c9714e1623cfe4b66ed6a69f0815b215182fbd373defb44054fc888

Pith citing papers

No inbound Pith citation observations are available.