Pith. sign in

Paper Citation Record · LEDGER

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing

As of 11 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2604.13565.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.13565 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:53:13.255412Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact20
  • verified fuzzy4
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10d5bb70-009c-487f-98bd-1d4a92bbec88 · outbound

This paper cites Phi-4 Technical Report.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Phi-4 Technical Report

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:34:04.627981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:eb6077b98c0cbf62edb7db214ef13a0f6e1e30495b923b12fb0042e2bbdc8876

Observation b743bf03-f52b-4f71-af2f-94a6c4b40e93 · outbound

This paper cites Qwen2.5-VL Technical Report.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Qwen2.5-VL Technical Report

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T13:55:28.943074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:adbdbb69f5db9646fd55575996e1f22395cfb9713aa1c5b4fdf29baf5cad75c9

Observation e9eb6281-4bf4-4163-942a-6f70a5b2385d · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:29:06.753688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:e48ef4ded1b958e496a309b01129f81fa6343a7369326cd4ad5f80f66d1f1fda

Observation be95497e-1c19-49f7-984d-9c1a7e3db8a4 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:55:28.882325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:449b3b0537c7796fe2a967246ffd4dde192acca0e91602be51d660c8aac819fb

Observation 66deedc9-2dc7-486d-9b70-61735715618a · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:27:52.297815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:dfd652b82ad1590304ac6b62bd4b28f44b9cbe5a8544c708f5b191e1782d4dc7

Observation e6bf72af-f900-4464-94c8-2f652b3890ac · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:55:28.886711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:712addaaaa63c265a71c38a3e9df0f8b5e4bbe28df36bc66bf2261962b0264bd

Observation 9ada7a52-d94c-42ab-8348-381645d93efc · outbound

This paper cites Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:55:28.864017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:7c88f65b6ae8cb1f357614d02ac4b03913279b1c16874ce5f8dd73b1bf7bfed8

Observation 41dffd84-efdf-410f-a7cf-d3cda9244948 · outbound

This paper cites Fuse-rsvlm: Feature fusion vision-language model for remote sensing.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Fuse-rsvlm: Feature fusion vision-language model for remote sensing

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:55:28.859406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:9a0d87494ee503b0b87d36349d089771146e5c23c6ac9694e8a0688e1a391ac7

Observation e02b6a43-6fbf-4be8-9971-81e4cdcd2f23 · outbound

This paper cites GPT-4o System Card.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing GPT-4o System Card

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T13:55:28.868449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:b0d17a62885b3da6092e46cdc93baf395fa8ee9fad4f997745adbd765239401b

Observation 43793d89-f142-4146-9280-b7ba2545f853 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:47:14.368858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:a499068c1725f52264d373fee9ca8a1e453995c441a402eb159ef92ef065037d

Observation 803264ab-af06-49ee-bafc-52a41ccd21ae · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing LLaVA-OneVision: Easy Visual Task Transfer

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:55:28.841281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:0141a2866d406b1b3ae33e239f87ff35304944fd66dcb21bc5ee9a50b6cb7ca3

Observation 758ae707-d055-446e-a5f7-fef6647e10f9 · outbound

This paper cites Zoomearth: Active perception for ultra-high-resolution geospatial vision-language tasks.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Zoomearth: Active perception for ultra-high-resolution geospatial vision-language tasks

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:55:28.877832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:6e26f6d191bd811590292ea75ba1660c7ac53aeff892a379be0b3375d6e5f983

Observation 4694d5be-3c81-497e-80d9-5804e5b594f3 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:58:54.896552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:c5f0fb506ca04e8cf0698e783b07152fb11d671e94d375f3a59858dd5262834f

Observation d7e78842-bf47-44bf-b04b-89513cee8a1e · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:55:28.929452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:bc0415eebca978318ada147bdcb73777d1e3955071e360efbfefca843fe9f806

Observation 05ac4714-57d8-430b-8b4c-7af0fb851665 · outbound

This paper cites Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:55:28.947914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:a09f460cd97668f8d77184c438030cb2d231b31210df618d799e47b1065ab6f9

Observation 1fe73360-2323-4e6d-b566-5b72059d795d · outbound

This paper cites Pang, C., Wu, J., Li, J., Liu, Y ., Sun, J., Li, W., Weng, X., Wang, S., Feng, L., Xia, G.-S., et al.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Pang, C., Wu, J., Li, J., Liu, Y ., Sun, J., Li, W., Weng, X., Wang, S., Feng, L., Xia, G.-S., et al

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:32:51.638803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:45f8bbbb75aa15c9d1cc918d02ad376f0dc50c1c9032c4df9b81af707f948aee

Observation c67f5314-4cc2-4240-a3f8-dcd1b3d58f48 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Gemini: A Family of Highly Capable Multimodal Models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T13:55:28.895324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:7acf34ccfaa7a51de185c9c267e879ed0838961f35926e5c59287b0ce8cefbd2

Observation e514c6c9-a609-403d-9c4c-39e1164a92fe · outbound

This paper cites Geollava-8k: Scaling remote-sensing multimodal large language models to 8k resolution.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Geollava-8k: Scaling remote-sensing multimodal large language models to 8k resolution

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:55:28.919701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:379d8f23b60e60282d6863a6f3281faf51b985635ec0bfa60186a0f3984d8718

Observation 3d400f2f-eff3-4772-a8dd-bb6c7efc96e3 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:55:28.900179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:72e036239adbb49a767cdbe3f3e9faefa9b2acd52972335d5e25ab53a51df353

Observation b093fa26-d4ce-4b8a-97aa-daa1ec2ec902 · outbound

This paper cites HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:19:51.127350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:ff1e089adfa5145ac6bbf2284ceb303465eb6c886e88b74944b6d47db00a5356

Observation 1e1f7aad-727b-41f6-a4f9-25ea935e6e7f · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.698795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:601e96c4fa5c6d1a2ca9927c635fdc2fc2066375a52a2330e3b57ca4dd6f8147

Observation f7842195-dade-46fb-95ff-89cf4b2f0476 · outbound

This paper cites TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:55:28.939019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:8b0c687c4efce84ceb514d22ff0de9b0211d4c3ac7680cdde2244be53b38e5bf

Observation 9317b03e-b415-4d34-9062-8c89946060ef · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:49.136768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:6a195484b03f18118f6fddca83903b40398b1b797131d2ffa96166db7f1eb264

Observation 3ff5b6d1-3237-43c4-840f-74ce213dd73e · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:29.265418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:e5eaa4c662c87414e55fd7da547cda61b14773fd81b38f5c2f22ef104c193568

Observation f4ffe41d-642d-4a2d-98f9-0c39c02f2819 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:37:02.028216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:293524fcf8e9d74ba36b28c2a0f2c78d4be92630a40d4b370bce668fca71d85c

Observation 813d64ae-fc82-4dce-abdf-2ebc52c07bc6 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:55:28.951954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:d8803a301d4931ada77da37647e3f7aa0ee72620d3b7d65b6647bbc3d5b41731

Observation f085f98e-51fb-4643-ae48-a07375c0a9be · outbound

This paper cites Background.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Background

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:32:51.644519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:846de1abcff7506633fba80289b1454406d8ae919d2810ac09171e3edc2e8e1e

Observation 6cc63cf4-b307-4574-afd0-0f5c3cc17ba4 · outbound

This paper cites Regarding our model, in the case of XLRS-Bench, the budget Bs is assigned as 180,1320,1600 , and 8000 corresponding to the four resolutions.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Regarding our model, in the case of XLRS-Bench, the budget Bs is assigned as 180,1320,1600 , and 8000 corresponding to the four resolutions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:32:51.641982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:b0a75cb4e93a1082809041cae942bc1fc04b275f527467f0c9c0f316c3a17b14

Observation 94aa12ee-4d6a-4e80-ae1f-c0c9e6d44b19 · outbound

This paper cites By integrating spatial distance, the clustering process effectively groups tokens that are both semantically similar and geographically close.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing By integrating spatial distance, the clustering process effectively groups tokens that are both semantically similar and geographically close

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:32:51.647864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:82f5e9497565100df1c4c273373444905609546e8d3e957e9a2a348ec9993061

Pith citing papers

No inbound Pith citation observations are available.