Pith. sign in

Paper Citation Record · LEDGER

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

As of 15 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 7 inbound Pith citation observations for arXiv:2501.06828.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06828 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:53:49.353753Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:09:55.854741Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T00:48:24.793323Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b57f24db-3cc5-4524-b60a-e8a3273e966d · outbound

This paper cites RSGPT: A Remote Sensing Vision Language Model and Benchmark.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:47.960477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:47.960477Z digest=sha256:b20455aff193213250b0be6e5f7be64c0ee03ff249ccf5b5ca2eab55b38a6cd9

Observation 479916e7-27cf-4b1e-82f4-0518bc6bb86e · outbound

This paper cites Geochat: Grounded large vision- language model for remote sensing,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Geochat: Grounded large vision- language model for remote sensing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.402256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.024344Z digest=sha256:2298dd1432c39c0f9146ce94c60341870daf40b2befa2203c1ea07086f7c5b00

Observation c4737594-d29f-47d6-881d-6c84d4b989a5 · outbound

This paper cites SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.029329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.029329Z digest=sha256:7372f10c9027f00f186d6bd1bb2cb4f575e3668c7a0ee240f7acd43d5ec5a144

Observation e517ff45-b20f-4bf6-8563-74498db1a163 · outbound

This paper cites Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.392167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.034921Z digest=sha256:baa4599425ce8a240e82342a9f7189d847d7a0fa1d3b52169e6d8060aba3df33

Observation 74407812-8a35-4aae-8fe8-744d5b6b82ba · outbound

This paper cites Bb-geogpt: A framework for learning a large language model for ge- ographic information science,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Bb-geogpt: A framework for learning a large language model for ge- ographic information science,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.382002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.039512Z digest=sha256:5bdb3c06b50c7d121e20bbe1a159f1a91bb2eaecc6b5fb5d426445590631eaa4

Observation d87e0878-d02d-47c0-82e8-3265553242e3 · outbound

This paper cites LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.124779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.124779Z digest=sha256:b7efb4c533a03a5b56b4f128ac5d90b6f6196c05d70065fe3420e0fc9fe597df

Observation 2db5c505-469c-4d4a-8d02-c746e5aa3b60 · outbound

This paper cites Mtp: Advancing remote sensing foundation model via multi-task pre- training,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Mtp: Advancing remote sensing foundation model via multi-task pre- training,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.373021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.225842Z digest=sha256:f706e67c5a3e439fc468466dc75b7f80c7856511a450bab1d04df2a6ed31106d

Observation ed8fa80b-baed-4c4e-bb80-6eaef91a6e75 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.229191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.229191Z digest=sha256:29bdf74fa3e5398d0b9f8f32be021dff993ed1675551a58fe484abda733508bd

Observation b4ed543d-7b2f-49b6-8fd0-eaf5c210dcdf · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.233807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.233807Z digest=sha256:a6416e4c942a601197b947e74f9680df240b4be935decb487615568aeffe0bc9

Observation 44b65459-5320-461b-bf2a-5d9080062463 · outbound

This paper cites Earthmarker: A visual prompting multi-modal large language model for remote sensing,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Earthmarker: A visual prompting multi-modal large language model for remote sensing,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.365058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.239203Z digest=sha256:4cc0457f38b1eab1e7f928b8afd7ed5178299211fdccfd1c0ccd3d02ad2bf66b

Observation fbd223b6-275f-4767-b043-163aaa67b744 · outbound

This paper cites Rsvg: Exploring data and models for visual grounding on remote sensing data,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Rsvg: Exploring data and models for visual grounding on remote sensing data,

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-08-10T20:53:48.243234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.243234Z digest=sha256:da6afe699467c4312d561a9461fa4745f8c1f0a66fbf79bba8f8fe6143b31bc6

Observation 74e8ace1-bf93-4ecc-9ce8-3337ffa99726 · outbound

This paper cites Samrs: Scaling- up remote sensing segmentation dataset with segment anything model,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Samrs: Scaling- up remote sensing segmentation dataset with segment anything model,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.355995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.247566Z digest=sha256:ef765c0035b465453a698ff96841952fdab1f25bcee2e66ef2dc4e85d223344d

Observation 676270de-e63e-4f45-8e34-f2115e1e0a37 · outbound

This paper cites Rrsis: Referring remote sensing image segmentation,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Rrsis: Referring remote sensing image segmentation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.347108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.251055Z digest=sha256:cb34209796d1865cf30bd8770675182bd36c3e03003903148bdabbdb5406be9a

Observation 0a0d5c78-5d9e-465f-a00d-8f79f4e83ca6 · outbound

This paper cites Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:53:50.015691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.254286Z digest=sha256:8c95ca1d724d86facdd56b043d34d1ba0d8a76914c759244d3cf4d53dbd920f2

Observation e286015a-c9ee-4903-94b0-ece529454917 · outbound

This paper cites VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.302408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.302408Z digest=sha256:be8eff6116c1b9e4c4362ec84329b686f1e2645fe2a0fc6de4296210dc99490c

Observation e734e1a0-b2b5-494a-b871-5c8184f1f067 · outbound

This paper cites an unresolved cited work.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:53:50.336893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.371796Z digest=sha256:dbc97a5e1f75d344c1cd8d1b1076086524ed1b5b6729f90bd9d54d2476029eb1

Observation 23036da1-0560-4222-9955-73812f1a3d9f · outbound

This paper cites an unresolved cited work.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:53:50.325240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.375743Z digest=sha256:f975d882248d8d773cc263b7b5270804edd1987409c934508ce6c308e511fe53

Observation 132a4d93-a91a-4b8d-bbd1-c496dcbb6acd · outbound

This paper cites an unresolved cited work.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:53:50.316077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.380182Z digest=sha256:2635c3425f3c156ac5ff30f39cbfc9526adc983b86b6e82999c644f558cd504d

Observation 60e82997-5cfc-43bd-a2fe-72a5345fbd19 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.385169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.385169Z digest=sha256:59bf14a14c2cf0f7b6566d21fbeb35832120500ed9682bbed626ccedba1dc2e9

Observation 72c10116-e0cb-43bf-af8a-f827f1acfd0d · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.389986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.389986Z digest=sha256:bbac89f2ec10886a60ff208ea91c7e1a12689b41d87e2a9625cfe1a02b9bc378

Observation 9f5833d6-7dd0-4bd2-9e0e-1dafa976c291 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.394343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.394343Z digest=sha256:1debe7cdcf857d3d1b9d6ebe6e872e7b4fcc4874e9d6e45f324b9b206a88a28c

Observation 2ad8e6e5-a33a-4a77-a1b2-694f80493fcc · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing CogVLM: Visual Expert for Pretrained Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.427182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.427182Z digest=sha256:8bd0c06f4b6718ae91f0565d1b7c5540e4bba7b7e484229f8d699510b6a729ac

Observation d2218fbd-ede2-4ba3-9b6f-0d325e5b5a67 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.493042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.493042Z digest=sha256:299b601060e0cc2d785821dce3d73dd02b58c7fba6f935e34517c841a867e9c4

Observation 4f6c3bab-f864-42f2-bede-2d92b567c5dd · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.528819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.528819Z digest=sha256:c59839119f41515bbea9cb0adb247604bb00fcc3346d47baee497065c8a0d6d1

Observation ed754b60-189c-4733-bfe6-b64d3a37e722 · outbound

This paper cites Instruction Tuning with GPT-4.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Instruction Tuning with GPT-4

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.533532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.533532Z digest=sha256:cfac5b284ee8f6b9dd4fa536831d4929b034da44387493376f8fc6f2f1644e00

Observation ddc7b7af-d2c7-45dc-a28a-874ae8c0d7c1 · outbound

This paper cites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.539652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.539652Z digest=sha256:df1044fb9459525fae87a114a96d2cfb3adef7271c0e14e9b2f99b225329a2fc

Observation 6a71dd86-2518-4f1c-b76f-f687a36fbd61 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.544572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.544572Z digest=sha256:6e31d5cab5aa939d99ac0e7b28cf45a33bdc1562982eac71b17ebb94b3ccba46

Observation 57de04f7-f32d-42a0-82ae-abf953042751 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.548802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.548802Z digest=sha256:6a59856e457cfc3670acd55b0b589d49404c8a92460c141252d4db0c7629f824

Observation b1f5dab7-0915-4a14-8676-0d8951c142fe · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing LISA: Reasoning Segmentation via Large Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.552690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.552690Z digest=sha256:fdd105444df98486f206d5dec9f851c6b5a54d50bc72905b3bcc2486e7684c80

Observation ad57b2d2-6bb5-4359-90fd-37aa725ec24e · outbound

This paper cites Segment Anything.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Segment Anything

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.556327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.556327Z digest=sha256:d0eb975ffda6ebbf111bcf6d518be93716a949ece5d8289c96990309ddbc04b9

Observation 86681f54-61b8-4a88-a6f5-273f4ef7fc9f · outbound

This paper cites PixelLM: Pixel Reasoning with Large Multimodal Model.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.653877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.653877Z digest=sha256:d27fab7be20fbb04ae3218065fa4d2415142d85d191b94d7bb3ac1d2a03295ef

Observation fac7f3f5-a830-40fd-8a84-133276b70349 · outbound

This paper cites GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.729794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.729794Z digest=sha256:b54da043fbc3101d10b6bce3ea7f88c515252a2a11223060cfc2abca6098845e

Observation a59bfb6a-f2d4-4226-9d05-6de0780ababb · outbound

This paper cites Object detection in optical remote sensing images: A 13 survey and a new benchmark,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Object detection in optical remote sensing images: A 13 survey and a new benchmark,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.306212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.733812Z digest=sha256:d6213bad08cf0f241593a6dcdabd95074e8366409aaea621c63853867694c9df

Observation 69fb0bfc-0a14-469f-b565-175e31227b04 · outbound

This paper cites Object detection in aerial images: A large-scale benchmark and challenges,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Object detection in aerial images: A large-scale benchmark and challenges,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.295806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.844459Z digest=sha256:1f10b36b120c580495ed686da818602f470fe0819184ff465df967498a863022

Observation c07902bd-c878-46fc-8958-fed9ce27a3a1 · outbound

This paper cites Fair1m: A bench- mark dataset for fine-grained object recognition in high- resolution remote sensing imagery,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Fair1m: A bench- mark dataset for fine-grained object recognition in high- resolution remote sensing imagery,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.915827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.915827Z digest=sha256:343603c2628daf29f780b1f674bb88b8d33ac7f3f84bbad59cea739472e8e065

Observation 74699657-9a64-4fbe-815b-0bff53a95b28 · outbound

This paper cites Aid: A benchmark data set for performance evaluation of aerial scene classification,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Aid: A benchmark data set for performance evaluation of aerial scene classification,

Reference 36

Resolution
metadata mismatch
raw_fallback, observed 2026-08-10T20:53:49.755682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.920809Z digest=sha256:ae5bd4942268ed55782718b76b03c5c2b2330c2868f496a88acc54f1435b6ee6

Observation 276bb595-2037-4757-a1ea-05da2a137c85 · outbound

This paper cites Nwpu-crowd: A large-scale benchmark for crowd counting and lo- calization,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Nwpu-crowd: A large-scale benchmark for crowd counting and lo- calization,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.924161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.924161Z digest=sha256:4c8d414a9e025cb988d7335db8c1b9e2d7b3b924f77a0d442ca79d0aea95163c

Observation 5cc974fa-04de-4841-94af-f937cc7041a8 · outbound

This paper cites Eu- rosat: A novel dataset and deep learning benchmark for land use and land cover classification,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Eu- rosat: A novel dataset and deep learning benchmark for land use and land cover classification,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.286416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.929800Z digest=sha256:977c6a534ea0d30808278b2f21142a6739c931acd0b855badbdcbc82f3ac17ab

Observation f6f8f5c1-212f-41f0-bf55-3d951d630960 · outbound

This paper cites Rsvqa: Visual question answering for remote sensing data,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Rsvqa: Visual question answering for remote sensing data,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.275459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.934859Z digest=sha256:a6d529143b8a73e957eba84ea9d83c861f3afa0eb3b66711f227bab44c3ec6cd

Observation 0b1b3e55-2252-4aa4-99c2-8052b752ed3e · outbound

This paper cites Memory Matching Networks for Genomic Sequence Classification.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Memory Matching Networks for Genomic Sequence Classification

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:53:49.604414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.943707Z digest=sha256:f56a67081de348dbd21aeb73d70f160d5e11e89b179907fff2ab396d1d94ed12

Observation f9f2c85f-b234-4ac9-8f1e-f914575d4ec1 · outbound

This paper cites Unsupervised learning using pretrained cnn and associative memory bank,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Unsupervised learning using pretrained cnn and associative memory bank,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.256559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.948787Z digest=sha256:93778f7003504490eea7123fdab87f60d1a2f664143816a6f8c291404f23ea81

Observation d64bc291-b7bb-4597-b758-eedc1136c231 · outbound

This paper cites Point cloud classi- fication via learnable memory bank,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Point cloud classi- fication via learnable memory bank,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.246020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.952127Z digest=sha256:5f3b683acf507242a4b51cf6751245f81823100a322373fac91a06557d8bc0c0

Observation 3e1a7e91-9290-4fcd-a985-29f759521b7d · outbound

This paper cites A new approach to automatic memory banking using trace-based address mining,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing A new approach to automatic memory banking using trace-based address mining,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.235618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.960698Z digest=sha256:4cd1a6aa9aec67187fcfdefa602e5f5664c7f40a3f41ec6b61fb2e94c47c2d46

Observation 2c9ba90b-a9b2-406f-8d28-2c936e290849 · outbound

This paper cites Visual anomaly detection via partition memory bank module and error estimation,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Visual anomaly detection via partition memory bank module and error estimation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.222736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:49.040394Z digest=sha256:3cd450aa84cdea97289eb77a48544faf7b460d4e49e38edfebb63e895af650f9

Observation ef5425ed-c87c-4bb3-ba1c-9a302af34009 · outbound

This paper cites Mamba: Multi-level aggregation via memory bank for video ob- ject detection,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Mamba: Multi-level aggregation via memory bank for video ob- ject detection,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.210886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:49.134468Z digest=sha256:126118eaf2fb10345308be291a8650b658fdb94f04ba40872633a0d749d9ebc0

Observation 5bfbfdac-8b31-4c20-8aca-15c4f3c3fc2a · outbound

This paper cites Semi-supervised se- mantic segmentation using unreliable pseudo-labels,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Semi-supervised se- mantic segmentation using unreliable pseudo-labels,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.198699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:49.138237Z digest=sha256:6e5a8cb061f455d9f6b6e590ca65c707c2d6b655d9eb17a22f36d390c2957544

Observation 415f4932-3d03-44c9-a949-187556a79f3b · outbound

This paper cites Weakly su- pervised semantic segmentation by pixel-to-prototype contrast,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Weakly su- pervised semantic segmentation by pixel-to-prototype contrast,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.187358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:49.142549Z digest=sha256:5aab4e3b464a42caf5aedc895a0c0cea351c3822c107af7d9abf14ef9627e757

Observation 684a14ed-537d-4607-87aa-0a3b54aafca4 · outbound

This paper cites Memory-based cross-image con- texts for weakly supervised semantic segmentation,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Memory-based cross-image con- texts for weakly supervised semantic segmentation,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.145828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.145828Z digest=sha256:e78e4c1a9b103912a0dad28a2e04b7da5123fc30310aabdb093a03757d3e592a

Observation 22fb036f-3f3b-49b7-9d3f-73c5c9f51bce · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing SAM 2: Segment Anything in Images and Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.150228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.150228Z digest=sha256:9cb214271a702036e639acc0d6ee63aace188a50398976e17265d182ebe98459

Observation 0c4f3726-bc66-403a-8186-5295f5245880 · outbound

This paper cites SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.155213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.155213Z digest=sha256:efa629aef8dd8efe5b09768075d81e8f06d803b02793eb2e3943bef58c21fb58

Observation acc63c6e-dc1a-491c-9f95-9c707cdbb79b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing LoRA: Low-Rank Adaptation of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.159740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.159740Z digest=sha256:67220f3df72468ced4a2c716fb7f2ea73791dcf619108d793d0dd662c39d70e2

Observation fb343ad8-1d33-493f-939c-733a0419130b · outbound

This paper cites an unresolved cited work.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:53:50.176587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:49.163538Z digest=sha256:2d8604f8dfa114e92a4a38974f7c9fef7c10e9fefd9c5b51c37bad9c3a96981c

Observation 27cc5b5b-8844-4ab4-9905-14ce799e1e9c · outbound

This paper cites Bleu: A method for automatic evaluation of machine translation,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Bleu: A method for automatic evaluation of machine translation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.164726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:49.175607Z digest=sha256:6cc95a368f3dad03856bf3d53f88a60d3cceb43241624a35f0f27de97257f2d7

Observation d9a9a7f3-0cd3-45c5-bb48-edcc708c2fcb · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Rouge: A package for automatic evaluation of summaries,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.236877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.236877Z digest=sha256:d65aa5cd8f7ef54d549d1cbc67b69ccf49bdfddc451a497cc40befe7f2b10975

Observation 2072881c-4e83-4d51-b916-1a177d2167bc · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.148712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:49.345436Z digest=sha256:b494bfb2f0f8cd4e34c1d276b1e452fbe8781867e0fbcc90452a0b40f1561fc9

Observation 6b402188-5f59-45d8-b7e8-43466cc73e3f · outbound

This paper cites Cider: Consensus-based image description evaluation,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Cider: Consensus-based image description evaluation,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.349243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.349243Z digest=sha256:7488ebd3447393c11bd5cd000a38328d0edb838cc81adeee91d12084a9366c9d

Observation e11a5396-91b5-45f2-a37e-e9b145a0dd63 · outbound

This paper cites CLAIR: Evaluating Image Captions with Large Language Models.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing CLAIR: Evaluating Image Captions with Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.353753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.353753Z digest=sha256:d387216439a6f78974473f2f2fc8a658a87ab1aa530c6f9d86776d7dc28ac538

Observation 4d9b901b-18c8-4fb0-ad40-bd1ab864714a · outbound

This paper cites 1109 / tgrs.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing 1109 / tgrs

Reference 644

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.265991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:53:48.939219Z digest=sha256:f880ea3aaf3fd13e804280b5c5f973cfc0b1227e7b6e29e82ad80e0da6d1cc7b

Observation 718daccd-d34f-4936-b586-5d5d9b13f84f · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.166936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.166936Z digest=sha256:d8783f0bdf6a152c916f429f0d4050feead1ba29bed4ecabfcfd9f8fd1d535db

Pith citing papers

Observation af11f5c5-d302-4dc7-a9f3-23d3a0e51e42 · inbound

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions cites this paper.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:55.854741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:55.854741Z digest=sha256:937f489e956b3fed6a48c015485e7b039a9346598cd538557ee3fb323eb8d354

Observation 834d3338-f0eb-4b07-8389-c1ba5296fd3f · inbound

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing cites this paper.

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:13.258836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:13.258836Z digest=sha256:d75940c6eb451cd8f31abbf85ba3b7f076b1cabaa3268519a841a6305a4623db

Observation 73294226-ce04-4cd9-bb65-ad36eefdba9f · inbound

SHARP: Spectrum-aware Highly-dynamic Adaptation for Resolution Promotion in Remote Sensing Synthesis cites this paper.

SHARP: Spectrum-aware Highly-dynamic Adaptation for Resolution Promotion in Remote Sensing Synthesis GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:48:24.795113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T00:46:34.777585Z digest=sha256:2757304ed7e86aaceedbcfe33d3019f69edc9287320ac58e6b935b3b18a993ac

Observation a421a2c2-4c4e-42fe-bacb-e066eecb8672 · inbound

SHARP: Spectrum-aware Highly-dynamic Adaptation for Resolution Promotion in Remote Sensing Synthesis cites this paper.

SHARP: Spectrum-aware Highly-dynamic Adaptation for Resolution Promotion in Remote Sensing Synthesis GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T20:17:28.243545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:17:28.243545Z digest=sha256:6f5a00ad1e759990f1ada17d6d36de9e6b50081df3b957ebe2e17dbe05842c6c

Observation 6f3903d3-6baa-450f-87e1-bcfa7ce2f8bd · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.112891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T09:20:54.635375Z digest=sha256:6a9aef02059d96d705d998c6a18e4337e0ff02070eea1d3fd4c1b5ae91767706

Observation 6b1db3a5-69b2-4226-ab85-d9179c01d41b · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:02.323102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:02.323102Z digest=sha256:04dd5ccf0e12c3f5d5fb9ff4ba4b8d30d75b35fbf5ff79c1c576c6eac2ee8ae3

Observation 1b903af9-48f8-45b4-8a74-25804925781b · inbound

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding cites this paper.

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T14:09:30.395518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:09:30.395518Z digest=sha256:2932d6cccefda1a54558ad271647692ffc9d94df93a42441474e15579ae090ec