Pith. sign in

Paper Citation Record · LEDGER

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features

As of 11 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2501.10144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10144 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:26:44.966467Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1838a2c-9fe3-4eb5-bfcd-0e56e5146ecd · outbound

This paper cites Canopy averaged chlorophyll content pre- diction using convolutional autoencoder on hyperspectral data,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Canopy averaged chlorophyll content pre- diction using convolutional autoencoder on hyperspectral data,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.545801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.847841Z digest=sha256:65e6f191887000befb16a7d25e5b554b0c1dcda2a6f7e727119e30ae5e83060f

Observation 1694854d-5b7e-4220-8dee-6b573e39c490 · outbound

This paper cites Enhancing deforestation monitoring in the brazilian amazon: A semi-automatic approach leveraging uncer- tainty estimation,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Enhancing deforestation monitoring in the brazilian amazon: A semi-automatic approach leveraging uncer- tainty estimation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.532352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.852903Z digest=sha256:84fe57738334602d9caa96987e8c61d08b6eec68bd6eb68e1fdb31deff4f9a32

Observation bc850173-9504-4c33-9c38-0a9ff8a53add · outbound

This paper cites Sen1floods11: a georeferenced dataset to train and test deep learning flood algorithms for sentinel-1,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Sen1floods11: a georeferenced dataset to train and test deep learning flood algorithms for sentinel-1,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.517400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.857234Z digest=sha256:0a101a572c95cff7f32925252a4e6716938c4481f0fa639c6c6275a69fe9d7d8

Observation dbe34249-0774-4a29-bdf4-dcb65b466d23 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Improved Baselines with Visual Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.861667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.861667Z digest=sha256:951578128ff2bd3d265dcac578d80eb2b3f9272ba2c38d717ec9e6d18d4cff9e

Observation 4cf2e9f3-cdc3-4d4f-a818-d66cdceaaecd · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.866767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.866767Z digest=sha256:3fa25924034749aa0caafc735e3a0967ef0c0cc42648b9f428f8bbc9ae37f589

Observation 4b4ef0d6-8697-41fc-9267-72625460c83d · outbound

This paper cites Visual instruction tuning,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Visual instruction tuning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.870831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.870831Z digest=sha256:9e6ec82724a60fa340d3f2fd90e54e01f7ed9d93cb48ec2a1ef7cd2a63a383bb

Observation eb33002a-ee4f-46f5-ad82-f8456ef1c552 · outbound

This paper cites Blip- 3: A family of open large multimodal models,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Blip- 3: A family of open large multimodal models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.875200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.875200Z digest=sha256:b776f1864993de9d0b8934106954e8c8e1629b1b459c9bc3693f81b3a2872e23

Observation 2af9673b-755f-4fd6-ba7a-cd1e8b138c42 · outbound

This paper cites Geochat: Grounded large vision- language model for remote sensing,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Geochat: Grounded large vision- language model for remote sensing,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.488360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.879243Z digest=sha256:c0ede1206c30bad0340fb32c043341bd30acc979257c6dba525daed76f48c1d5

Observation e99ded90-b835-4b0e-adb9-5f79cfd646c5 · outbound

This paper cites Geollava: Efficient vision-language models for temporal change detection in remote sensing,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Geollava: Efficient vision-language models for temporal change detection in remote sensing,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.476126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.883227Z digest=sha256:ec16621b36ec8ebbd43d2fa81de8c6affab9b6029246df393a08436af8690388

Observation ca4e9245-c9bb-4459-b843-cc1052393104 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.887132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.887132Z digest=sha256:c211606a65a17123b66c8a20c14c0dd35e3def191266b855a9a255663e514a0e

Observation fd2f19f5-4f7e-45d9-8907-80706657e817 · outbound

This paper cites reBEN: Refined BigEarthNet Dataset for Remote Sensing Image Analysis.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features reBEN: Refined BigEarthNet Dataset for Remote Sensing Image Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.891341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.891341Z digest=sha256:5b58bd19763f930ebc2bbe58ca60a4a55936cf975ab31083ba8804b8bf841de0

Observation 4863534a-e817-46a7-9594-825c2c60a759 · outbound

This paper cites First principles residual resistivity using locally self-consistent multiple scattering method.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features First principles residual resistivity using locally self-consistent multiple scattering method

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T19:26:45.085727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.895190Z digest=sha256:7358001537477f12950f37d58bdfea3e7b2a755ba63a770b93f991258e9e8522

Observation 58760e58-4351-44e1-a8ab-6582f79edd20 · outbound

This paper cites Spectral reconstruction from satellite multispectral imagery using convolution and transformer joint network,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Spectral reconstruction from satellite multispectral imagery using convolution and transformer joint network,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.463752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.899023Z digest=sha256:70dba7f3dd192c9fdf510201bcb04835edf21889d34891375b2a43047ceaccb2

Observation 061acac1-f5ce-41fb-b4cb-fef4bbe9b770 · outbound

This paper cites Ringmo: A remote sensing foundation model with masked image modeling,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Ringmo: A remote sensing foundation model with masked image modeling,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.451307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.902304Z digest=sha256:df6e81b356cce59a25fbc00199833491cf81cc8aabe37b8b4a7838fabe47761c

Observation 2cbe90f8-8d2f-4d07-94e2-66cf27b0fd8a · outbound

This paper cites Masked autoencoders are scalable vision learners,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Masked autoencoders are scalable vision learners,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.438591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.906082Z digest=sha256:50ec9a24360874c2c1cb4fe0678b301cac078df4b54ac88a2d54394df4906383

Observation bda0b774-b894-48e7-aaed-7e210aad771b · outbound

This paper cites Late-time Evolution and Instabilities of Tidal Disruption Disks.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Late-time Evolution and Instabilities of Tidal Disruption Disks

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T19:26:45.067195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.909901Z digest=sha256:05dcc6aefdd829cc24c3ebf140969e6328e4580a91b872f46e09b43a6e8a1dae

Observation 41ff7c79-c082-4b5c-9f4e-158b0e998fa9 · outbound

This paper cites Spectralgpt: Spectral remote sensing foundation model,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Spectralgpt: Spectral remote sensing foundation model,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.425343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.913939Z digest=sha256:4fdf09727c4d7da3a44ad7bd669c2dad09dc7bc307a00048662b31be98927aca

Observation 21f93b25-e620-4a28-beb6-81179c3601da · outbound

This paper cites Rsgpt: A remote sensing vision-language model and benchmark,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Rsgpt: A remote sensing vision-language model and benchmark,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.412442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.917729Z digest=sha256:18ead74f94917d4592b54f7ef7ab8457238a7c0862258794dddcd993fbab6bd4

Observation 1937ae6e-5625-4d8a-b76b-53a8f2514bc6 · outbound

This paper cites Earthmarker: A visual prompting multi-modal large language model for remote sensing,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Earthmarker: A visual prompting multi-modal large language model for remote sensing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.400090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.921340Z digest=sha256:8c4f43e8ccd2820baeca4f4e8d39860e01926304ae1edaa2b83e8f915462c19e

Observation afed98c1-59ec-4a35-877e-4be148390dd5 · outbound

This paper cites Combinatorics of generalized orthogonal polynomials of type $R_{II}$.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Combinatorics of generalized orthogonal polynomials of type $R_{II}$

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.925055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.925055Z digest=sha256:dc3ff7c4d49117a5ded7207e6e20b9cf0818eec5053cf7dcdeadede95957d2ee

Observation fb283d58-5694-487e-bbb6-c1bfaa93ed9c · outbound

This paper cites LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.930836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.930836Z digest=sha256:15f4cbd6ffe7356f53a5a799e373d139fc72d80d04554686439cfd18d4d5027c

Observation 107ffc3c-f686-4b15-b1a2-65da33d38e76 · outbound

This paper cites GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.935014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.935014Z digest=sha256:a1be6b7a24e6fda8d186627a340a0eee757bb79fc29642130a55e6c1355663c0

Observation d72f046e-e008-4199-937e-5e13bb5abb3f · outbound

This paper cites Skyeyegpt: Unifying remote sensing vision-language tasks via instruction tun- ing with large language model,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Skyeyegpt: Unifying remote sensing vision-language tasks via instruction tun- ing with large language model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.281077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.939029Z digest=sha256:37430e20b6e74910aff331d94f4a836abbbbf2462273b4082d6c88fa34a73645

Observation d51d1951-6a09-4315-942c-46bc30b80357 · outbound

This paper cites Ringmogpt: A unified remote sensing foundation model for vision, language, and grounded tasks,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Ringmogpt: A unified remote sensing foundation model for vision, language, and grounded tasks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.266758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.942758Z digest=sha256:625aa24dcdafe00a01e9e88f2535a9eb9a9474476b58a1e459db79ecb71ca592

Observation f2d47d54-df15-40e7-9b89-6f41eef3f49c · outbound

This paper cites Teochat: A vision-language assistant for earth observation data,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Teochat: A vision-language assistant for earth observation data,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.252934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.946669Z digest=sha256:19ab3ebfbfdc0685d012cdf36fee9a1ee52e101a39be20fb1201b11eabfec7c4

Observation b576194a-f80a-4478-a2ef-8059ee3c040e · outbound

This paper cites VisionTrap: Vision-Augmented Trajectory Prediction Guided by Textual Descriptions.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features VisionTrap: Vision-Augmented Trajectory Prediction Guided by Textual Descriptions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.950567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.950567Z digest=sha256:63a8d832532119e8c2c0f2e83e0abd671bbd5692d7047b6dded83a47c145ac5d

Observation eeac5fb0-3a0c-4606-993f-0df4e692bf1e · outbound

This paper cites Functional map of the world,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Functional map of the world,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.239515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.954718Z digest=sha256:bf708a14c15065198146b92a0060c95ddd0c979d1c08dbf8a63a022b7c7283d8

Observation 5901d53d-5e2b-4278-b5c2-7c087f02401d · outbound

This paper cites BigEarthNet: A large-scale benchmark archive for re- mote sensing image understanding,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features BigEarthNet: A large-scale benchmark archive for re- mote sensing image understanding,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.227417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.958497Z digest=sha256:b41e8a1ae440c852750b9d38c998ab3598c4fe9c4ac0fcf0353612c3e82a9b83

Observation 1d9f7c85-4c4c-4afb-8b45-e0d0ccc77b26 · outbound

This paper cites Distributionally Robust Receive Combining.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Distributionally Robust Receive Combining

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.962285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.962285Z digest=sha256:d9be05f9d1b203e10bef2d0cd8ee6d73e412280cb89093e5ab78b3560f8b46a6

Observation 8df9fe71-2f7c-403c-83cb-3bf198841104 · outbound

This paper cites Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.214254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:26:44.966467Z digest=sha256:1874dd0723b7b8e92b64b6b5636d5dd16fdaa602d6deb75603cd551551edc5c7

Pith citing papers

No inbound Pith citation observations are available.