Pith. sign in

Paper Citation Record · LEDGER

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing

As of 19 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2607.17386.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17386 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:11:04.192167Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c911503-5255-4be9-9c32-4a06b43b0c4a · outbound

This paper cites One token to seg them all: Language instructed reasoning seg- mentation in videos.Advances in Neural Information Processing Systems, 37:6833–6859, 2024.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing One token to seg them all: Language instructed reasoning seg- mentation in videos.Advances in Neural Information Processing Systems, 37:6833–6859, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.691639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.691639Z digest=sha256:5ceebcf902505d9d8bca7765dfb93480b6b374710a444ae17524e76d101930b6

Observation e726ec18-5a18-4e01-a66d-acf19c22c573 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.744246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.744246Z digest=sha256:da50352719070c490982dcb2f7bfa7d40b3537c80e4e324b4d4fb4680505d4de

Observation bc7c76ca-1343-4e91-9975-31bbadbdd027 · outbound

This paper cites End-to-end referring video object segmentation with multimodal transformers.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing End-to-end referring video object segmentation with multimodal transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.790110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.790110Z digest=sha256:96692bf747374796e76091e31a2ff0443b617a1adf618a1850529a92e405ad90

Observation ebeeb569-da85-4acb-8d79-5b96c294852e · outbound

This paper cites Streamingtom: Streaming token compres- sion for efficient video understanding.arXiv preprint arXiv:2510.18269, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Streamingtom: Streaming token compres- sion for efficient video understanding.arXiv preprint arXiv:2510.18269, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.854409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.854409Z digest=sha256:30cdd72342ab9d8941e4b04ed805ddaf794a6749624d1f173ca8792d714b1508

Observation aac304ac-f50d-4d06-89f3-224d7201ef3f · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.919449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.919449Z digest=sha256:a617e708182bc84036789650e9f57f1010bd75af4a5a8228796fee3d806f3c33

Observation f0495812-3451-4c19-85ed-ad679685c4c3 · outbound

This paper cites The unmanned aerial vehicle benchmark: Object detection and tracking.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing The unmanned aerial vehicle benchmark: Object detection and tracking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.984738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.984738Z digest=sha256:6794fde81dcf51eea54de329e60b29565c9af6ce18f1c272caf119317a27f4f3

Observation f6ace37c-9484-43c0-8418-585971a9f62c · outbound

This paper cites Framefusion: Combining similarity and importance for video token reduction on large vision language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Framefusion: Combining similarity and importance for video token reduction on large vision language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.082341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.082341Z digest=sha256:7d59d010e14fcfa58bc5f9f19056fa54011eceb28a2c56ec9c04360d8d156a81

Observation e7486384-4444-47a1-b64c-6ab04edcd40b · outbound

This paper cites Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.204751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.204751Z digest=sha256:6a213f6166eb4ad8f163f4f0aa5e3e16d7f923482deed47756d2c125f0acd800

Observation 954877f9-8f77-44ff-bdd2-48c34a27fa1d · outbound

This paper cites The devil is in temporal token: High quality video reasoning segmentation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing The devil is in temporal token: High quality video reasoning segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.296336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.296336Z digest=sha256:707088f65289b5bb96aeb1c4f237a9b993dd69a814d92d7371fe164c14a65661

Observation c2dae05a-a68c-4d7b-9543-b872f27f70de · outbound

This paper cites Rsgpt: A remote sensing vision language model and benchmark.ISPRS Journal of Photogrammetry and Remote Sensing, 224:272–286, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Rsgpt: A remote sensing vision language model and benchmark.ISPRS Journal of Photogrammetry and Remote Sensing, 224:272–286, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.367418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.367418Z digest=sha256:48c3e529cca25790314e290477a89807035c9f461a8b6bacd0f3865b4a916399

Observation 69ff8e28-3474-4e55-8196-1d2bf37d04b2 · outbound

This paper cites Prunevid: Visual token pruning for efficient video large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Prunevid: Visual token pruning for efficient video large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.457054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.457054Z digest=sha256:b201eff3f604df868a4011986b8ac8c0afc54f0e671a517224ced43738e5e570

Observation dfecb740-b81a-49cf-a800-cfeac3bca07f · outbound

This paper cites Multi-granular spatio-temporal token merging for training-free acceleration of video llms.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Multi-granular spatio-temporal token merging for training-free acceleration of video llms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.521166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.521166Z digest=sha256:28abccfe413ee549b9a16cfe0426e09cc61146c4f2054d69e45da49cf3e0d674

Observation c57608d3-d131-4995-8b1c-eff11246edac · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.580630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.580630Z digest=sha256:6f71703d8d851927a764487ad7a270441543a2e64ce262b21da0428e6aafa5a2

Observation a485416f-d454-4b76-9ef5-7fc4ede0b8cf · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Lisa: Reasoning segmentation via large language model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.696273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.696273Z digest=sha256:0698e19ea4fc48a40f73acb9e21a4a36973a276258642427c89c522bff49a04b

Observation 83cd5b5c-dd0b-47d7-a5e5-ee9ef889736f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.761048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.761048Z digest=sha256:cf850ea4ebf401055ffc4e88eb75ba59dd212ff092f9cee2382d5cb32d441906

Observation 6731052e-dbfa-41ad-814f-7b040ed21374 · outbound

This paper cites Referdino: Referring video object segmentation with visual grounding foundations.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Referdino: Referring video object segmentation with visual grounding foundations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.849069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.849069Z digest=sha256:69a40e796ece16b0fc723b1455b3822057dd3f879deece58647299807151e032

Observation 4f45cead-57fa-4193-934d-3aad75ef807a · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Video-llava: Learning united visual representation by alignment before projection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.971650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.971650Z digest=sha256:bd47aca9b7061cfa2a5d404d86fa3145c4872a891e6b74d466d6cc7d413d5422

Observation 27c684a3-7358-48c6-886a-e007211e6294 · outbound

This paper cites Glus: Global-local reasoning unified into a single large language model for video segmentation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Glus: Global-local reasoning unified into a single large language model for video segmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.034375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.034375Z digest=sha256:a51cece2eb55b45498aee3e4d17e543a2b3ff6e7dac2476a34d020fc70e90dfd

Observation 1e1e3f91-2614-4465-aaeb-bde5dc3df7a4 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.108301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.108301Z digest=sha256:c9dfbbf6e5a5ac950b93b615f08795b6b07ea95ffc445ffd3d215d2ad6a566dd

Observation bf66e690-c950-4d14-b84c-12030a074936 · outbound

This paper cites RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.265647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.265647Z digest=sha256:7362dee329ee4f2c7aef7b886c280b3eb9f17d0fe3b145fb91b88358877799ef

Observation 7da24c0e-c78c-49c0-915a-672bacdb4bf6 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.304478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.304478Z digest=sha256:320da0ac296b9491c27849dd110693eee5286d4ee9c55c203c155d79fdfe493d

Observation 0ebe0656-7020-49d2-88e4-e00b58c04323 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.369009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.369009Z digest=sha256:7761babe99ce268b5881ea78066e2cf2b510966677cc55e209fa9c02e8f618aa

Observation b7b0d143-8b00-4c13-8bfd-e131ad13230a · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual grounding in videos.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Videoglamm: A large multimodal model for pixel-level visual grounding in videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.428042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.428042Z digest=sha256:c204b7d2022b48e2f59820db143c880738172b8dae9549da12344a9dcb15ef1f

Observation 113d5682-d67d-40cb-b88c-a2e3a48dad5a · outbound

This paper cites Geopix: A multimodal large language model for pixel-level image understanding in remote sensing.IEEE Geoscience and Remote Sensing Magazine, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Geopix: A multimodal large language model for pixel-level image understanding in remote sensing.IEEE Geoscience and Remote Sensing Magazine, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.501383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.501383Z digest=sha256:2aead24a7cb78db04bb6366ab9b25fc8b127561310556378487930c9949558bb

Observation 65940e87-c7ca-4c12-b52c-4d4054028358 · outbound

This paper cites Vhm: Versatile and honest vision language model for remote sensing image analysis.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Vhm: Versatile and honest vision language model for remote sensing image analysis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.552720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.552720Z digest=sha256:16cf626eee8ad45f49bfb77055b8ed17a11ece9b3fe0414b8bde94b68eb3ffb5

Observation 9ddf7593-9019-48ce-b84c-fa5f2f6664a9 · outbound

This paper cites Llava++: extending visual capabilities with llama-3 and phi-3 (2024).URL https://github.com/mbzuai-oryx/LLaVA- pp, 2024.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Llava++: extending visual capabilities with llama-3 and phi-3 (2024).URL https://github.com/mbzuai-oryx/LLaVA- pp, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.593053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.593053Z digest=sha256:35341bf2dc54b565c23c1ef923a0f3b92f0f684269eb7f8d13b9d6386eb714d0

Observation 086ae9bd-1812-4e22-acf0-874e4e16621b · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Glamm: Pixel grounding large multimodal model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.644608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.644608Z digest=sha256:f1c68422bb0ad610b538363faf33c064c4b95611e234eafba8071f930fe42912

Observation fb476229-664f-4803-89e2-2fd277f44817 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SAM 2: Segment Anything in Images and Videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.721970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.721970Z digest=sha256:6fbc6b0faedec65cb627eeff2f6449c7c3b21e792d9aeb1378dc36197bdef657

Observation eec54fbe-cb03-45c0-b031-d8fcb329e4d4 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Pixellm: Pixel reasoning with large multimodal model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.803839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.803839Z digest=sha256:9938cd074ab68ba127fda4eac7af6582aa36df04e0b9301f5d8ec7110014f8ef

Observation 867c4e28-917a-4a41-93be-31ea3459c95a · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Moviechat: From dense token to sparse memory for long video understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.860093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.860093Z digest=sha256:1e13ae9baf1bb6358e406eec1f7a3883ad3689f1bb6eece26520a7ac133ba822

Observation d83d7eb4-6c63-4448-94a0-952dadbdb274 · outbound

This paper cites Earthdial: Turning multi-sensory earth observations to interactive dialogues.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Earthdial: Turning multi-sensory earth observations to interactive dialogues

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.915593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.915593Z digest=sha256:27e3880ca21948d6ac17117186a74de689791f4d57b9564379c027397c45fd77

Observation 52daab6d-b37e-48b3-9d2e-d8222f199e69 · outbound

This paper cites Drone-based rgb-infrared cross-modality vehicle detection via uncertainty-aware learning.IEEE Transactions on Circuits and Systems for Video Technology, 32(10):6700–6713, 2022.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Drone-based rgb-infrared cross-modality vehicle detection via uncertainty-aware learning.IEEE Transactions on Circuits and Systems for Video Technology, 32(10):6700–6713, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.973608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.973608Z digest=sha256:901c84e452f68831e99d0e7985f09c882f533f84d4e8d5a5fc9376485f75fb49

Observation 17863fb7-2119-4768-8b96-53666e87d913 · outbound

This paper cites Adaptive keyframe sampling for long video understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Adaptive keyframe sampling for long video understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.030943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.030943Z digest=sha256:f50993b657f73189be13ae1af207aef0d83f15e5ad3ca4a875f4b3c52288451c

Observation 4e899a43-c9ad-4aeb-9ac4-0626688eb142 · outbound

This paper cites Cider: Consensus-based image description evaluation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Cider: Consensus-based image description evaluation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.074265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.074265Z digest=sha256:c8f1b8275c7a3dce6fa170a7cf057b1de11886d242283559563f90bb25e34d14

Observation a24a7245-aef3-4bb0-aa6b-10394c3e1801 · outbound

This paper cites Geollava-8k: scaling remote-sensing multimodal large language models to 8k resolution.arXiv preprint arXiv:2505.21375, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Geollava-8k: scaling remote-sensing multimodal large language models to 8k resolution.arXiv preprint arXiv:2505.21375, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.124523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.124523Z digest=sha256:4707ce78bd58656f654d7b45f2f492c56d1f10b5bc4b2e8a6f53490d6c822ecf

Observation 2dd4ab56-621a-491c-bbf2-07dda40d7475 · outbound

This paper cites Instructseg: Unifying instructed visual segmentation with multi-modal large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Instructseg: Unifying instructed visual segmentation with multi-modal large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.221579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.221579Z digest=sha256:2606ea7e440cbf024f1e9cbe11db8cc1de4928159a9dfbe1e4b93b34b9bc89f7

Observation aa94ce04-88b5-493b-9d0e-b29f1600024d · outbound

This paper cites Longvlm: Efficient long video understanding via large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Longvlm: Efficient long video understanding via large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.350493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.350493Z digest=sha256:c46c1c921b2a42a71881a49764562c59e4d72edab366fd75a46cd2af77819b72

Observation 1a711df3-dda3-4d74-9bfd-c7b3a4daab95 · outbound

This paper cites Language as queries for referring video object segmentation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Language as queries for referring video object segmentation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.422889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.422889Z digest=sha256:f13d92799127a48b3a00f9ab1bd73b03f16f65d045273f1fe71a87bf578eb9cc

Observation 9ce06e3c-e938-4b4b-ab75-acf4c575595c · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.486399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.486399Z digest=sha256:e087fa1bdbcff965f72900e9b345ba904bc8a05fd25c7524aefdff2ce97bb420

Observation 32849a1b-d5f7-4a20-b197-fc29806e3be2 · outbound

This paper cites Visa: Reasoning video object segmentation via large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Visa: Reasoning video object segmentation via large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.576528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.576528Z digest=sha256:ffc628d60e903b264dbdae9aabe076aad5b252fdb547d80ef83d234735408223

Observation 126070c8-ffb6-4120-bcfd-8e74f9716739 · outbound

This paper cites Referred by multi-modality: A unified temporal transformer for video object segmentation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Referred by multi-modality: A unified temporal transformer for video object segmentation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.579616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.579616Z digest=sha256:767ecaa4728591bd59c5ad24fd9a79a5b107668f5dcda2ceb3a5a0d5f5f0dc55

Observation f4431f6f-4b71-442c-8f56-39b9b7774506 · outbound

This paper cites Self-chained image-language model for video localization and question answering.Advances in Neural Information Processing Systems, 36:76749–76771, 2023.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Self-chained image-language model for video localization and question answering.Advances in Neural Information Processing Systems, 36:76749–76771, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.612401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.612401Z digest=sha256:60c0695eaa48c1b702eeedc1bef8979f422bdaf8221ddf2969d005e1cb79e3d9

Observation 4a2ee8db-e65b-4904-8f56-7f723e0d407d · outbound

This paper cites Frame-Voyager: Learning to Query Frames for Video Large Language Models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Frame-Voyager: Learning to Query Frames for Video Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.661242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.661242Z digest=sha256:82b6ac08adaa1b48efec44d013cf72d20f11cbf7fef3215933d6f27ce97fffe5

Observation 4817844c-a271-4236-a712-ef3db908dedc · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.743258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.743258Z digest=sha256:09b24224b0851c2a409486ffab82f393443d72a3c0cecd142ea01dbaa2cce051

Observation bb233953-67a4-4e45-9c3c-ed97252e8ef0 · outbound

This paper cites Skyeyegpt: Unifying remote sensing vision- language tasks via instruction tuning with large language model.ISPRS Journal of Photogram- metry and Remote Sensing, 221:64–77, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Skyeyegpt: Unifying remote sensing vision- language tasks via instruction tuning with large language model.ISPRS Journal of Photogram- metry and Remote Sensing, 221:64–77, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.814064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.814064Z digest=sha256:2f47da7d7332b4792066dfe95777cfbb1a00f39849db619c91854fe693211015

Observation bd172da8-7b7c-4156-9dcf-d6053158fda9 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.885115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.885115Z digest=sha256:9387a0598bd839e7263e2905df6096703d868cb44fef5918cfc0061f7751b837

Observation 0cd98b9a-d26e-4ea7-90ad-819a85d924ee · outbound

This paper cites GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.944636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.944636Z digest=sha256:2064f856a422206ee6a8a33093aa0d9a2ac2113739ebac1b2c333dad35fd7837

Observation a1b883f2-6312-4707-a5a4-da08e0549566 · outbound

This paper cites Tifre: Text-guided video frame reduction for efficient video multi-modal large language models.arXiv preprint arXiv:2602.08861, 2026.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Tifre: Text-guided video frame reduction for efficient video multi-modal large language models.arXiv preprint arXiv:2602.08861, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:04.007847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:04.007847Z digest=sha256:327280d8b963621cfe57cadeabf5d79e92756672d6e7023d25bfe330d6483014

Observation 9860a9dd-5836-408c-b37e-fe9a0a0d5ded · outbound

This paper cites Reason: Reinforced causal search with information bottleneck for video understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Reason: Reinforced causal search with information bottleneck for video understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:04.079355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:04.079355Z digest=sha256:b1642ae470b8884ac1c1e271b026f2861d5bfdde0998eb9895de98c27964d409

Observation 890e5b91-6ce5-49a7-9e45-11915eb46598 · outbound

This paper cites Detection and tracking meet drones challenge.IEEE transactions on pattern analysis and machine intelligence, 44(11):7380–7399, 2021.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Detection and tracking meet drones challenge.IEEE transactions on pattern analysis and machine intelligence, 44(11):7380–7399, 2021

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:04.130946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:04.130946Z digest=sha256:ed2c84cc8c197c1db60b2bcfddf9db147c72aeda622fc0b2c77f754f28d246c5

Observation 42c5d8b8-44ec-4295-ac63-564a1914705d · outbound

This paper cites Focus: Efficient keyframe selection for long video understanding.arXiv preprint arXiv:2510.27280, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Focus: Efficient keyframe selection for long video understanding.arXiv preprint arXiv:2510.27280, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:04.192167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:04.192167Z digest=sha256:8ee9bb09c35073d956828174240297c03c01dc565b72c1beb46e3c3f19c360f7

Pith citing papers

No inbound Pith citation observations are available.