Pith. sign in

Paper Citation Record · LEDGER

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing

As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2607.17386.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17386 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:11:04.192167Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c911503-5255-4be9-9c32-4a06b43b0c4a · outbound

This paper cites One token to seg them all: Language instructed reasoning seg- mentation in videos.Advances in Neural Information Processing Systems, 37:6833–6859, 2024.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing One token to seg them all: Language instructed reasoning seg- mentation in videos.Advances in Neural Information Processing Systems, 37:6833–6859, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.691639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.691639Z digest=sha256:559dfba864a89dc6e09c104b41f49521928ef2dfe7579a25ba2e6702d6354614

Observation e726ec18-5a18-4e01-a66d-acf19c22c573 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.744246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.744246Z digest=sha256:cc0b7239d66555ee4052c679ac90be165bffb2c55f6fdbd3d1f1a364d6e2b8d3

Observation bc7c76ca-1343-4e91-9975-31bbadbdd027 · outbound

This paper cites End-to-end referring video object segmentation with multimodal transformers.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing End-to-end referring video object segmentation with multimodal transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.790110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.790110Z digest=sha256:778a4210b1eb33453244f781586a78313c5a964df0fded66073020bda72082ee

Observation ebeeb569-da85-4acb-8d79-5b96c294852e · outbound

This paper cites Streamingtom: Streaming token compres- sion for efficient video understanding.arXiv preprint arXiv:2510.18269, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Streamingtom: Streaming token compres- sion for efficient video understanding.arXiv preprint arXiv:2510.18269, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.854409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.854409Z digest=sha256:ec8c4a3d0caa88b780f9c1c154fa8ca0d35c446c5bbde7dd8df0d665d4a2c227

Observation aac304ac-f50d-4d06-89f3-224d7201ef3f · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.919449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.919449Z digest=sha256:0f3eb55d91759fb25127aec1d891e8f8e4b03b9dd8f16ef5198b533759fe756c

Observation f0495812-3451-4c19-85ed-ad679685c4c3 · outbound

This paper cites The unmanned aerial vehicle benchmark: Object detection and tracking.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing The unmanned aerial vehicle benchmark: Object detection and tracking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:00.984738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:00.984738Z digest=sha256:7d9ca874c9783a3c1b40eff27bad466f7fa03f7c5ce2c4c06d0816ac36e7b663

Observation f6ace37c-9484-43c0-8418-585971a9f62c · outbound

This paper cites Framefusion: Combining similarity and importance for video token reduction on large vision language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Framefusion: Combining similarity and importance for video token reduction on large vision language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.082341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.082341Z digest=sha256:3b1a8eaba8422b74c7ab85de42167b288b82f9ed771f83961ed2a420d92a8c6f

Observation e7486384-4444-47a1-b64c-6ab04edcd40b · outbound

This paper cites Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.204751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.204751Z digest=sha256:a646a9115397c1991ba8e333d37edc89844d827f835844ca32fbecc5fa514617

Observation 954877f9-8f77-44ff-bdd2-48c34a27fa1d · outbound

This paper cites The devil is in temporal token: High quality video reasoning segmentation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing The devil is in temporal token: High quality video reasoning segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.296336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.296336Z digest=sha256:ea60008ea440e19654d74b81463511e4a32ae78db1693189659c69e8ec0ca980

Observation c2dae05a-a68c-4d7b-9543-b872f27f70de · outbound

This paper cites Rsgpt: A remote sensing vision language model and benchmark.ISPRS Journal of Photogrammetry and Remote Sensing, 224:272–286, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Rsgpt: A remote sensing vision language model and benchmark.ISPRS Journal of Photogrammetry and Remote Sensing, 224:272–286, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.367418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.367418Z digest=sha256:4a80b92abfddc465dee837159b65a95d068298a966dae7fc536379aca6852c1f

Observation 69ff8e28-3474-4e55-8196-1d2bf37d04b2 · outbound

This paper cites Prunevid: Visual token pruning for efficient video large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Prunevid: Visual token pruning for efficient video large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.457054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.457054Z digest=sha256:b049f73118ca8dcd91c569b012fdf3e020b3482b595bd264c7a0c70a45298939

Observation dfecb740-b81a-49cf-a800-cfeac3bca07f · outbound

This paper cites Multi-granular spatio-temporal token merging for training-free acceleration of video llms.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Multi-granular spatio-temporal token merging for training-free acceleration of video llms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.521166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.521166Z digest=sha256:23c4592fe765ea045d7f02c27557aa47009a41a87fbc2a9fb6123f54506686af

Observation c57608d3-d131-4995-8b1c-eff11246edac · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.580630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.580630Z digest=sha256:6d694265ff7f00bfb973372478119bed1fd4dd21dfa0d00d3b2efc276451e0ad

Observation a485416f-d454-4b76-9ef5-7fc4ede0b8cf · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Lisa: Reasoning segmentation via large language model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.696273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.696273Z digest=sha256:05e8b38938a2055b4fd1b1b53b2680dd150a9434c3d09a6d6b9d725fa059c410

Observation 83cd5b5c-dd0b-47d7-a5e5-ee9ef889736f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.761048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.761048Z digest=sha256:ad16651272b8d159801fe4db5824ccd58d8f1a6a9fd5bdbfc0bd1a2e1c82f692

Observation 6731052e-dbfa-41ad-814f-7b040ed21374 · outbound

This paper cites Referdino: Referring video object segmentation with visual grounding foundations.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Referdino: Referring video object segmentation with visual grounding foundations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.849069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.849069Z digest=sha256:693e51ad49e0fe92dfff56d63cf4dd7a4026d674b2a49b04d83bed055dc085cc

Observation 4f45cead-57fa-4193-934d-3aad75ef807a · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Video-llava: Learning united visual representation by alignment before projection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:01.971650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:01.971650Z digest=sha256:8b8033ed5140d2a306e37002969400c0d1bc1149443d2a5c5638dcf8715386d7

Observation 27c684a3-7358-48c6-886a-e007211e6294 · outbound

This paper cites Glus: Global-local reasoning unified into a single large language model for video segmentation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Glus: Global-local reasoning unified into a single large language model for video segmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.034375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.034375Z digest=sha256:151306b2378e4228e842f795555a78f0bbe76583a4fded9527687ae4ffe2c0a2

Observation 1e1e3f91-2614-4465-aaeb-bde5dc3df7a4 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.108301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.108301Z digest=sha256:31b9822a64bd95e17b381ac7c3239bfae7b71df25e7d9a7dbb23a44c7d239e54

Observation bf66e690-c950-4d14-b84c-12030a074936 · outbound

This paper cites RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.265647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.265647Z digest=sha256:2d56de3070eb7401949c60a755a41426eae9075cb49a246a8b2b92062078e983

Observation 7da24c0e-c78c-49c0-915a-672bacdb4bf6 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.304478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.304478Z digest=sha256:761bb5b052cd06c661b616c510c45ab168e120747f3018e576cc6712f4494e8c

Observation 0ebe0656-7020-49d2-88e4-e00b58c04323 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.369009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.369009Z digest=sha256:a39b144f8c44b1cf8ef49468e5a5b2bdd1c23235e9acc9d9b705da4159b20dfb

Observation b7b0d143-8b00-4c13-8bfd-e131ad13230a · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual grounding in videos.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Videoglamm: A large multimodal model for pixel-level visual grounding in videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.428042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.428042Z digest=sha256:44bdf3dd48bdb85ef8dda3bd4df728e14bedcf59e66d831ff288d3644df32af2

Observation 113d5682-d67d-40cb-b88c-a2e3a48dad5a · outbound

This paper cites Geopix: A multimodal large language model for pixel-level image understanding in remote sensing.IEEE Geoscience and Remote Sensing Magazine, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Geopix: A multimodal large language model for pixel-level image understanding in remote sensing.IEEE Geoscience and Remote Sensing Magazine, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.501383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.501383Z digest=sha256:6b2f482027b56e03a9bee7f3c2392f7b543e15148215161ba79fcdd8cb2d1a46

Observation 65940e87-c7ca-4c12-b52c-4d4054028358 · outbound

This paper cites Vhm: Versatile and honest vision language model for remote sensing image analysis.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Vhm: Versatile and honest vision language model for remote sensing image analysis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.552720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.552720Z digest=sha256:4310dade618d80c9f93a260d3dba570b2cb862606750d1b1b3b01a2f50280c82

Observation 9ddf7593-9019-48ce-b84c-fa5f2f6664a9 · outbound

This paper cites Llava++: extending visual capabilities with llama-3 and phi-3 (2024).URL https://github.com/mbzuai-oryx/LLaVA- pp, 2024.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Llava++: extending visual capabilities with llama-3 and phi-3 (2024).URL https://github.com/mbzuai-oryx/LLaVA- pp, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.593053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.593053Z digest=sha256:a7190c4bec6c5df9a87ee70db9070213b667dcab2961cbde5df285a5727372a5

Observation 086ae9bd-1812-4e22-acf0-874e4e16621b · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Glamm: Pixel grounding large multimodal model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.644608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.644608Z digest=sha256:f9ff9c6a4d9e4a2513ed0ca32bf22d9f68647b64c961122e26ef7f7601681a6a

Observation fb476229-664f-4803-89e2-2fd277f44817 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SAM 2: Segment Anything in Images and Videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.721970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.721970Z digest=sha256:cf597c4ce10822ac6c936db19f28aa369200495bcb9e6efbe859fd808e1558c0

Observation eec54fbe-cb03-45c0-b031-d8fcb329e4d4 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Pixellm: Pixel reasoning with large multimodal model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.803839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.803839Z digest=sha256:b3ed360f47678a858f8c356de9acb21795253a9eaad86e2ec2e10472460e3478

Observation 867c4e28-917a-4a41-93be-31ea3459c95a · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Moviechat: From dense token to sparse memory for long video understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.860093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.860093Z digest=sha256:6591295478b26ddfc697652e5f389fdedcd9bb3c6b0b44d3f484919f2c4b75ff

Observation d83d7eb4-6c63-4448-94a0-952dadbdb274 · outbound

This paper cites Earthdial: Turning multi-sensory earth observations to interactive dialogues.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Earthdial: Turning multi-sensory earth observations to interactive dialogues

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.915593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.915593Z digest=sha256:319cc80866720cca85a743e40b7f6b9d1d178bdcaf399d0364e2c86961e677c1

Observation 52daab6d-b37e-48b3-9d2e-d8222f199e69 · outbound

This paper cites Drone-based rgb-infrared cross-modality vehicle detection via uncertainty-aware learning.IEEE Transactions on Circuits and Systems for Video Technology, 32(10):6700–6713, 2022.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Drone-based rgb-infrared cross-modality vehicle detection via uncertainty-aware learning.IEEE Transactions on Circuits and Systems for Video Technology, 32(10):6700–6713, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:02.973608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:02.973608Z digest=sha256:2e3961e024661a31540a358e3660d536c4096e08b93dd611a7815b2b78613b3f

Observation 17863fb7-2119-4768-8b96-53666e87d913 · outbound

This paper cites Adaptive keyframe sampling for long video understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Adaptive keyframe sampling for long video understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.030943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.030943Z digest=sha256:1e1b03031879c06e2a6fd85329bc694be270b2e9ee973fccbf565fd31ab06e9e

Observation 4e899a43-c9ad-4aeb-9ac4-0626688eb142 · outbound

This paper cites Cider: Consensus-based image description evaluation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Cider: Consensus-based image description evaluation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.074265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.074265Z digest=sha256:6f08436c71baa93f3da8dbd76aedf18d7b56a346fdda23b3605d27c5eb1b862f

Observation a24a7245-aef3-4bb0-aa6b-10394c3e1801 · outbound

This paper cites Geollava-8k: scaling remote-sensing multimodal large language models to 8k resolution.arXiv preprint arXiv:2505.21375, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Geollava-8k: scaling remote-sensing multimodal large language models to 8k resolution.arXiv preprint arXiv:2505.21375, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.124523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.124523Z digest=sha256:24876ac0e76f201f923ece73e159e24cec924fc90d97638721e4a740f8070655

Observation 2dd4ab56-621a-491c-bbf2-07dda40d7475 · outbound

This paper cites Instructseg: Unifying instructed visual segmentation with multi-modal large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Instructseg: Unifying instructed visual segmentation with multi-modal large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.221579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.221579Z digest=sha256:d840692f02ae2bcefa16038f47744ec8192108d821717d206083a0add9d512f3

Observation aa94ce04-88b5-493b-9d0e-b29f1600024d · outbound

This paper cites Longvlm: Efficient long video understanding via large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Longvlm: Efficient long video understanding via large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.350493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.350493Z digest=sha256:6eacfd0e43f9f6d0ac8d8a5c0360935e2a27708dd3e40148b8641bb60a446ffa

Observation 1a711df3-dda3-4d74-9bfd-c7b3a4daab95 · outbound

This paper cites Language as queries for referring video object segmentation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Language as queries for referring video object segmentation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.422889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.422889Z digest=sha256:c216041cc66f006a3ad98f86ceb6885dd96dfa0f588f7da747983fcd36132bc4

Observation 9ce06e3c-e938-4b4b-ab75-acf4c575595c · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.486399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.486399Z digest=sha256:0ae4f1e1afa5bd0337dce25ba6bd2a6955c78120fdba1ee13384390a8c1cd671

Observation 32849a1b-d5f7-4a20-b197-fc29806e3be2 · outbound

This paper cites Visa: Reasoning video object segmentation via large language models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Visa: Reasoning video object segmentation via large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.576528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.576528Z digest=sha256:c05a7e591359a09978079b6cf7fbf6f975c94317c9abda5a800d6027dfa9efe9

Observation 126070c8-ffb6-4120-bcfd-8e74f9716739 · outbound

This paper cites Referred by multi-modality: A unified temporal transformer for video object segmentation.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Referred by multi-modality: A unified temporal transformer for video object segmentation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.579616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.579616Z digest=sha256:6b849a13abe5aa9c8dd54021aab6448045e03446f090cadce86dc601cc5532d7

Observation f4431f6f-4b71-442c-8f56-39b9b7774506 · outbound

This paper cites Self-chained image-language model for video localization and question answering.Advances in Neural Information Processing Systems, 36:76749–76771, 2023.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Self-chained image-language model for video localization and question answering.Advances in Neural Information Processing Systems, 36:76749–76771, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.612401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.612401Z digest=sha256:8fb27872574138f140c818e01c1aa8b8900983324a8733150eb686790a90b4a4

Observation 4a2ee8db-e65b-4904-8f56-7f723e0d407d · outbound

This paper cites Frame-Voyager: Learning to Query Frames for Video Large Language Models.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Frame-Voyager: Learning to Query Frames for Video Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.661242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.661242Z digest=sha256:4689590de02132076597ae670b474b0772016a328ab1e7a4aa58a8d9c00ab9de

Observation 4817844c-a271-4236-a712-ef3db908dedc · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.743258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.743258Z digest=sha256:b26d03a3091e1d086cc9c39eaf8d91dbacd359c9bcf4275a8a429a9420e91567

Observation bb233953-67a4-4e45-9c3c-ed97252e8ef0 · outbound

This paper cites Skyeyegpt: Unifying remote sensing vision- language tasks via instruction tuning with large language model.ISPRS Journal of Photogram- metry and Remote Sensing, 221:64–77, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Skyeyegpt: Unifying remote sensing vision- language tasks via instruction tuning with large language model.ISPRS Journal of Photogram- metry and Remote Sensing, 221:64–77, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.814064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.814064Z digest=sha256:b594a192da69bf145d9c6d70ba5b6458acb5769e218bf5d951138a8bac450628

Observation bd172da8-7b7c-4156-9dcf-d6053158fda9 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.885115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.885115Z digest=sha256:0a7fd5853203d519958259eff1cab1332e8a5371860b46c0083b9b5dc5def3bf

Observation 0cd98b9a-d26e-4ea7-90ad-819a85d924ee · outbound

This paper cites GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:03.944636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:03.944636Z digest=sha256:5b2a8a2b7ad62022424e29ed7642623d29ad1d6fa3caf04ec630fd6a82c8fe55

Observation a1b883f2-6312-4707-a5a4-da08e0549566 · outbound

This paper cites Tifre: Text-guided video frame reduction for efficient video multi-modal large language models.arXiv preprint arXiv:2602.08861, 2026.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Tifre: Text-guided video frame reduction for efficient video multi-modal large language models.arXiv preprint arXiv:2602.08861, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:04.007847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:04.007847Z digest=sha256:6356fd78eb8bb388293765fad2e64729040c9fef3c2636ca62f2af51cc40a18d

Observation 9860a9dd-5836-408c-b37e-fe9a0a0d5ded · outbound

This paper cites Reason: Reinforced causal search with information bottleneck for video understanding.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Reason: Reinforced causal search with information bottleneck for video understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:04.079355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:04.079355Z digest=sha256:db42f9fe6feba6d864cb01767fea9956f51cf59ac61d1f16eb5938bd5ebeb23c

Observation 890e5b91-6ce5-49a7-9e45-11915eb46598 · outbound

This paper cites Detection and tracking meet drones challenge.IEEE transactions on pattern analysis and machine intelligence, 44(11):7380–7399, 2021.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Detection and tracking meet drones challenge.IEEE transactions on pattern analysis and machine intelligence, 44(11):7380–7399, 2021

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:04.130946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:04.130946Z digest=sha256:860940f70a73bda3dd76218699b7f28c7ee14c9c8c843318062bb6f05f18e8ad

Observation 42c5d8b8-44ec-4295-ac63-564a1914705d · outbound

This paper cites Focus: Efficient keyframe selection for long video understanding.arXiv preprint arXiv:2510.27280, 2025.

SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing Focus: Efficient keyframe selection for long video understanding.arXiv preprint arXiv:2510.27280, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T18:11:04.192167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:11:04.192167Z digest=sha256:b0510b8f98ddc3cbbb4cde75dfcd6ac657f6a7853f547c864406c10fcaf1a819

Pith citing papers

No inbound Pith citation observations are available.