Pith. sign in

Paper Citation Record · LEDGER

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos

As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.19857.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19857 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:33:44.814543Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b995d4e-2191-4b15-b83f-b8fc2466296d · outbound

This paper cites The small-drone revolution is coming—scientists need to ensure it will be safe,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos The small-drone revolution is coming—scientists need to ensure it will be safe,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.010000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.010000Z digest=sha256:1099cbc9039b8b8df7420812590519f3edd8646166a5d9ea79d0ce49bffc2a02

Observation 49b07a8c-e707-4b7c-9a55-2c743b824f6b · outbound

This paper cites Champion-level drone racing using deep reinforcement learning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Champion-level drone racing using deep reinforcement learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.077507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.077507Z digest=sha256:91bfe6252583d8c1e17289d095e69291e1e3937c4f3080f722c49788cdcd1c4f

Observation 761e62ab-1b1c-4192-b68f-2dad832a8038 · outbound

This paper cites Video object segmentation without tem- poral information,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Video object segmentation without tem- poral information,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.210959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.210959Z digest=sha256:3976e61e6144f5f5483a61d0e58202e38aa32a551b1fbd4b6e5c1c01137d177d

Observation 5c39917d-b99c-406e-b8f3-61d24b26056c · outbound

This paper cites 3d question answering for city scene understanding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos 3d question answering for city scene understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.290989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.290989Z digest=sha256:af9e3178819558bcde57019a66d9b3703e1a547b81881bb5b9645a00fb7752db

Observation 43efa3e9-286a-4e1b-be86-a3197c84754d · outbound

This paper cites Phenobench: A large dataset and benchmarks for semantic image interpretation in the agricultural do- main,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Phenobench: A large dataset and benchmarks for semantic image interpretation in the agricultural do- main,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.391352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.391352Z digest=sha256:7086493d84ed07b2c0cb3f007d727c3b9cff9bee3966a79c0a14f70f8b421e9d

Observation 01ecd4ed-e87b-4b06-8dec-7ad48ad22ef2 · outbound

This paper cites Detecting flying objects using a single moving camera,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Detecting flying objects using a single moving camera,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.479024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.479024Z digest=sha256:e4d185a1b423c96b7498d9eacaf8bc6af0b6b8caed25592d090f62e6d3209fde

Observation 4f4d291c-5418-4571-8780-b59da9a6df49 · outbound

This paper cites Revisiting image-language networks for open-ended phrase detection,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Revisiting image-language networks for open-ended phrase detection,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.574070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.574070Z digest=sha256:4f5d0cf4e0171c1f81854cfc70786644530a14d35055da7067b33e6c30852c3b

Observation d2b93844-3363-4b7a-aff3-0037efbdcd01 · outbound

This paper cites Mevis: A large- scale benchmark for video segmentation with motion expressions,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Mevis: A large- scale benchmark for video segmentation with motion expressions,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.647764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.647764Z digest=sha256:fc27af091c3001a66019a473db42b7611b3db80507f6c982c764a1fbc2d9f2d3

Observation d5431214-f946-448e-974c-5886b8e16458 · outbound

This paper cites Lamot: Language- guided multi-object tracking,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Lamot: Language- guided multi-object tracking,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.750581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.750581Z digest=sha256:4fdda37488341c05c86ade5de2239f4c42e623d1c0e7f78b79311ab5e75cc1fa

Observation a6822768-0774-409a-b3c4-642633a0a9f6 · outbound

This paper cites Skyfind: A large-scale benchmark unveiling referring expres- sion comprehension for uav.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Skyfind: A large-scale benchmark unveiling referring expres- sion comprehension for uav

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.793599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.793599Z digest=sha256:c74cecff457fe8f3aebcd1849c5214e163aacca5ff926a2a2db707e716f2ebbf

Observation 50c4b8f7-f820-4c1e-96ac-d23aae4cef2d · outbound

This paper cites Aerialmind: Towards referring multi-object tracking in uav scenarios,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Aerialmind: Towards referring multi-object tracking in uav scenarios,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.860865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.860865Z digest=sha256:71b371bba494c4be739740f29464123090ff6850394879bd4cb531402e41dc4e

Observation 6da0cf3b-dd9b-4492-acd3-8afc41de44ce · outbound

This paper cites Event-aware instructed assistant for referring video segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Event-aware instructed assistant for referring video segmentation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:41.918717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:41.918717Z digest=sha256:e2aba73ffc21a63412028c5104d49fa66d680c8548c4085cbac4b29efd41514d

Observation cd07ae9b-fc66-4d12-97e5-44f84799e2fb · outbound

This paper cites City-vlm: Towards multidomain perception scene understanding via multimodal incomplete learning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos City-vlm: Towards multidomain perception scene understanding via multimodal incomplete learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.000557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.000557Z digest=sha256:ea13ea8e2eecdf8ce4baca9c4fb7b571a3a7839d36a0f62626c9a9948e5b0437

Observation 71600a21-e282-4709-8a33-42faab77c5fb · outbound

This paper cites Qwen2.5 technical report,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5 technical report,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.085352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.085352Z digest=sha256:674438224a7b253e6f778d8d33a8e5902134c69c078e9317f290c5de6472497f

Observation 366e8549-1cbc-4afa-984b-b330d1f213b2 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SAM 3: Segment Anything with Concepts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.259681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.259681Z digest=sha256:0500ce11066ca87c209d1b034f15a6bb93899aa2b09b88e29a27db4db16ceb09

Observation 305d387f-576c-47d0-99f8-484771a6ffed · outbound

This paper cites React: Streaming video analytics on the edge with asynchronous cloud support,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos React: Streaming video analytics on the edge with asynchronous cloud support,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.350558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.350558Z digest=sha256:6c2619bbd22c1154b886756692b25e1797f0ced35e69f518cb5da1e98f9dc0bf

Observation 5b29eab7-6a4d-4271-8344-f237d491cea0 · outbound

This paper cites Multiple-object-tracking algorithm based on dense trajectory voting in aerial videos,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Multiple-object-tracking algorithm based on dense trajectory voting in aerial videos,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.439765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.439765Z digest=sha256:b0c7f550105f3b67e79967f45b34ab7dbda09df13ed7334fb2cff257206bd8ac

Observation 08fa439c-bdb6-4acf-8ed9-9f15f82ab321 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Urvos: Unified referring video object segmentation network with a large-scale benchmark,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.515394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.515394Z digest=sha256:d5a1ab448ed5e50a9476c0e65cc3b1c53bd47e993619cc686d5f7d35f734a0ab

Observation b14fc791-78fa-4110-bcd4-66ae0c8a6cf8 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Sharegpt4video: Improving video understanding and generation with better captions,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.604426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.604426Z digest=sha256:f671638ecac5f09e0763eaaebc53588092207a1a5d90c81f72c7f891fab3270d

Observation 79b21db3-821c-4694-8687-61f39cbdae15 · outbound

This paper cites Mevis: A multi-modal dataset for referring motion expression video segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Mevis: A multi-modal dataset for referring motion expression video segmentation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.675678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.675678Z digest=sha256:accd53cc487ae3ddd885cd4cd380e128d4475fd45115f9685124b35039550340

Observation 844c35ca-a119-472a-8294-da0f114f5e6f · outbound

This paper cites Jtd-uav: Mllm-enhanced joint tracking and description framework for anti-uav systems,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Jtd-uav: Mllm-enhanced joint tracking and description framework for anti-uav systems,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.778744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.778744Z digest=sha256:3c66519215b8ae52fd0fa8d809c0faed93afdf539d9afe2343a2c48513a668f1

Observation 0e892caa-5ce7-4f56-bdb7-f50fa6c94cd4 · outbound

This paper cites Visa: Reasoning video object segmentation via large lan- guage models,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Visa: Reasoning video object segmentation via large lan- guage models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.829375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.829375Z digest=sha256:4504fe52f7e82174a93125546712985b87c8550b8c08229c89c442396ce6d9ba

Observation eed37ea5-009d-4755-b182-9cc0b34b6b93 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.889990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.889990Z digest=sha256:8117237f866380de06c7528f516246c9d4920dfe530a9c713a905872e51aa27e

Observation 26c53c5e-2219-4323-a01b-152283af7035 · outbound

This paper cites Unipixel: Unified object referring and segmentation for pixel-level visual reason- ing,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Unipixel: Unified object referring and segmentation for pixel-level visual reason- ing,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.985730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.985730Z digest=sha256:3ea230685e3074477e466741008a0f74d1e23760eda51f61746a333003243656

Observation 166f5840-37de-4732-b725-fca92a964b90 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SAM 2: Segment Anything in Images and Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.065979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.065979Z digest=sha256:9769da756390920db83c6638e6f0a899c99607bf8f6a052c465519e1d1f12b87

Observation 4af657c0-610a-4099-a718-52299e96ebc4 · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14 grounding in videos,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Videoglamm: A large multimodal model for pixel-level visual JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14 grounding in videos,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.135856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.135856Z digest=sha256:4a338f608963d5092124409999902500f710ec2495ca43570771bab7435f9f61

Observation 7a0db379-3b2a-4df5-b8ca-796f9e2f9a06 · outbound

This paper cites Glus: Global-local reasoning unified into a single large language model for video segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Glus: Global-local reasoning unified into a single large language model for video segmentation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.233507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.233507Z digest=sha256:6bae2a2d74a09f9a0da33a601fa1cac8473fdb72e192e0aac8d8bcac43a1246a

Observation 5c1cf94f-b519-42a1-b153-da195082896f · outbound

This paper cites The devil is in temporal token: High quality video reasoning segmentation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos The devil is in temporal token: High quality video reasoning segmentation,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.328278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.328278Z digest=sha256:1d50dcb71894d7b7562e780b4d33e03147f5864ed634fb9a3515430b74ec25c4

Observation 1fc75169-9d8b-4957-8153-a284702ec0b5 · outbound

This paper cites Instructseg: Unifying instructed visual segmentation with multi-modal large language models,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Instructseg: Unifying instructed visual segmentation with multi-modal large language models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.435658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.435658Z digest=sha256:7a4945f9da64782f91375fe309b21ecb4b0674943b9014be695ab48a644c661f

Observation df21c549-3a59-43ec-94da-f40429c71d43 · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Geochat: Grounded large vision-language model for remote sensing,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.532666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.532666Z digest=sha256:ea94bd8bd50251919963daea4e4ebb2241a744974ef5c2b70a6fd7c51b683724

Observation aa2add77-efb9-4878-be89-fdcd1b629294 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.613930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.613930Z digest=sha256:d2287900b4385ce723536ba31e89063447ed4aca31e2fddea78aabb2ffd05866

Observation 2ef00eab-20ef-407f-b9a5-a1423abd28ee · outbound

This paper cites Earthgpt: A universal multimodal large language model for multisensor image comprehension in remote sensing domain,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Earthgpt: A universal multimodal large language model for multisensor image comprehension in remote sensing domain,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.690981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.690981Z digest=sha256:8450da1688c02494cf384d81b894f5ff989c6306e272af39b9e6af3558ff27b1

Observation 1b7777f8-6659-4ffb-9d4d-397c763d51cd · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Kimi K2.5: Visual Agentic Intelligence

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.758425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.758425Z digest=sha256:13a668185d6dc121bbc6e4788208c5c437502f6c42729dd71ebebb44bd7e360e

Observation beedfb6f-dfe7-4f2b-9f22-bf52ce0350a9 · outbound

This paper cites Qwen3.6-Plus: Towards real world agents,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen3.6-Plus: Towards real world agents,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.831961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.831961Z digest=sha256:237717041e419c0603f4096e05c86a95cd197d682799ded4c513c59d17e154d6

Observation 211f0754-f661-4860-b3d8-3eb118d94675 · outbound

This paper cites Visual instruction tuning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Visual instruction tuning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.921225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.921225Z digest=sha256:e1d4d6912a65206d16b129b2f8b4daf299cd18858075debdf098309c7bb78aa3

Observation 49e040f5-7cb0-4289-9175-9dbffc0da9c2 · outbound

This paper cites Qwen2.5-VL Technical Report.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5-VL Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:43.990599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:43.990599Z digest=sha256:ea043c7a56cfab6a0447aae7aa4ecfbe0a2046f3f611ce359e44dbf285e7cb2c

Observation e7b2fe0a-3f59-41ca-9521-652d554e1f04 · outbound

This paper cites StreamingVLM: Real-Time Understanding for Infinite Video Streams.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos StreamingVLM: Real-Time Understanding for Infinite Video Streams

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.071908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.071908Z digest=sha256:f8eed95f4a4623cc9fd7c57254dd4331ad1f0b607682c1282a53164b74ffdc9b

Observation 9f512ed1-8155-4f28-bd60-74bb222a0f54 · outbound

This paper cites A fast and accurate one-stage approach to visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos A fast and accurate one-stage approach to visual grounding,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.133296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.133296Z digest=sha256:187f9bfe47bdbf7205f62b4bbc007c625c1886398619a50ec07070e86756485e

Observation dfd9d33c-78e5-40bc-b1bf-c29bba6295cf · outbound

This paper cites Improving one-stage visual grounding by recursive sub-query construction,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Improving one-stage visual grounding by recursive sub-query construction,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.252921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.252921Z digest=sha256:8f23eeab5bf2cb2181b6fe26cda74f4a4d9b01ec99672190b263c15a2aa860cc

Observation 994a86ff-27c6-4364-b2da-a43282072bfd · outbound

This paper cites Referring transformer: A one-step approach to multi-task visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Referring transformer: A one-step approach to multi-task visual grounding,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.355563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.355563Z digest=sha256:aaf3f4f26378201f6c244b542d8a0edb63b837b1361be25426bdd771e344da0f

Observation ed070ee0-2232-42c8-b209-4f4b6c51f151 · outbound

This paper cites Transvg: End-to-end visual grounding with transformers,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Transvg: End-to-end visual grounding with transformers,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.438359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.438359Z digest=sha256:edbe9cdfa648a407ac9ee89ce48c446e835aa53f0acefd1eabd610de754b55a9

Observation a3617427-5543-4c34-81cf-13884fbc05a6 · outbound

This paper cites Improving visual grounding with visual-linguistic verification and iterative reasoning,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Improving visual grounding with visual-linguistic verification and iterative reasoning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.518896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.518896Z digest=sha256:54a8d7ebdfb842b93cb6208f2b3a0c035e21a8f2b16aa11bfb1ee1b9b0d7fc9c

Observation 15da3c5a-576d-4eba-b291-7e1adf5e5b66 · outbound

This paper cites Seqtr: A simple yet universal network for visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Seqtr: A simple yet universal network for visual grounding,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.607850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.607850Z digest=sha256:ec4ab6d73564dac5695f1619cbb6876de046e6513d0fa72101fd5f0ca4da3e57

Observation 2c406c7c-c960-4212-b552-37e5b05d6790 · outbound

This paper cites Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.682676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.682676Z digest=sha256:c50387f8d24cad3637696da28bc65a3483953416b05a398734371c878bd960c3

Observation 8f6255a7-0317-45e2-98b1-3e35f30b41d7 · outbound

This paper cites A survivor in the era of large- scale pretraining: An empirical study of one-stage referring expression comprehension,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos A survivor in the era of large- scale pretraining: An empirical study of one-stage referring expression comprehension,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.745993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.745993Z digest=sha256:e10c688ddc43a8d438dc9e91a04534096e42bb2dcc99ef78df6e717d3dad2d1a

Observation 2bc27e84-0bb9-4c6c-9f13-e9a704e31956 · outbound

This paper cites Polyformer: Referring image segmentation as sequential polygon generation,.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Polyformer: Referring image segmentation as sequential polygon generation,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:44.814543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:44.814543Z digest=sha256:3f4c9d8b17be870469234601f24c4b629c2d721a129928c2ca04e17ffc192b98

Observation c2c19032-b299-4504-a5f6-cad0810c8554 · outbound

This paper cites Qwen2.5 Technical Report.

Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:42.188553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:42.188553Z digest=sha256:92bd837f940324c49f72a4fef277e7aa6e976a2802fb5124dac0d2fbf67a8a2a

Pith citing papers

No inbound Pith citation observations are available.