Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T11:33:44.814543Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.19857.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T11:33:44.814543Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9b995d4e-2191-4b15-b83f-b8fc2466296d · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos The small-drone revolution is coming—scientists need to ensure it will be safe,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49b07a8c-e707-4b7c-9a55-2c743b824f6b · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Champion-level drone racing using deep reinforcement learning,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 761e62ab-1b1c-4192-b68f-2dad832a8038 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Video object segmentation without tem- poral information,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c39917d-b99c-406e-b8f3-61d24b26056c · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos 3d question answering for city scene understanding,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43efa3e9-286a-4e1b-be86-a3197c84754d · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Phenobench: A large dataset and benchmarks for semantic image interpretation in the agricultural do- main,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01ecd4ed-e87b-4b06-8dec-7ad48ad22ef2 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Detecting flying objects using a single moving camera,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f4d291c-5418-4571-8780-b59da9a6df49 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Revisiting image-language networks for open-ended phrase detection,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b93844-3363-4b7a-aff3-0037efbdcd01 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Mevis: A large- scale benchmark for video segmentation with motion expressions,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5431214-f946-448e-974c-5886b8e16458 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Lamot: Language- guided multi-object tracking,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6822768-0774-409a-b3c4-642633a0a9f6 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Skyfind: A large-scale benchmark unveiling referring expres- sion comprehension for uav
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50c4b8f7-f820-4c1e-96ac-d23aae4cef2d · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Aerialmind: Towards referring multi-object tracking in uav scenarios,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6da0cf3b-dd9b-4492-acd3-8afc41de44ce · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Event-aware instructed assistant for referring video segmentation,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd07ae9b-fc66-4d12-97e5-44f84799e2fb · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos City-vlm: Towards multidomain perception scene understanding via multimodal incomplete learning,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71600a21-e282-4709-8a33-42faab77c5fb · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5 technical report,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 366e8549-1cbc-4afa-984b-b330d1f213b2 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SAM 3: Segment Anything with Concepts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 305d387f-576c-47d0-99f8-484771a6ffed · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos React: Streaming video analytics on the edge with asynchronous cloud support,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b29eab7-6a4d-4271-8344-f237d491cea0 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Multiple-object-tracking algorithm based on dense trajectory voting in aerial videos,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08fa439c-bdb6-4acf-8ed9-9f15f82ab321 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Urvos: Unified referring video object segmentation network with a large-scale benchmark,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b14fc791-78fa-4110-bcd4-66ae0c8a6cf8 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Sharegpt4video: Improving video understanding and generation with better captions,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79b21db3-821c-4694-8687-61f39cbdae15 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Mevis: A multi-modal dataset for referring motion expression video segmentation,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 844c35ca-a119-472a-8294-da0f114f5e6f · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Jtd-uav: Mllm-enhanced joint tracking and description framework for anti-uav systems,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e892caa-5ce7-4f56-bdb7-f50fa6c94cd4 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Visa: Reasoning video object segmentation via large lan- guage models,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed37ea5-009d-4755-b182-9cc0b34b6b93 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c53c5e-2219-4323-a01b-152283af7035 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Unipixel: Unified object referring and segmentation for pixel-level visual reason- ing,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 166f5840-37de-4732-b725-fca92a964b90 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SAM 2: Segment Anything in Images and Videos
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af657c0-610a-4099-a718-52299e96ebc4 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Videoglamm: A large multimodal model for pixel-level visual JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14 grounding in videos,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a0db379-3b2a-4df5-b8ca-796f9e2f9a06 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Glus: Global-local reasoning unified into a single large language model for video segmentation,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c1cf94f-b519-42a1-b153-da195082896f · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos The devil is in temporal token: High quality video reasoning segmentation,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fc75169-9d8b-4957-8153-a284702ec0b5 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Instructseg: Unifying instructed visual segmentation with multi-modal large language models,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df21c549-3a59-43ec-94da-f40429c71d43 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Geochat: Grounded large vision-language model for remote sensing,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2add77-efb9-4878-be89-fdcd1b629294 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ef00eab-20ef-407f-b9a5-a1423abd28ee · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Earthgpt: A universal multimodal large language model for multisensor image comprehension in remote sensing domain,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b7777f8-6659-4ffb-9d4d-397c763d51cd · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Kimi K2.5: Visual Agentic Intelligence
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beedfb6f-dfe7-4f2b-9f22-bf52ce0350a9 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen3.6-Plus: Towards real world agents,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 211f0754-f661-4860-b3d8-3eb118d94675 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Visual instruction tuning,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e040f5-7cb0-4289-9175-9dbffc0da9c2 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5-VL Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7b2fe0a-3f59-41ca-9521-652d554e1f04 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos StreamingVLM: Real-Time Understanding for Infinite Video Streams
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f512ed1-8155-4f28-bd60-74bb222a0f54 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos A fast and accurate one-stage approach to visual grounding,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfd9d33c-78e5-40bc-b1bf-c29bba6295cf · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Improving one-stage visual grounding by recursive sub-query construction,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994a86ff-27c6-4364-b2da-a43282072bfd · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Referring transformer: A one-step approach to multi-task visual grounding,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed070ee0-2232-42c8-b209-4f4b6c51f151 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Transvg: End-to-end visual grounding with transformers,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3617427-5543-4c34-81cf-13884fbc05a6 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Improving visual grounding with visual-linguistic verification and iterative reasoning,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15da3c5a-576d-4eba-b291-7e1adf5e5b66 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Seqtr: A simple yet universal network for visual grounding,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c406c7c-c960-4212-b552-37e5b05d6790 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f6255a7-0317-45e2-98b1-3e35f30b41d7 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos A survivor in the era of large- scale pretraining: An empirical study of one-stage referring expression comprehension,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bc27e84-0bb9-4c6c-9f13-e9a704e31956 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Polyformer: Referring image segmentation as sequential polygon generation,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c19032-b299-4504-a5f6-cad0810c8554 · outbound
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos Qwen2.5 Technical Report
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.